Multi-granularity semantic tree construction system and method for training distributed large models
By constructing a system with multi-granularity semantic trees and utilizing edge-cloud information fusion and model interaction, the problem of insufficient utilization of data granularity in distributed large model training is solved, achieving efficient and accurate heterogeneous model training and improving the model's performance on data of different granularities.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-16
- Publication Date
- 2026-04-03
AI Technical Summary
In existing technologies, the training of distributed large models lacks effective utilization of different data granularities, resulting in poor model performance on data of different granularities. Furthermore, traditional clustering methods struggle to handle the chaotic complexity of multidimensional data, leading to incomplete and uneven data granularity generation.
A multi-granularity semantic tree construction system is adopted. Through edge-cloud information fusion and model interaction, hypergraph clustering and multimodal feature fusion are used to generate data granularity suitable for different model sizes. Combined with Fedrated Learning, efficient and accurate distributed training is achieved.
It significantly enhances the ability of large models to guide small models, achieves efficient and accurate distributed training, improves the performance of models on data of different granularities, and solves the problems of incomplete and uneven data granularity generation.
Smart Images

Figure CN119494388B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of large model collaborative processing technology, and particularly relates to a multi-granularity semantic tree construction system and method for training distributed large models. Background Technology
[0002] With the rapid development of artificial intelligence technology, large models such as GPT4 and BERT have emerged, driving significant breakthroughs in various fields. However, large models face significant challenges in terms of energy consumption and edge deployment, limiting their application and development. To address this issue, collaborative heterogeneous distributed training of models of different sizes within a system has become necessary. To fully leverage the potential of this collaborative training, recent research has focused primarily on models and training, including lightweight models, enhanced quantization performance, and personalization. However, a crucial factor that has received less attention is the varying granularity of data.
[0003] In distributed training, data exhibits multiple granularities, meaning samples are similar but have different labels, leading to diversity in training data. Most studies use the same data granularity, ignoring the potential benefits of utilizing different granularities. In reality, models of different sizes perform differently on data of varying granularities. Smaller models, due to their limited capacity to handle complex tasks, show lower accuracy on fine-grained data but exhibit higher model and data utilization on coarse-grained data. Conversely, larger models fail to fully utilize the potential of coarse-grained data, wasting their inherent capabilities. These findings highlight the necessity of matching appropriate data granularity to different model scales, especially in the complex scenarios of distributed training in the era of large models. Therefore, exploring how to fully utilize data granularity may represent a new breakthrough in advancing accurate and efficient heterogeneous distributed training of models of different sizes.
[0004] Existing technologies have drawbacks:
[0005] 1. The fixed granularity of current datasets cannot meet the matching needs of various models ranging from billions of parameters to lightweight models. To maximize the benefits of distributed training and achieve high model utilization, suitable data granularity must be generated for different model sizes. Furthermore, due to the decentralized nature of the construction system, the challenges of large-scale data, and the privacy issues of edge-sensitive data, centralized data granularity generation methods are impractical.
[0006] 2. Due to a lack of comprehensive information, it is difficult to acquire all the knowledge needed to generate transformation relationships at different data granularities, resulting in incomplete generation. In a distributed architecture, a single device may only contain partial fine-grained data, limiting the construction of comprehensive data granularity and hindering the expansion of granularity for new data. Furthermore, using unimodal features for generation may obfuscate multi-granular hierarchical structures, as image features alone may struggle to distinguish visually similar but semantically different data. Therefore, fully utilizing the multimodal features of comprehensive data is crucial for generating complete and robust data granularity.
[0007] 3. The data granularity generation process based on clustering is significantly affected by the initial state. However, due to the chaotic complexity of multidimensional data information, obtaining an effective initial state is challenging. Integrating heterogeneous data from all edges is difficult, and the data itself is large-scale and high-dimensional, posing challenges to traditional clustering methods. The chaotic and complex information from the edges to the data makes it difficult to achieve stable and comprehensive initial spatial clustering. Summary of the Invention
[0008] To address the challenges of existing technologies, this invention provides a multi-granularity semantic tree construction system and method for training distributed large models. In this invention, appropriate data granularity is collaboratively generated in the construction system to meet the matching requirements of heterogeneous models. At the same time, by adopting a distributed mechanism for data and model interaction, this invention significantly enhances the ability of large models to guide small models, achieving efficient and accurate distributed training.
[0009] To address the problems in the existing technology, the present invention adopts the following technical solution.
[0010] A multi-granularity semantic tree construction system for distributed large-scale model training is provided. The system includes a cloud server and edge devices. The system comprises an edge-cloud information fusion space, a model information interaction space, an edge-cloud collaborative interaction space, a multi-granularity semantic tree, and a globally granular generation model.
[0011] The edge-cloud information fusion space embeds samples and text tags from cloud servers. and Collect edge devices in the form of Multimodal data fusion is performed using public dataset information;
[0012] The model information interaction space embeds samples and text labels from the cloud server. and use
[0013] Hypergraph clustering algorithm obtains global initial cluster centers O l ∈R d×k ;
[0014] The cloud server is based on the initial cluster center O. l Filtering words to construct a semantic space in WordNet
[0015] The cloud server will initialize cluster center O l Text embedding in semantic space Distribute to various edge devices to build an edge-cloud collaborative interaction space;
[0016] The cloud server will initialize cluster center O l The data is distributed to each edge device to initialize the granular generation model of each edge device.
[0017] The edge-cloud collaborative interaction space embeds text from the semantic space. This is distributed to each edge device to guide subsequent local learning and training of the edge device to build a global granular generation model, namely:
[0018]
[0019] Wherein: it maps the label from the l-th granularity to the (l+1)-th layer; This represents the sample embedding on edge device n. This represents the l-th layer text label embedding of a sample on edge device n. It is the (l+1)th layer text label of the sample on edge device n; φ n This refers to the parameters of the local granularity generation model deployed on edge device n to build a multi-granularity semantic tree.
[0020] Furthermore, the construction system consists of N edge devices and 1 cloud server, denoted as follows: and C;
[0021] The multi-granularity semantic tree employs a distributed, layer-by-layer generation of data granularities to construct an L-layer tree structure corresponding to models of different scales; that is... Where: the number of nodes in the l-th level tree is denoted as L. l ;
[0022] The edge device in the edge pair Each edge device, i.e. As a local sample dataset, where: M n ∈N + Let the set of text labels corresponding to the number of samples at the l-th granularity be denoted as . The edge device uses a sample encoder g(·) to generate sample embeddings. The edge device for each tag The corresponding text embedding is obtained using the text encoder z(·). in: For the edge The cloud server is for collecting a public dataset. Wherein: contains representative samples from each edge device; M C ∈N + express The number of samples in the sample; As a set of sample embeddings, The text embedding is used as the l-th level granularity label, where: This represents the text label corresponding to the sample in the l-th layer of the multi-granularity semantic tree.
[0023] Furthermore, the model information interaction space embeds samples and text labels from the cloud server.
[0024] and The global initial cluster centers O are obtained using the hypergraph clustering algorithm. l ∈R d×k The process includes:
[0025] Based on sample embedding By V v Compared to its most recent top-k I Connecting the neighboring vertices yields the sample hyperedge. Right now:,
[0026] Text hyperedge By using text embedding-based similarity at the current input l-th layer data granularity, V v Its top-k T The nearest neighbor vertex is connected; that is:
[0027] Apply the following formula to ε respectively I and ε T Normalize the weights; given the correlation matrix H, define
[0028] The degree is:
[0029]
[0030] Where, ∈[0,1] represents the weights of the balanced samples and text hyperedges;
[0031] The normalized Laplacian matrix of the hypergraph can be constructed using the following formula:
[0032]
[0033] Among them: Order and Let W represent the diagonal matrices for vertex degree and hyperedge degree, respectively; W = diag(w(E1), ..., w(E2)). |ε| )) represents the diagonal matrix of hyperedge weights;
[0034] From the l-th clustering layer to the (l+1)-th layer, spectral clustering is applied on Λ; and the feature vector is calculated. L corresponding to Λ l+1 The smallest non-zero eigenvalues;
[0035] These eigenvectors form the spectral embedding matrix.
[0036] Represents the low-dimensional encoding of the hypergraph structure;
[0037] right Applying K-means clustering, we obtain the Ll of the l-th layer. l+1 An initial cluster, denoted as .
[0038] Furthermore, the cloud server will initialize the cluster center O. l Distributed to each edge device to initialize the granular generation model of each edge device; including:
[0039] Based on the l-th layer of labels, construct the (l+1)-th layer of labels, using the currently input text labels of the l-th layer. Obtain the sample descriptor;
[0040] This set is expanded by searching for synonyms and hypernyms in WordNet, creating an initial semantic space.
[0041] Through the initial cluster centers O l Improve upon the foundation using the Faiss library Among them: based on For each cluster center, select the γ nouns that are closest to it, i.e.:
[0042]
[0043] Where: sim(·,·) is a similarity measure based on Euclidean distance; z(t) represents Semantic embedding;
[0044] The final semantic space of the l-th layer is formed by the union of these nearest neighbor words. Where: T is the total number of noun phrases selected from WordNet;
[0045] For each noun Construct a descriptive sentence template "a photo of a Obtain semantic embedding Furthermore, the edge-cloud collaborative interaction space embeds text. This information is distributed to various edge devices to guide the subsequent local learning and training process of building a global granular generative model; including:
[0046] A local multi-granularity semantic tree is trained using a consistency loss method with joint balanced regularization.
[0047] The upper layer of the multi-granularity semantic tree is trained using the FL framework;
[0048] Edge devices optimize the parameters in their granular generation model locally according to the following loss function and then transmit them to the cloud server;
[0049]
[0050] Where: η is the learning rate, which is the edge device's local model parameters at the end of each global iteration e. Uploaded to the cloud server;
[0051] The cloud server uses the FedAvg method to aggregate these parameters and builds a global granular generation model through E rounds of iterations, and updates the model parameters.
[0052]
[0053] in: Let n be the number of data items on the edge device. This represents the total amount of data across all edge devices.
[0054] The global granularity generation model takes the first-layer data as input and ultimately maps the l-th layer data to the (l+1)-th layer labels, i.e.
[0055] Enable interaction between heterogeneous edge device models.
[0056] Furthermore, the edge-cloud collaborative interaction space trains a local multi-granularity semantic tree process through a joint balanced regularization consistency loss, including:
[0057] Through the parameter Φ in the granularity generation model n Embedded u n,i Mapping to soft clustering assignment probability
[0058] Right now
[0059] Based on the relationship between sample embeddings and semantic embeddings, pseudo-labels for the sample set are generated after normalization.
[0060]
[0061] in: The c-th element is calculated as τ i,c / ∑ c τ i,c , indicating u n,i arrive The soft clustering probability of the mapping;
[0062] The semantic enhancement loss function is obtained based on the pseudo-labels according to the following formula;
[0063]
[0064] Where: CE(·) is the cross-entropy loss function; M n This represents the number of local samples for edge device n.
[0065] Indicates embedding u n,i To the semantic center Pseudo-labels formed by mapping Indicates embedding u n,i The soft cluster allocation result mapped by the local granularity generator;
[0066] The balanced regularization loss function is obtained based on the semantic enhancement loss function according to the following formula;
[0067]
[0068] Where α is a trade-off parameter.
[0069] Beneficial effects
[0070] This invention addresses the problems existing in the prior art:
[0071] 1. This invention addresses the distributed training of large and small models by collaboratively generating appropriate data granularity within the construction system to meet the matching requirements of heterogeneous models. By employing a distributed mechanism for data and model interaction, it significantly enhances the ability of large models to guide small models, achieving efficient and accurate distributed training.
[0072] 2. To overcome the problem of incomplete granular generation, this invention proposes a multi-dimensional information fusion strategy that combines image features and text labels from both the edge and cloud. This strategy enables the invention to comprehensively capture the deep semantics of images, providing rich guidance for granular generation models. Furthermore, this invention employs Optimal Transport (OT) theory to establish a soft pseudo-label mapping between images and semantics, and designs a semantically enhanced consistency learning model based on multimodal fusion. By utilizing Federated Learning (FL) methods, this invention achieves efficient collaborative model interaction between farmers across the edge and cloud, effectively integrating data information.
[0073] 3. To address the challenge of uneven granularity generation caused by the chaotic complexity of multimodal data, this invention introduces a hypergraph-based multimodal clustering method. The cloud collects and processes multimodal features from public data at the edge and uses Hypergraph Clustering (HGC) to capture complex high-order relationships within the data. This makes generating stable and comprehensive initial cluster centers a robust guide for the dynamic learning of the granularity generation model at the edge.
[0074] This invention validates the single-layer granularity order of the semantic tree generated from the baseline by matching it to a specific model. For example, this invention validates the performance of the first layer of the semantic tree (50 nodes on CIFAR100 and 100 nodes on TinyImageNet) on AlexNet.
[0075] like Figure 5 As shown, on CIFAR100 and TinyImageNet, the granular generation model exhibits significant accuracy across all layers, at 14.98% and 17.43%, respectively. Simultaneously, the accuracy density of Cultivator also shows a significant improvement, increasing by 12.47% on CIFAR100 and 11.49% on TinyImageNet. Compared to the centralized baseline, Cultivator (centralized) generates more appropriate data granularity for different layers and datasets by enhancing multimodal information fusion and semantic learning. However, Cultivator (centralized) is prone to overfitting because it uses all data and more model parameters for training.
[0076] When generating the first layer of granularity, the differences between classes may be blurred, leading to incomplete data and underfitting. As higher layers of granularity are generated, i.e., the number of classes decreases, the classes become easier to distinguish. In this case, due to overfitting, the centralized cultivator loses important general patterns, which it could have captured better with partial data. This demonstrates that the cultivator generates data granularity suitable for the model, which is beneficial for achieving efficient and accurate distributed training. Attached Figure Description
[0077] Figure 1 This invention features a construction system architecture with data and model interaction;
[0078] Figure 2 This is the initial HGC method of the present invention;
[0079] Figure 3 This is the semantic space determination process of the present invention;
[0080] Figure 4 This invention relates to the design of the multimodal consistency learning loss;
[0081] Figure 5 This invention compares the effectiveness of the generated granularity with baseline schemes. Detailed Implementation
[0082] The following is in conjunction with the appendix Figures 1-4 The present invention is described as follows:
[0083] I. System Overview
[0084] The objective of this invention is to achieve distributed construction of multi-granularity semantic trees within a construction system to adapt to heterogeneous edge environments. The construction system consists of N edge devices and one cloud server, denoted as follows: And C. This invention proposes a distributed construction method for multi-granularity semantic trees called Cultivator. Cultivator generates different data granularities layer by layer, corresponding to models of different scales, and constructs a hierarchical tree structure called MGTree; that is, a multi-granularity semantic tree.
[0085] 1) MGTree in the system: The specific architecture of the semantic tree constructed by the Cultivator is as follows: Figure 1 As shown at the bottom. This invention requires generating a multi-granularity semantic tree, called an MGtree with L layers, i.e. The number of nodes in the l-th level tree is denoted as L. l Without loss of generality, this invention assumes that the finest granularity of semantic tags constitutes the 0th layer of the MGtree. As the number of layers increases, the granularity of the semantic tags decreases.
[0086] 2) Edge: For For each edge device, the present invention defines As a local sample dataset, M n ∈N + Let be the number of samples. The set of text labels corresponding to the l-th granularity is denoted as . This invention uses a sample encoder g(·) to generate sample embeddings. For each tag This invention constructs a descriptive sentence (e.g., "a sheet of paper"). (The photo). Then, the corresponding text embedding is obtained using the text encoder z(·). in For the edge This invention defines a local granularity generation model. It maps labels from the l-th granularity to the (l+1)-th layer. A global granularity generation model, obtained by aggregating parameters from the local granularity generation model, interprets and processes the labels at the current layer, producing labels with higher semantic information. These generated labels will be used by upper-layer applications or other edge devices. Simultaneously, there exists a model specifically designed for sample-related tasks (such as image recognition), using model parameters θ. n express.
[0087] 3) Cloud server: The cloud server collects a public dataset. This contains representative samples from each edge device. To protect privacy, the edge devices only transmit samples and text embeddings. Let M... C ∈N + express The number of samples in the dataset. This invention defines... As a set of sample embeddings, As the text embedding of the l-th granularity label, where This represents the corresponding text label in MGtree.
[0088] II. Building the Interaction Process in the System
[0089] The distributed generation of multi-granularity semantic trees in the system involves data and model interaction, such as... Figure 1 As shown. Here, the present invention takes the construction of the data granularity of the first to second layers of MGTree as an example:
[0090] 1) Data interaction: The data interaction in the method of this invention includes the following two main parts:
[0091] ① Multimodal data fusion: Cloud servers embed samples and text labels. and Collection in the form of
[0092] Edge devices Use public dataset information to create an edge-cloud information fusion space.
[0093] ② Information interaction for guiding the model: In the edge cloud information fusion space, this invention uses a clustering algorithm.
[0094] Obtain global initial sample centers Where d is the embedding dimension and k is the number of clusters. Then, the cloud server constructs an appropriate semantic space. To capture the deep semantics of the samples. The cloud server will... Text embedding in semantic space Distribute to edge devices to establish an edge-cloud collaborative interaction space. This approach ensures that edge devices initialize their granular generative models (Cultivato) with collaborative multimodal information. r And optimize it under unified semantic guidance.
[0095] 2) Model Interaction: Model interaction is achieved through the following main components:
[0096] ③ Edge-Cloud Collaborative Learning: In the edge-cloud collaborative interaction space, this invention uses the FL framework for joint training.
[0097] The second layer is the granular generative model. Let E represent the total number of iterations in the global training. For each iteration e∈{1,…,E}, the edge devices locally optimize their granular generative model parameters Φ. n
[0098] The data is then transmitted to a cloud server, which aggregates these parameters using the FedAvg strategy. After E rounds of iterations, this invention obtains a global granularity generation model Φ. This granularity generation model integrates the granularity mining capabilities of distributed edge devices, mapping the first-layer data to the second-layer labels, i.e.
[0099] III. Hypergraph-based Multimodal Clustering
[0100] When computational efficiency is paramount and the data's feature structure is very simple, traditional K-means is a viable alternative to HGC. However, in real-world scenarios, many datasets are heterogeneous and high-dimensional, leading to complex and confusing information that is difficult to process. To handle complex multimodal relationships in large-scale datasets, this invention uses HGC on a cloud server, such as... Figure 2 As shown. This invention defines a weighted hypergraph. in Let be a set of vertices, where each vertex represents a common sample. It is a super-edge set that satisfies Super-edge E e The relevant vertices of ∈ε are A subset of the set whose vertices are associated with the hyperedge is equivalent to the set of the set. except, It is a weighting function that assigns positive real weights to the hyperedge. correlation matrix Defined by function h:
[0101]
[0102] Where: h v,e It is the element in the v-th row and e-th column of matrix H. This invention defines the hyperedge E. e The degree is This indicates the number of vertices it contains. This formula enables the present invention to model complex high-order relationships between samples, thereby promoting more effective clustering in multimodal spaces.
[0103] This invention uses a Gaussian kernel function to define the hyperedge E. e Superedge weights:
[0104]
[0105] In the formula: V i V j ∈E e Indicates the superedge E e All distinct vertex pairs within. σ e It is E e The median distance between all pairs of vertices in the equation. It is a distance metric for calculating vertex similarity. This formula ensures that hyperedges connecting closely related vertices receive higher weights, thus emphasizing strong local relationships in the multimodal feature space.
[0106] Based on the hypergraph described above This invention constructs two types of homogeneous hyperedges: the sample hyperedge set ε I and text hyper-edge set ε T To model the multimodal relationships between samples and their semantics. For each vertex Construction of the invention:
[0107] • Based on sample embedding The similarity, by V v Compared to its most recent top-k I Connecting the neighboring vertices yields the sample hyperedge. formal,
[0108] Text hyperedge By using text embedding-based similarity at the current input l-th layer data granularity, V v Its top-k T Connect the nearest neighbor vertices. The form is:
[0109] This invention addresses ε respectively I and ε T The weights are normalized. Given the incidence matrix H, define... The degree is:
[0110]
[0111] Among them, ∈[0,1] represents the weights of the balanced samples and text hyperedges.
[0112] make and These are diagonal matrices representing vertex degree and hyperedge degree, respectively.
[0113] W = diag(w(E1), ..., w(E) |ε| () represents the diagonal matrix of the hyperedge weights. This invention can define the normalized Laplacian matrix of the hypergraph as:
[0114]
[0115] To cluster from layer l to layer l+1, this invention uses spectral clustering on Λ. Specifically, this invention calculates feature vectors. L corresponding to Λ l+1 The smallest non-zero eigenvalues. These eigenvectors form the spectral embedding matrix. This represents the low-dimensional encoding of the hypergraph structure. Then, this invention... Applying K-means clustering, we obtain the Ll of the l-th layer. l+1 An initial cluster, denoted as .
[0116] 1) Semantic Space Determination: The cloud server constructs a semantic space within the edge-cloud information fusion space, facilitating sample-semantic mapping and restricting the semantic labels of upper-level tree nodes. The method of this invention utilizes a vocabulary database containing a wide range of word domains as a semantic dataset, such as WordNet, a vocabulary database containing over 82,000 nouns. Figure 3 As shown, in order to select the most relevant nouns from a given dataset, this invention proposes a two-step filtering process:
[0117] To construct the (l+1)th layer of labels based on the l-th layer of labels, this invention utilizes the currently input text labels of the l-th layer. As a valid sample descriptor, this invention utilizes WordNet.
[0118] Searching for synonyms and hypernyms in the search terminology expands this set, creating an initial semantic space.
[0119] The present invention focuses on the initial cluster center O. l Further improvements based on Exploit Faiss
[0120] Library, this invention is based on For each cluster center, select the γ nouns that are closest to it, i.e.:
[0121]
[0122] Where sim(·,·) is a similarity measure based on Euclidean distance. z(t) represents The semantic embedding is then used. The present invention then forms the final semantic space of the l-th layer based on the union of these nearest neighbor words, i.e. Where T is the total number of noun phrases selected from WordNet.
[0123] For each noun This invention constructs a descriptive sentence template "a photo of a “”, thus obtaining its semantic embedding 2) Multimodal Consistency Learning: In the edge-cloud collaborative learning process, to mine higher-level data granularity, the edge uses multimodal consistency learning to generate models at a local training granularity, such as... Figure 4 As shown. Furthermore, global granularity generation model parameters Φ are generated through edge-cloud collaborative training. For edge device n, parameters Φ are used... n The granularity generator embeds the sample into u n,i Mapping to soft clustering assignment probability Right now At this stage, only the network parameter Φ is updated. n Meanwhile, CLIP encoder g(·) and z(·) remain frozen.
[0124] a. Balanced Regularized Consistency Learning: To achieve better clustering, this invention introduces the manifold assumption, which states that if two samples lie in the local neighborhood of each other's low-dimensional manifold, they will have similar soft clustering assignments. Based on the manifold assumption, this invention defines the sample consistency loss as:
[0125]
[0126] Where <·,·> denote the dot product, and Nκ(x) n,i) is sample x n,i The set of κ nearest neighbors, i.e.
[0127] Furthermore, to avoid the granularity generation model assigning all samples to a single cluster, this invention adds a balanced regularization term and introduces the commonly used negative entropy loss. This invention defines the sample consistency loss with the balanced regularization term as:
[0128]
[0129] in To balance the regularization term. in Indicates sample x n,i The probability assigned to cluster c. β is a trade-off parameter for balancing regularization. Sample consistency loss. This indicates consistency between adjacent samples in the sample dimension. However, some of these samples may not belong to the same semantic cluster, which could lead to visually similar but semantically different clustering results, causing confusion.
[0130] b. Semantic Enhancement Consistency Learning: To address the limitations of sample-based consistency learning, this invention proposes a semantic space-based approach. This paper presents a semantically enhanced consistency learning method. This method leverages the relationship between sample embeddings and semantic embeddings to generate meaningful semantically enhanced pseudo-labels to supplement and guide the self-supervised clustering learning process.
[0131] Given a sample set middle In this invention, the soft cluster allocation results are obtained by first selecting the top-ξ samples for each cluster and storing the selection results as a binary value matrix. If a sample is selected As a cluster c∈{1, ..., L l+1 If the sample is B, then B n element B in row i and column c n,i,c B equals 1; otherwise B n,i,c =0. Sample center set Calculated as:
[0132] In order to obtain the semantic center set This invention will center the sample Map to it The latest semantic embeddings in China
[0133]
[0134] This invention will With sample embedding The mapping between them is expressed as an OT problem to generate pseudo-labels.
[0135] set up This is the cost matrix. In this invention, the transmission cost is defined as... Each semantic center and each embedding in The negative cosine similarity between samples, i.e.
[0136] Then, this invention solves the entropy-regularized OT problem:
[0137]
[0138] In the formula, ∏(a,b) is the set of effective transportation plans that satisfy the flow constraints, H(τ) is the entropy regularizer, and v>
[0139] 0 represents the regularization parameter. 1 k Let be a k-dimensional vector of size 1, where + indicates an edge constraint, which is generally uniformly distributed.
[0140] After obtaining the plan τ, normalize it to obtain the sample set x. n pseudo-tags The c-th element is calculated as τ i,c / ∑ c τ i,c , indicating u n,i arrive The soft clustering probability of the mapping. For the entropy-regularized OT problem defined above, this invention uses the Sinkhorn-Knopp algorithm for efficient solution. Based on pseudo-labels. This invention defines semantic enhancement in self-supervised learning.
[0141] loss:
[0142]
[0143] Where CE(·) is the cross-entropy loss function. The loss proposed in this term is used to align samples and their semantics.
[0144] Therefore, this invention defines a network based on balancing regularization consistency loss and semantic enhancement loss.
[0145] for:
[0146]
[0147] Here, α is a trade-off parameter. Before self-supervised training, this invention utilizes the initial HGC results to provide effective guidance for the initial state of unsupervised clustering, ensuring that the self-supervised process follows the correct trajectory and accelerating model convergence. The weights of the cultivated network are initialized to W in this invention. l =2ψO l , bias is Where ψ is the temperature parameter, O l This represents the initial HGC result for layer l.
[0148] Edge-Cloud Collaborative Learning: Each edge device trains a local model based on the multimodal consistency learning design described above. To integrate edge data and construct a comprehensive multi-granular semantic tree, this invention uses an FL framework in the granular generation model. The process is as follows: In each global iteration e∈{1,…,
[0149] In E}, each edge device n uses local data and Training parameters The goal of edge device n is to update parameters by optimizing the consistency loss function.
[0150]
[0151] In the formula, η is the learning rate, and g∈{1,…,G} is the local iteration round. At the end of each global iteration e, the edge device sets its local model parameters. Uploaded to the cloud. The cloud uses FedAvg policy.
[0152] These local parameters are aggregated to update the global model parameters:
[0153]
[0154] in Let n be the number of data items on the edge device. The total number of data points on all edges. Then, the updated global parameter φ is calculated. e+1 The data is distributed back to the edge as the initial state for the next local training. This iterative process continues until the global model parameters converge. Each global iteration integrates distributed edge device data, optimizes the global model, and progressively builds a robust MGTree.
[0155] Example:
[0156] This invention aims to use this technology to generate data with granularity that meets the matching needs of models of different sizes. Taking Cultivator, a distributed construction method for multi-granularity semantic trees, as an example, this invention compares it with five centralized clustering methods: SIC, SCAN, NNM, CC, and Cultivator (centralized), demonstrating the technical advantages of this solution. It shows that by utilizing multi-dimensional information fusion and cross-model interaction, as well as hypergraph-based clustering, Cultivator solves the challenges of incomplete and uneven generation.
[0157] Although the present invention has been described above, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many modifications under the guidance of the present invention without departing from the spirit of the present invention, and these modifications are all within the protection scope of the present invention.
Claims
1. A multi-granularity semantic tree construction system for training distributed large models, the system comprising cloud servers and edge devices; characterized in that: The construction system includes an edge-cloud information fusion space, a model information interaction space, an edge-cloud collaborative interaction space, a multi-granularity semantic tree, and a global granularity generation model; wherein: The construction system consists of An edge device and It consists of several cloud servers, denoted as follows: and ; The multi-granularity semantic tree is generated layer by layer in a distributed manner to correspond to the construction of models of different scales. A layered tree structure; that is , of which: The number of nodes in a layered tree is denoted as ; The edge device in the edge pair Each edge device, i.e. = As a local sample dataset, where: For the number of samples, in the th... The set of text tags corresponding to the layer granularity is denoted as = The edge device employs a sample encoder. Generate sample embeddings = The edge device for each tag Using a text encoder Obtain the corresponding text embedding = ,in: For the edge ; The cloud server is for collecting a public dataset. , where: contains representative samples from each edge device; express The number of samples in the sample; = As a set of sample embeddings, = As the first Text embedding of layer-level tags, where: = Represents the multi-granularity semantic tree. The text labels corresponding to the layer samples; The edge-cloud information fusion space embeds samples and text tags from cloud servers. and Collect edge devices in the form of Multimodal data fusion is performed using public dataset information; The model information interaction space embeds samples and text labels from the cloud server. and The hypergraph clustering algorithm is used to obtain the global initial cluster centers. ;include: Based on sample embedding By Compared to its recent top- Connecting the neighboring vertices yields the sample hyperedge. , Text hyperedge By entering the current number Similarity based on text embedding at the layer data granularity will Rather than top- The nearest neighbor vertex is connected; that is: ; According to the following formulas, respectively and Normalize the weights; given the correlation matrix ,definition The degree is: in, Balance the sample and text hyperedge weights; The normalized Laplacian matrix of the hypergraph can be constructed using the following formula: Among them: Order and These are diagonal matrices representing vertex degree and hyperedge degree, respectively. A diagonal matrix representing the weights of the hyperedges; From the Layer clustering to the first +1 layer, apply spectral clustering on Λ; and compute feature vectors. Corresponding to of The smallest non-zero eigenvalues; These eigenvectors form the spectral embedding matrix. , representing the low-dimensional encoding of the hypergraph structure; right Applying K-means clustering, we obtain the first... Layer An initial cluster, denoted as . ; The cloud server is based on the initial cluster center. Filtering words to construct a semantic space in WordNet The cloud server will initialize the cluster center. Text embedding in semantic space Distribute to various edge devices to build an edge-cloud collaborative interaction space; among which: The cloud server will initialize the cluster center. Distributed to each edge device to initialize the granular generation model of each edge device; including: In the Building the first layer based on the layer tags Layer label, using the current input layer Layer text labels Obtain the sample descriptor; This set is expanded by searching for synonyms and hypernyms in WordNet, creating an initial semantic space. Through initial cluster centers Improve upon the foundation using the Faiss library based on Select the best for each cluster center. One noun, namely: Where: sim( () is a similarity measure based on Euclidean distance; express semantic embedding; The union of these nearest neighbor words will form the first... The final semantic space of the layer, i.e. ,in: This represents the total number of noun phrases selected from WordNet. For each noun Construct a descriptive sentence template "a photo of a Obtain semantic embedding ; The edge-cloud collaborative interaction space embeds text from the semantic space. This is distributed to each edge device to guide subsequent local learning and training of the edge device to build a global granular generation model, namely: : Where: it will change the label from the first Layer granularity mapping to the first layer; Indicates edge device On the sample embedding, Indicates edge device The first sample Layered text tag embedding, It is an edge device The first sample Layer text labels; That is, edge devices The local granularity generation model parameters for constructing multi-granularity semantic trees are deployed on top; where: The edge-cloud collaborative interaction space embeds text. This information is distributed to various edge devices to guide the subsequent local learning and training process of building a global granular generative model; including: A local multi-granularity semantic tree is trained using a consistency loss method with joint balanced regularization. The upper layer of the multi-granularity semantic tree is trained using the FL framework; Edge devices optimize the parameters in their granular generation model locally according to the following loss function and then transmit them to the cloud server; in: The learning rate is used in each global iteration. At the end, the edge device sets its local model parameters Upload to the cloud server; The cloud server uses the FedAvg method to aggregate these parameters. The process iterates through rounds to build a global granular generation model and updates the model parameters. in: For edge devices The number of data points on the screen This represents the total amount of data across all edge devices. The global granularity generation model takes the first layer of data as input and ultimately generates the second layer. Layer data mapping to the first 1-layer label, i.e. Enable interaction between heterogeneous edge device models.
2. The multi-granularity semantic tree construction system for distributed large model training according to claim 1, characterized in that: The edge-cloud collaborative interaction space trains a local multi-granularity semantic tree through a joint balanced regularization consistency loss process, including: Through parameters in the granularity generation model embed Mapping to soft clustering assignment probability ,Right now ; Based on the relationship between sample embeddings and semantic embeddings, pseudo-labels for the sample set are generated after normalization. ; in: The Each element is calculated as ,express arrive The soft clustering probability of the mapping; The semantic enhancement loss function is obtained based on the pseudo-labels according to the following formula; in The cross-entropy loss function; Indicates embedding To the semantic center Pseudo-labels formed by mapping Indicates embedding The soft cluster allocation result mapped by the local granularity generator; The balanced regularization loss function is obtained based on the semantic enhancement loss function according to the following formula; in: It is a trade-off parameter.
3. A method for constructing a multi-granularity semantic tree system for training distributed large models, characterized in that: The method is based on the system implementation of any one of claims 1-2 to construct system data and model interaction.
Citation Information
Patent Citations
Global and local low-noise training method for guaranteeing edge computing data privacy
CN111475848A
Method based on multi-granularity federated learning
CN114154647A