Government affair SaaS architecture and method based on cloud edge collaboration

Through the cloud-edge collaboration government SaaS architecture, using scene semantic representation and federated learning technology, the problems of cross-scene knowledge integration and scene specialization services in government data analysis are solved, and the accuracy and resource efficiency of the model are improved under the premise of protecting privacy.

CN120598757APending Publication Date: 2025-09-05HEFEI WEIQINGLUO NETWORK TECH CO LTD

Patent Information

Application Number
CN202511116339.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

The existing government data analysis system cannot effectively integrate multi-scene knowledge while protecting data privacy, and cannot provide precise services for specific scenarios, resulting in poor performance of the model in specific scenarios.

Method used

Through the government affairs SaaS architecture based on cloud-edge collaboration, it uses scene semantic representation, similarity calculation and dynamic weight adjustment, combined with federated learning and knowledge distillation technology to achieve efficient integration and migration of cross-scene knowledge and optimize the target scenario model.

Benefits of technology

On the premise of protecting data privacy, the accuracy of the model in specific scenarios is improved, efficient integration and migration of multi-scene knowledge is realized, the consumption of model training resources is reduced, and the adaptability of zero-sample scenarios is demonstrated.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120598757A_ABST
    Figure CN120598757A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of government affair data analysis, and discloses a government affair SaaS architecture and method based on cloud edge collaboration, and the government affair data processing method based on cloud edge collaboration comprises the steps: obtaining government affair scene feature data, and generating scene semantic representation; the scene semantic similarity is calculated, and the knowledge migration potential is evaluated; analyzing semantic features of a target scene, and screening clients participating in training based on scene similarity; processing the client model, and dynamically adjusting the aggregation weight according to the scene similarity and the model contribution degree; global model knowledge is extracted, and a target scene local model is optimized; through organic combination of federated learning technologies of scene semantic understanding and scene perception, the limitation that data privacy protection and scene specialization model training cannot be realized at the same time in the traditional technology is overcome, and the technical problems of cross-scene knowledge integration and scene specialization service in government affair data analysis are effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of government data analysis technology, and more specifically, to a government SaaS architecture and method based on cloud-edge collaboration. Background Art

[0002] With the development and progress of science and technology, government data analysis has been widely used in application fields such as government service optimization, economic regulation and monitoring, social governance, and ecological and environmental protection, improving the efficiency of government services.

[0003] At present, government data not only has distribution differences among different departments, but also has significant differences in semantic characteristics in different scenarios, such as urban management, emergency response, and people's livelihood services.

[0004] Cloud-edge collaboration is the fusion of cloud computing and edge computing, aiming to achieve efficient data processing and transmission. Government data within this cloud-edge collaboration requires the use of federated learning systems to protect data privacy. While existing federated learning systems can protect data privacy, they lack the ability to understand contextual semantics, resulting in poor model performance in specific scenarios. While contextual understanding technologies can optimize model performance for a single scenario, they cannot integrate knowledge from multiple scenarios while protecting privacy. This technological fragmentation prevents government data analysis from fully leveraging the value of data across multiple scenarios and hinders the provision of precise services tailored to specific scenarios. Summary of the Invention

[0005] The present invention provides a government SaaS architecture and method based on cloud-edge collaboration to solve the limitations of traditional technologies in related technologies that cannot simultaneously achieve data privacy protection and scenario-specific model training, and the technical problems of cross-scenario knowledge integration and scenario-specific services in government data analysis.

[0006] The present invention provides a government data processing method based on cloud-edge collaboration, comprising the following steps:

[0007] Obtain government affairs scenario feature data and generate scenario semantic representation;

[0008] Calculate scene semantic similarity and evaluate knowledge transfer potential;

[0009] Analyze the semantic features of the target scene and select clients participating in the training based on scene similarity;

[0010] Process the client model and dynamically adjust the aggregation weight according to the scene similarity and model contribution;

[0011] Extract global model knowledge and optimize the local model of the target scene;

[0012] Among them, the scene semantic representation is generated by a feature extraction component, which includes a text feature extraction unit and a structured data processing unit, and is used to map the government scene to a unified semantic space.

[0013] Furthermore, the step of obtaining government affairs scene feature data and generating scene semantic representation specifically includes: collecting structured and unstructured data of government affairs scenes;

[0014] Preprocess the collected data;

[0015] Extract text features using pre-trained language models;

[0016] Combined with graph neural networks to process structured information;

[0017] Generate fixed-dimensional scene semantic vector representations.

[0018] Furthermore, the step of calculating scene semantic similarity specifically includes:

[0019] Select the vector similarity calculation method;

[0020] Normalize the scene feature vector;

[0021] Calculate the semantic similarity matrix between the target scene and other scenes;

[0022] Construct a scene affinity graph based on the similarity matrix.

[0023] Furthermore, the step of analyzing the semantic features of the target scene and screening the clients participating in the training specifically includes:

[0024] Obtain the semantic vector representation of the target scene;

[0025] For each candidate client, obtain its scene semantic vector;

[0026] Calculate the semantic similarity between the target scene and each candidate client scene;

[0027] Determine the client selection threshold or select the top k clients based on application requirements;

[0028] Filter clients that meet the conditions to form a set of clients participating in training.

[0029] Furthermore, the step of processing the client model and dynamically adjusting the aggregation weight according to the scene similarity and model contribution specifically includes:

[0030] Get local model updates from each participating client;

[0031] Calculate the contribution of each client model, taking into account factors such as data volume, model accuracy, and loss value;

[0032] For each client, calculate its aggregate weight;

[0033] Perform weighted aggregation using the calculated weights to generate a global model;

[0034] Distribute the generated global model to each client.

[0035] Furthermore, in the step of calculating the contribution of each client model, the contribution calculation formula is:

[0036] Contribution = first coefficient × normalized data volume + second coefficient × model accuracy + third coefficient × (1-normalized loss value);

[0037] The sum of the first coefficient, the second coefficient and the third coefficient is 1;

[0038] The first coefficient controls the weight of the normalized data volume in the contribution calculation. A larger data volume means more samples for model training, which generally provides more reliable model updates.

[0039] The second coefficient controls the weight of the model accuracy in the contribution calculation. A higher accuracy indicates better model performance and a greater contribution to the global model.

[0040] The third coefficient controls the weight of the normalized loss value in the contribution calculation. The lower the loss value, the better the model fit.

[0041] The specific selection of coefficients should be adjusted according to the needs and characteristics of the actual application scenario:

[0042] When data quality and quantity are primary considerations, the value of the first coefficient can be increased.

[0043] When model performance (accuracy) is the primary consideration, the value of the second coefficient can be increased.

[0044] When model fit (loss value) is the main consideration, you can increase the value of the third coefficient.

[0045] Furthermore, the steps of extracting global model knowledge and optimizing the local model of the target scene specifically include:

[0046] The global model is used as the teacher model and the local model of the target scene is used as the student model;

[0047] Prepare teacher-student model training samples based on the feature data of the target scene;

[0048] Calculate the distillation loss function;

[0049] Adjust the distillation temperature parameters to control the degree of knowledge transfer;

[0050] Iteratively optimize the student model parameters to minimize the distillation loss;

[0051] Deploy the optimized local model to the target scene.

[0052] Furthermore, the method further includes the steps of analyzing the semantic associations between scenes and constructing a scene semantic knowledge graph, which specifically includes:

[0053] Get the scene similarity matrix;

[0054] Set a similarity threshold to filter out scene pairs with significant similarity;

[0055] Execute community discovery algorithms to identify scene clusters;

[0056] Execute association rule mining algorithms to discover implicit associations between scenarios;

[0057] Build scene semantic knowledge graph based on community structure and association rules;

[0058] Generate visual representations to support exploration and knowledge sharing of scene contexts.

[0059] Furthermore, the step of executing the association rule mining algorithm includes:

[0060] Count the co-occurrence frequencies of scenes;

[0061] Calculate scene association confidence;

[0062] Filter high-confidence association rules.

[0063] The present invention provides a government SaaS architecture based on cloud-edge collaboration, which is used to execute the aforementioned government data processing method based on cloud-edge collaboration, including:

[0064] The scenario semantic representation module is used to obtain government scenario feature data and generate scenario semantic representation;

[0065] The scene similarity calculation module is used to calculate the semantic similarity of scenes and evaluate the potential of knowledge transfer;

[0066] The client selection module is used to analyze the semantic features of the target scene and select the clients participating in the training;

[0067] Model aggregation module, used to process client models and dynamically adjust aggregation weights;

[0068] Knowledge distillation module, used to extract knowledge from the global model and optimize the local model of the target scene;

[0069] The scene knowledge graph module is used to analyze the semantic associations between scenes and construct a scene semantic knowledge graph;

[0070] Among them, a data flow relationship is formed between each module, and the output of the previous module serves as the input of the next module.

[0071] The beneficial effects of the present invention are: by introducing a federated learning mechanism of scene semantic representation and scene perception, it solves the problem that government data cannot effectively integrate multi-scenario knowledge under the premise of protecting privacy, improves the accuracy of the model in specific scenarios, realizes the efficient integration and migration of multi-scenario knowledge, reduces the resource consumption of model training, demonstrates zero-sample scene adaptability, and automatically discovers potential connections between different government scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] Figure 1 It is a flow chart of the government data processing method based on cloud-edge collaboration of the present invention;

[0073] Figure 2 This is a two-dimensional visualization scatter plot of the semantic vectors of government affairs scenarios of the present invention, which intuitively shows the distribution and clustering of different government affairs scenarios in the semantic space, demonstrating the effectiveness of the semantic representation learning of scenarios in this patent;

[0074] Figure 3 This is the heat map of the similarity of government affairs scenarios in the present invention, which intuitively shows the semantic similarity relationship between different government affairs scenarios;

[0075] Figure 4 This is a line chart showing the changes in client aggregation weights during the federated learning iteration process of the present invention, which shows the dynamic change trend of the aggregation weights of different clients during the federated learning iteration process in the urban waterlogging prediction scenario.

[0076] Figure 5 This is a bar chart comparing resource consumption of different methods of the present invention, demonstrating the technical effect of this patent in significantly reducing resource consumption through scenario-specific knowledge distillation technology;

[0077] Figure 6 It is the urban government affairs scenario association network diagram of the present invention, which demonstrates the effect of the scenario semantic knowledge graph construction steps in this patent, intuitively presents the potential associations between scenarios, and provides data support for cross-departmental business collaboration. DETAILED DESCRIPTION

[0078] The subject matter described herein will now be discussed with reference to example embodiments. It should be understood that these embodiments are discussed solely to enable those skilled in the art to better understand and implement the subject matter described herein, and that the functions and arrangements of the elements discussed may be varied without departing from the scope of this specification. Various examples may omit, substitute, or add various processes or components as needed. Furthermore, features described in some examples may be combined in other examples.

[0079] Example 1

[0080] like Figure 1 As shown, this embodiment provides a government data processing method based on cloud-edge collaboration, including the following steps:

[0081] Step 100, obtain the characteristic data of government affairs scenarios and generate the semantic representation of the scenarios. This step obtains the characteristic data of government affairs scenarios and uses the scenario semantic representation learning model to map different government affairs scenarios (such as urban management, emergency response, and people's livelihood services) into the semantic space. Specifically, first, the structured and unstructured data of government affairs scenarios are collected, including business process descriptions, data field definitions, business rules and other information; then the collected data is preprocessed, including word segmentation, stop word removal, word frequency-inverse document frequency (TF-IDF) calculation of text data, as well as missing value filling, outlier processing and feature standardization of structured data; then the pre-trained language model is used to extract text features, and the structured information is processed in combination with the graph neural network; finally, a fixed-dimensional scene semantic vector representation is generated, which can effectively capture the core semantic features of government affairs scenarios. The scene semantic vector representation can be mathematically expressed as:

[0082]

[0083] in, Representation scene The semantic vector representation of Representation scene The characteristic data, represents the semantic encoding function, which can be expressed as:

[0084]

[0085] in, represents the text encoding function, Text data representing the processing scenario; represents the graph structure encoding function, Structured data representing the processing scenario; Represents the weight matrix in semantic encoding; Represents the bias term, which is used for linear transformation in neural networks; Represents nonlinear activation functions, such as ReLU, sigmoid, etc.; Represents the feature concatenation operation, which connects feature vectors from different sources into one vector.

[0086] Furthermore, the scene semantic representation learning model in this embodiment includes the following components:

[0087] Text feature extraction component: This component includes a text encoding unit, a context understanding unit, and a feature fusion unit, and is used to process unstructured text data in government scenarios.

[0088] Structured data processing component: includes graph structure encoding unit, node representation unit and edge representation unit, used to process structured data in government scenarios;

[0089] Multimodal feature fusion component: This includes a feature alignment unit, a weight calculation unit, and a fusion transformation unit, which is used to fuse the feature representations of different types of data into a unified vector space.

[0090] Semantic vector generation component: Contains a dimensionality reduction unit, a normalization unit, and an output unit, and is used to generate a fixed-dimensional scene semantic vector representation.

[0091] Step 200: Calculate the semantic similarity of scenarios and evaluate the potential for knowledge transfer. This step is based on the scenario semantic representation generated in step 1 and calculates the semantic similarity between different government scenarios to evaluate the potential for knowledge transfer and the risk of conflict between scenarios. Specifically, first select an appropriate vector similarity calculation method (such as cosine similarity, Euclidean distance, etc.); then normalize the scenario feature vector to eliminate the impact of dimensionality differences. The specific normalization method uses L2 norm normalization, which scales the vector to a unit vector. , ensuring that features of different dimensions contribute evenly to similarity calculations; then calculating the semantic similarity matrix between the target scene and other scenes; finally, constructing a scene affinity graph based on the similarity matrix to guide subsequent model training and knowledge transfer. Scene similarity calculation can be expressed as:

[0092]

[0093] in, Representation scene With scene The similarity between ; and Represents scenes respectively and scenes Semantic vector of Represents a vector and The dot product is calculated as:

[0094]

[0095] in is the vector dimension, Represents a vector No. A portion.

[0096] Step 300: Analyze the semantic features of the target scene and select clients for training. This step prioritizes clients with similar scenes for training based on the semantic features of the target scene using a scenario-aware federated learning client selection algorithm. The scenario-aware federated learning client selection algorithm is executed as follows:

[0097] Step 301: Obtain the semantic vector representation of the target scene ;

[0098] Step 302: For each candidate client , get its scene semantic vector ;

[0099] Step 303: Calculate the semantic similarity between the target scene and each candidate client scene. ;

[0100] Step 304: Determine the client selection threshold based on application requirements. Or select the top k clients;

[0101] Step 305: Filter clients that meet the conditions to form a client set participating in training.

[0102] The client selection result can be expressed as:

[0103]

[0104] in, Indicates the selected client set, Represents the semantic vector of the target scene, Represents the client The scene semantic vector, Indicates the client selection threshold, used to filter clients whose similarity meets the conditions.

[0105] Furthermore, the scenario-aware federated learning client selection algorithm also includes a dynamic threshold adjustment step, which adaptively adjusts the selection threshold based on system performance feedback and scenario changes. , to optimize the number and quality of clients participating in training.

[0106] Step 400: Process the client model and dynamically adjust the aggregation weight. This step executes a scene-aware model aggregation weight dynamic adjustment algorithm based on scene semantic similarity and model contribution. The scene-aware model aggregation weight dynamic adjustment algorithm is executed as follows:

[0107] Step 401: Get local model updates from each participating client ;

[0108] Step 402: Obtain the semantic similarity between the target scene and each client scene from step 2. ;

[0109] Step 403: Calculate the contribution of each client model ,Considerations include data volume, model accuracy, etc.;

[0110]

[0111] in, Represents the client The normalized data volume is calculated as follows:

[0112]

[0113] in, and Client and the original data volume of j, Indicates the maximum amount of data among all clients. For client index, Indicates from 1 to The integer range of n represents the total number of clients; Indicates the accuracy of the model on the validation set, with a value range of ; Represents the client The normalized loss value of , calculated as:

[0114]

[0115] in and Respectively represent the maximum and minimum loss values ​​of all clients, Represents the original loss value of the model of client i on the validation set; 、 and Is the weight coefficient for contribution calculation, satisfying ;

[0116] Step 404, for each client , calculate its aggregation weight, the calculation formula is:

[0117]

[0118] in, is the sum variable, n represents the total number of clients, represents the semantic similarity between the target scene and the j-th candidate client scene, Indicates the contribution of the j-th client model;

[0119] Step 405: Perform weighted aggregation using the calculated weights to generate a global model:

[0120]

[0121] in, Represents the global model, which is obtained by weighted aggregation of client models; Represents the client The aggregation weight is calculated based on scene similarity and model contribution;

[0122] Step 406: Generate the global model Distribute to each client for the next round of training.

[0123] Furthermore, the global model and client-side local models Both contain the following components:

[0124] Feature extraction component: responsible for extracting effective features from raw data;

[0125] Representation learning component: converts the extracted features into a representation suitable for model processing;

[0126] Decision component: Make predictions or classifications based on the learned representations;

[0127] Adaptation component: adjusts model output according to specific scenarios.

[0128] Step 500: Extract global model knowledge and optimize the local model of the target scene. This step uses a scene-specific knowledge distillation module to extract knowledge related to the target scene from the global model to assist in local model optimization. The knowledge distillation process is performed according to the following process:

[0129] Step 501: Use the global model as the teacher model and the local model of the target scene as the student model;

[0130] Step 502: Prepare teacher-student model training samples based on the target scene’s feature data and perform preprocessing on the sample data, including feature normalization (converting each feature to a standard normal distribution with a mean of 0 and a standard deviation of 1), category encoding (converting discrete category features to one-hot encoding or embedding vectors), and data augmentation (expanding the training set by adding small perturbations or generating synthetic samples).

[0131] Step 503: Calculate the distillation loss function to guide the student model to learn the decision boundary and probability distribution of the teacher model;

[0132] Step 504, controlling the degree of knowledge transfer by adjusting the distillation temperature parameter;

[0133] Step 505, iteratively optimize the student model parameters to minimize the distillation loss;

[0134] Step 506: deploy the optimized local model to the target scene.

[0135] The loss function in the knowledge distillation process can be expressed as:

[0136]

[0137] in, represents the total distillation loss, Represents the cross entropy loss function, which is calculated as:

[0138]

[0139] in, Represents the one-hot encoding of the true label, that is, the standard answer; Indicates the true label one-hot encoding elements, the value is 0 or 1; Indicates the number of categories, that is, the total number of categories in the classification task; Represents the natural logarithm function, the default is natural logarithm For the bottom; Represents the KL divergence loss function, which is calculated as:

[0140]

[0141] and Represent the output of the teacher model and the student model, namely the logits value; Represents the teacher model’s response to the first The predicted probability of each category; The student model after temperature scaling is The predicted probability of each category; Represents the softmax function, and the calculation formula is:

[0142]

[0143] in, represents the exponential function with natural logarithm as base, that is of power, where Approximately 2.71828; can also be written as ; and Respectively represent Output values ​​of the jth and jth categories; Represents the temperature parameter, which controls the smoothness of the output distribution. Values ​​will produce smoother probability distributions; and It is the distillation loss balance weight, which is used to adjust the ratio of cross entropy loss and KL divergence loss.

[0144] Furthermore, the scenario-specific knowledge distillation module contains the following components:

[0145] Model interface component: responsible for connecting the teacher model and the student model and unifying the input and output formats;

[0146] Feature selection component: Filters scene-related features and knowledge based on the characteristics of the target scene;

[0147] Distillation loss calculation component: includes cross entropy loss calculation unit and KL divergence loss calculation unit;

[0148] Temperature regulation component: dynamically adjusts the softness or hardness of knowledge transfer;

[0149] Model optimization component: updates the student model parameters based on the distillation loss.

[0150] Furthermore, this embodiment also includes the step of constructing a scene semantic knowledge graph:

[0151] Step 600: Analyze the semantic associations between scenarios and construct a scenario semantic knowledge graph. This step automatically discovers and constructs a semantic association network between different government scenarios using long-term accumulated scenario similarity data. The scenario semantic knowledge graph construction algorithm is executed as follows:

[0152] Step 601, obtaining the scene similarity matrix calculated in step 200;

[0153] Step 602: Set a similarity threshold, select scene pairs with significant similarity, and construct the initial structure of the scene semantic graph;

[0154] Step 603: executing a community discovery algorithm, including:

[0155] Initialize each scene as an independent community;

[0156] Calculate the similarity between communities using the following formula:

[0157]

[0158] in, and Represents two different communities, and Respectively represent communities and The number of inner scene nodes, is the inter-scene similarity calculated in step 200;

[0159] Merge the community pairs with the highest similarity;

[0160] Iterate until the termination condition is met.

[0161] Step 604, executing an association rule mining algorithm, includes:

[0162] Statistical scene co-occurrence frequency, that is, calculation scene and scenes Frequency of co-occurrence in similar scenes ;

[0163] The original frequency data can be further logarithmically transformed to reduce the impact of long-tail distribution. The conversion formula is: , represents the frequency data after transformation;

[0164] Calculate the scene association confidence, the calculation formula is:

[0165]

[0166] in, Indicates that from the scene Deriving the scenario The confidence level, Representation scene Frequency of occurrence; in order to balance the differences in frequency of occurrence in different scenarios, the adjusted confidence index is also calculated , taking both confidence and support into account;

[0167] Filter high-confidence association rules, that is, select association rules with confidence greater than the preset threshold The scene of Form association rules.

[0168] Step 605: constructing a scene semantic knowledge graph based on the community structure and association rules;

[0169] Step 606 , generate a visual representation to support exploration of scene associations and knowledge sharing.

[0170] The scene association in the knowledge graph can be expressed as:

[0171]

[0172] in, Represents a scene knowledge graph, consisting of a set of nodes , edge set and weights composition; Represents a collection of scene nodes, All scene nodes in; Represents the edge set between scenes, indicating the association relationship between scenes; express The weight set of the middle edge represents the similarity between scenes.

[0173] Furthermore, the scene semantic knowledge graph contains the following components:

[0174] Scene node component: stores the semantic features and attribute information of the scene;

[0175] Relationship edge component: represents the type and strength of semantic associations between different scenarios;

[0176] Community structure component: identifies clusters of scenarios with close connections;

[0177] Graph index component: accelerates graph query and scene retrieval;

[0178] Visualization component: Converts complex graph structures into intuitive visual representations.

[0179] The following is an example of an application of the present invention:

[0180] In this example, this implementation method is applied to the multi-department collaborative governance system of a provincial capital city, involving three government affairs scenarios: urban management department, traffic management department, and emergency management department.

[0181] like Figure 2-6 As shown, the implementation process is as follows:

[0182] Scenario Semantic Representation Construction: Structured and unstructured data for each scenario was collected from the business systems of the three departments, including business process descriptions, data dictionaries, and historical disposal cases. The collected data was then preprocessed, including word segmentation, stop word removal, term frequency-inverse document frequency (TF-IDF) calculation, and missing value filling and feature standardization for structured data. The following table shows sample feature data for each scenario:

[0183] Table 1: Feature data samples of three government affairs scenarios

[0184]

[0185] Next, we use the pre-trained language model to extract text features and use a graph neural network to process structured information. After processing it through the scene semantic representation learning model, we obtain a fixed-dimensional scene semantic vector representation:

[0186] Table 2: Scene semantic vector representation results (truncated some dimensions)

[0187]

[0188] Scene similarity calculation and client selection: Based on semantic vector representation, we first selected cosine similarity as the vector similarity calculation method. Then, we performed L2 norm normalization on the scene feature vectors, scaling the vectors to unit vectors to ensure that features of different dimensions contribute evenly to the similarity calculation. We then calculated the similarity between each scene and constructed a scene similarity matrix:

[0189] Table 3: Scene similarity matrix

[0190]

[0191] In practical applications, using "urban waterlogging analysis and prediction" as the target scenario, the system employed a scenario-aware federated learning client selection algorithm. The algorithm first acquired a semantic vector representation of the target scenario, then calculated the semantic similarity between the target scenario and each candidate client scenario. Finally, the city management department and the traffic management department were selected as clients for training because their semantic similarity with the target scenario exceeded a set threshold of 0.6. The system also dynamically adjusted the selection threshold to optimize the number and quality of participating clients based on the changing characteristics of different scenarios.

[0192] Model aggregation and knowledge distillation: During multiple rounds of federated learning iterations, the system first obtains the local model updates of each participating client. , then executes a scenario-aware dynamic adjustment algorithm for model aggregation weights based on scenario similarity and model contribution. The system considers factors such as the client's data volume, model accuracy, and loss value to calculate contribution, and then calculates the aggregation weight based on scenario similarity and model contribution. The following table shows the aggregation weight calculation results for the 10th iteration:

[0193] Table 4: Model aggregation weight calculation for the 10th iteration

[0194]

[0195] Using the calculated weights to perform weighted aggregation, the system generates a global model , and distributes it to each client for the next round of training. Then, through the scenario-specific knowledge distillation module, the system extracts knowledge related to urban flooding scenarios from the global model and optimizes the local model of the target scenario. Specifically, the system uses the global model as the teacher model and the local model of the target scenario as the student model. It prepares and preprocesses the teacher-student model training samples, and calculates the distillation loss function to guide the student model learning. The knowledge distillation process uses a temperature parameter τ=2.5, weight coefficients α=0.3 and β=0.7, so that the student model maintains classification accuracy and learns the decision boundary of the teacher model at the same time. Finally, the distillation loss is minimized by iteratively optimizing the student model parameters.

[0196] Construction of a scene semantic knowledge graph: In addition to the above process, the system also automatically discovers and constructs a semantic association network between different government scenarios through long-term accumulation of scene similarity data. Specifically, the system sets a similarity threshold of 0.6, screens scene pairs with significant similarity, and constructs the initial structure of the scene semantic graph. It then executes a community discovery algorithm, initializing each scene as an independent community, calculating the similarity between communities, merging the community pairs with the highest similarity, and iterating until the termination condition is met. Next, the association rule mining algorithm is executed to count the co-occurrence frequency of scenes, calculate the confidence level of scene associations, and screen high-confidence association rules. Finally, based on the community structure and association rules, a scene semantic knowledge graph is constructed, and a visual representation is generated to support the exploration of scene associations and knowledge sharing.

[0197] The following table shows some of the scenario association rules mined by the system:

[0198] Table 5: Results of scene association rule mining (partial)

[0199]

[0200] It is understandable that the data preprocessing methods known to those skilled in the art include data cleaning, data conversion, and data reduction, among which data conversion includes type conversion and normalization and standardization. Although the dimension and type of the data are ignored in the description of the previous embodiment, data preprocessing is technical knowledge known to those skilled in the art and a prerequisite for data processing. Therefore, the well-known data preprocessing steps are not independently described in the previous content.

[0201] The above describes an embodiment of the present invention, but this embodiment is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Ordinary technicians in this field can also make more forms of equivalent embodiments based on the inspiration of this embodiment, all of which are protected by this embodiment.

Claims

1. A government data processing method based on cloud-edge collaboration, characterized in that: The following steps are involved: Obtain government affairs scenario feature data and generate scenario semantic representation; Calculate scene semantic similarity and evaluate knowledge transfer potential; Analyze the semantic features of the target scene and select clients participating in the training based on scene similarity; Process the client model and dynamically adjust the aggregation weight according to the scene similarity and model contribution; Extract global model knowledge and optimize the local model of the target scene; Among them, the scene semantic representation is generated by a feature extraction component, which includes a text feature extraction unit and a structured data processing unit, and is used to map the government scene to a unified semantic space.

2. The method according to claim 1, characterized in that The steps of obtaining government affairs scenario feature data and generating scenario semantic representation specifically include: Collect structured and unstructured data in government affairs scenarios; Preprocess the collected data; Extract text features using pre-trained language models; Combined with graph neural networks to process structured information; Generate fixed-dimensional scene semantic vector representations.

3. The method according to claim 1, characterized in that The step of calculating scene semantic similarity specifically includes: Select the vector similarity calculation method; Normalize the scene feature vector; Calculate the semantic similarity matrix between the target scene and other scenes; Construct a scene affinity graph based on the similarity matrix.

4. The method according to claim 1, wherein The steps of analyzing the semantic features of the target scene and selecting the clients to participate in the training specifically include: Obtain the semantic vector representation of the target scene; For each candidate client, obtain its scene semantic vector; Calculate the semantic similarity between the target scene and each candidate client scene; Determine the client selection threshold or select the top k clients based on application requirements; Filter clients that meet the conditions to form a set of clients participating in training.

5. The method according to claim 1, characterized in that The step of processing the client model and dynamically adjusting the aggregation weight according to the scene similarity and model contribution specifically includes: Get local model updates from each participating client; Calculate the contribution of each client model, taking into account factors such as data volume, model accuracy, and loss value; For each client, calculate its aggregate weight; Perform weighted aggregation using the calculated weights to generate a global model; Distribute the generated global model to each client.

6. The method according to claim 5, characterized in that In the step of calculating the contribution of each client model, the contribution calculation formula is: Contribution = first coefficient × normalized data volume + second coefficient × model accuracy + third coefficient × (1-normalized loss value); The sum of the first coefficient, the second coefficient and the third coefficient is 1; The first coefficient: controls the weight of the normalized data volume in the contribution calculation; The larger the amount of data, the more samples the model can be trained on, which can provide more reliable model updates; The second coefficient controls the weight of the model accuracy in the contribution calculation; the higher the accuracy, the better the model performance and the greater the contribution to the global model; The third coefficient: controls the weight of the normalized loss value in the contribution calculation; The lower the loss value, the better the model fit; The selection of the first coefficient, the second coefficient and the third coefficient is adjusted according to the needs: When data quality and quantity are the main considerations, increase the value of the first coefficient; When model performance is the primary consideration, increase the value of the second coefficient; When model fit is the primary consideration, increase the value of the third coefficient.

7. The method according to claim 1, characterized in that The steps of extracting global model knowledge and optimizing the local model of the target scene specifically include: The global model is used as the teacher model and the local model of the target scene is used as the student model; Prepare teacher-student model training samples based on the feature data of the target scene; calculate the distillation loss function; Adjust the distillation temperature parameters to control the degree of knowledge transfer; Iteratively optimize the student model parameters to minimize the distillation loss; Deploy the optimized local model to the target scene.

8. The method according to claim 1, characterized in that The method also includes analyzing the semantic associations between scenes and constructing a scene semantic knowledge graph, which specifically includes: Get the scene similarity matrix; Set a similarity threshold to filter out scene pairs with significant similarity; Execute community discovery algorithms to identify scene clusters; Execute association rule mining algorithms to discover implicit associations between scenarios; Build scene semantic knowledge graph based on community structure and association rules; Generate visual representations to support exploration and knowledge sharing of scene contexts.

9. The method according to claim 8, characterized in that The steps of executing the association rule mining algorithm include: Count the co-occurrence frequencies of scenes; Calculate scene association confidence; Filter high-confidence association rules.

10. A government SaaS architecture based on cloud-edge collaboration, used to execute a government data processing method based on cloud-edge collaboration as described in any one of claims 1-9, characterized in that: include: The scenario semantic representation module is used to obtain government scenario feature data and generate scenario semantic representation; The scene similarity calculation module is used to calculate the semantic similarity of scenes and evaluate the potential of knowledge transfer; The client selection module is used to analyze the semantic features of the target scene and select the clients participating in the training; Model aggregation module, used to process client models and dynamically adjust aggregation weights; Knowledge distillation module, used to extract knowledge from the global model and optimize the local model of the target scene; The scene knowledge graph module is used to analyze the semantic associations between scenes and construct a scene semantic knowledge graph; Among them, a data flow relationship is formed between each module, and the output of the previous module serves as the input of the next module.

Citation Information

Patent Citations

  • Specific scene model upgrading method and system based on federated learning

    CN111444848A

  • Federated learning modeling optimization method and device, medium and computer program product

    CN113095512A

  • Federal learning fairness improvement method for medical data heterogeneous scene

    CN117764199A

  • Privacy protection contribution evaluation method in horizontal federated learning scene

    CN118114296A

  • Personalized federal learning method based on self-adaptive local model initialization and double knowledge distillation

    CN119514727A

Cited By

  • Three-dimensional modeling water surface cavity filling method and system based on federal learning

    CN122156498A

  • Three-dimensional modeling water surface hole filling method and system based on federated learning

    CN122156498B