Large model analysis method and system for multi-source heterogeneous data fusion
Through edge-cloud collaborative architecture and lightweight model technology, the problems of data silos, privacy leakage and computing bottlenecks in multi-source heterogeneous data fusion are solved, real-time data analysis and decision support at the edge are realized, and the adaptability of the model and the reliability of decision-making are improved.
Patent Information
- Application Number
- CN202510778069.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-09-26
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In existing technologies, the fusion of multi-source heterogeneous data faces the risks of data silos and privacy leakage, computing and storage bottlenecks, and model complexity and deployment difficulties, making it difficult to achieve real-time analysis on edge devices.
By adopting an edge-cloud collaborative architecture and combining trusted data space technology and intelligent agent decision-making technology, a lightweight spatiotemporal sequence model is constructed. Through technologies such as model pruning, knowledge distillation, small sample learning, and federated learning, data fusion analysis is performed at the edge to ensure data privacy and computing efficiency.
It realizes real-time multi-source heterogeneous data fusion at the edge, meets real-time requirements, improves the adaptability and generalization ability of the model, and ensures data security and decision-making reliability.
Smart Images

Figure CN120705799A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of data analysis technology, and specifically relates to a large-scale model analysis method and system for multi-source heterogeneous data fusion. Background Art
[0002] With the rapid development of the Internet of Things (IoT), mobile internet, and sensor technology, data is being generated at an unprecedented rate and scale. This data is often multi-source, heterogeneous, and highly real-time. Sources may include sensors, mobile devices, social media, and enterprise systems, and its formats and semantics can vary widely. Effectively integrating this multi-source, heterogeneous data, extracting valuable information, and enabling rapid, accurate, and real-time decision-making are key challenges facing big data applications.
[0003] The existing technology has the following defects:
[0004] 1) Data silos and privacy leakage risks: Data from different sources are often stored in isolated systems, making them difficult to share. Directly transmitting raw data to the center for integration may lead to privacy leakage and data security risks.
[0005] 2) Computing and storage bottlenecks: Large-scale, high-dimensional heterogeneous data fusion analysis requires enormous computing resources and storage space. Centralizing all data for processing in the cloud may result in excessive latency and fail to meet real-time requirements.
[0006] 3) Model complexity and deployment difficulties: Traditional complex models are difficult to deploy and run on resource-constrained edge devices, limiting the application scope of real-time analysis. Summary of the Invention
[0007] In order to solve the problems of data silos and privacy leakage risks, computing and storage bottlenecks, and model complexity and deployment difficulties in the existing technology, the purpose of the present invention is to provide a large-scale model analysis method and system for multi-source heterogeneous data fusion.
[0008] The technical solution adopted in the present invention is:
[0009] A large-scale model analysis method for multi-source heterogeneous data fusion includes the following steps:
[0010] In the cloud data center of the edge-cloud collaborative architecture, we use trusted data space technology and intelligent agent decision-making technology to build a large spatiotemporal sequence model;
[0011] At the edge of the edge-cloud collaborative architecture, the spatiotemporal sequence large model is deployed in a lightweight manner to obtain a lightweight spatiotemporal sequence large model;
[0012] Use a lightweight spatiotemporal sequence model to fuse and analyze real-time multi-source heterogeneous data collected at the edge to obtain real-time application scenario decisions.
[0013] Furthermore, the spatiotemporal sequence model includes a standardized access and quality verification layer, a cross-domain entity association layer, a fusion analysis layer, and an intelligent decision-making layer, which are connected in sequence;
[0014] The standardized access and quality verification layer includes a data isolation module, a data evidence audit module, a blockchain rights confirmation module, and a data source filtering module;
[0015] The cross-domain entity association layer includes domain knowledge graph and semantic mapping modules;
[0016] The fusion analysis layer includes multi-source homogeneous data fusion model and fusion data analysis model;
[0017] The intelligent decision-making layer includes decision-making agents and post-verification modules.
[0018] Furthermore, in the cloud data center of the edge-cloud collaborative architecture, we use trusted data space technology and intelligent agent decision-making technology to build a large spatiotemporal sequence model, including the following steps:
[0019] Distribute connections between several edge devices and cloud data centers to build an edge-cloud collaborative architecture;
[0020] In the cloud data center of the edge-cloud collaborative architecture, we use trusted data space technology and intelligent agent decision-making technology to build a large-scale spatiotemporal sequence model architecture;
[0021] Using a privacy-enhanced fusion algorithm, the spatiotemporal sequence large model architecture is improved to obtain an improved spatiotemporal sequence large model architecture;
[0022] The optimization algorithm is used to optimize the initial large model parameters of the improved spatiotemporal sequence large model architecture to obtain the spatiotemporal sequence large model.
[0023] Furthermore, at the edge of the edge-cloud collaborative architecture, a lightweight spatiotemporal sequence model is deployed to obtain a lightweight spatiotemporal sequence model, including the following steps:
[0024] Extract the large model metadata of the spatiotemporal sequence large model in the cloud data center, and use model pruning technology to structurally prune the large model metadata to obtain lightweight large model metadata, and send it to all edge terminals of the edge-cloud collaborative architecture;
[0025] Based on the metadata of the lightweight large model, the spatiotemporal sequence large model is lightweight deployed at the edge of the edge-cloud collaborative architecture to obtain the initial lightweight spatiotemporal sequence large model.
[0026] At each edge, an initial training dataset is collected and dynamically optimized using knowledge distillation and small sample learning mechanisms to obtain an optimized training dataset.
[0027] Based on the optimized training data set, the initial lightweight spatiotemporal sequence large model is trained and optimized, and the large model parameters of the optimized lightweight spatiotemporal sequence large model are extracted and uploaded to the cloud data center;
[0028] In the cloud data center, based on the big model parameters uploaded by all edge terminals of the edge-cloud collaborative architecture, a federated learning algorithm is used to adjust the spatiotemporal sequence big model, extract the adjusted big model parameters of the adjusted spatiotemporal sequence big model, and send them to all edge terminals of the edge-cloud collaborative architecture;
[0029] At each edge, the optimized lightweight spatiotemporal sequence large model is updated according to the adjusted large model parameters to obtain the final lightweight spatiotemporal sequence large model.
[0030] Furthermore, a lightweight spatiotemporal sequence model is used to integrate and analyze real-time multi-source heterogeneous data collected at the edge to obtain real-time application scenario decisions. This includes the following steps:
[0031] Use the edge to collect users' real-time multi-source heterogeneous data and pre-process it to obtain real-time multi-source homogeneous data;
[0032] Use the standardized access and quality verification layer of the lightweight spatiotemporal series large model to perform standardized access and quality verification on real-time multi-source homogeneous data to obtain standardized real-time multi-source homogeneous data;
[0033] Use the cross-domain entity association layer of the lightweight spatiotemporal sequence large model to perform cross-domain entity association on standardized real-time multi-source homogeneous data to obtain real-time multi-source homogeneous data after association;
[0034] Use the fusion analysis layer of the lightweight spatiotemporal sequence large model to perform fusion analysis on the real-time multi-source homogeneous data after correlation to obtain real-time data analysis results;
[0035] Using the intelligent decision-making layer of the lightweight spatiotemporal sequence large model, intelligent decision-making is generated based on real-time application scenario requirements and real-time data analysis results to obtain real-time application scenario decisions.
[0036] Furthermore, the standardized access and quality verification layer of the lightweight spatiotemporal series large model is used to perform standardized access and quality verification on the real-time multi-source homogeneous data to obtain standardized real-time multi-source homogeneous data, including the following steps:
[0037] Input real-time multi-source homogeneous data into the standardized access and quality verification layer of the lightweight spatiotemporal series large model;
[0038] A lightweight isolation sandbox using a data isolation module with standardized access and quality verification layers isolates incoming real-time multi-source homogeneous data.
[0039] Use the data evidence audit module of the standardized access and quality verification layer to audit the real-time data evidence of real-time multi-source homogeneous data in the lightweight isolation sandbox;
[0040] If the audit passes, the real-time quality verification result is output as qualified and the process proceeds to the next step. Otherwise, the real-time quality verification result is output as unqualified and the large model analysis process ends.
[0041] Use the blockchain rights confirmation module of the standardized access and quality verification layer to verify the legitimacy of real-time multi-source homogeneous data;
[0042] If the legality verification passes, the real-time quality verification result is output as qualified and the next step is entered; otherwise, the real-time quality verification result is output as unqualified and the large model analysis process ends;
[0043] The data source filtering module of the standardized access and quality verification layer is used to filter low-quality data sources in real-time multi-source homogeneous data according to the node reputation evaluation mechanism to obtain standardized real-time multi-source homogeneous data.
[0044] Furthermore, the cross-domain entity association layer of the lightweight spatiotemporal sequence model is used to perform cross-domain entity association on the standardized real-time multi-source homogeneous data to obtain the associated real-time multi-source homogeneous data, including the following steps:
[0045] Input standardized real-time multi-source homogeneous data into the cross-domain entity association layer of the lightweight spatiotemporal sequence model;
[0046] Use the semantic mapping module of the cross-domain entity association layer to extract real-time cross-domain entities from standardized real-time multi-source homogeneous data;
[0047] According to the real-time cross-domain entities, the domain knowledge graph of the cross-domain entity association layer is used to perform cross-domain entity association on the standardized real-time multi-source homogeneous data to obtain the associated real-time multi-source homogeneous data.
[0048] Furthermore, the fusion analysis layer of the lightweight spatiotemporal sequence model is used to perform fusion analysis on the real-time multi-source homogeneous data after correlation to obtain real-time data analysis results, including the following steps:
[0049] Input the correlated real-time multi-source homogeneous data into the fusion analysis layer of the lightweight spatiotemporal sequence model;
[0050] Use the multi-source homogeneous data fusion model of the fusion analysis layer to extract and fuse features from the real-time multi-source homogeneous data after association to obtain real-time fusion features;
[0051] Use the fusion data analysis model of the fusion analysis layer to analyze the real-time fusion features and obtain real-time data analysis results.
[0052] Furthermore, the intelligent decision layer of the lightweight spatiotemporal sequence model is used to generate intelligent decisions based on the real-time application scenario requirements and real-time data analysis results, and obtain real-time application scenario decisions, including the following steps:
[0053] Input the real-time data analysis results into the intelligent decision-making layer of the lightweight spatiotemporal sequence model;
[0054] According to the requirements of real-time application scenarios, the action space, policy network, and reward function of the decision-making agent in the intelligent decision-making layer are adjusted to obtain the adjusted action space, adjusted policy network, and adjusted reward function;
[0055] According to the real-time data analysis results, the state space of the decision-making agent is updated to obtain the updated state space;
[0056] Based on the adjusted policy network and the adjusted reward function, a decision agent is used to generate intelligent decisions based on the adjusted action space and the updated state space to obtain real-time application scenario decisions.
[0057] The post-verification module of the intelligent decision-making layer is used to post-verify the initial real-time application scenario decision. If the post-verification passes, the final real-time application scenario decision is output; otherwise, the fusion analysis step is repeated.
[0058] A large-scale model analysis system for multi-source heterogeneous data fusion is used to implement a large-scale model analysis method. The system includes a cloud data center and several edge terminals. The cloud data center is provided with a large-scale spatiotemporal sequence model. The several edge terminals are communicatively connected to the cloud data center, and the edge terminals are provided with a lightweight large-scale spatiotemporal sequence model.
[0059] The beneficial effects of the present invention are:
[0060] The present invention discloses a large-scale model analysis method and system for multi-source heterogeneous data fusion, which provides security protection in data access, storage and sharing through trusted data space technology; combines data isolation, evidence auditing, blockchain confirmation and other mechanisms to effectively prevent data leakage and unauthorized access; the application of federated learning algorithm enables large-scale model parameters to be globally optimized without exposing the original data, further protecting the privacy of edge data; the edge-cloud collaborative architecture combines the powerful computing power of the cloud with the low latency and near real-time characteristics of the edge, deploys lightweight models on the edge, and can directly process local collected real-time data to meet the real-time requirements. Application scenarios with high requirements for accuracy; through technologies such as model pruning, knowledge distillation, small sample learning, and federated learning, not only the lightweight deployment of the model is achieved, but also local data can be used for fine-tuning at the edge, improving the model's adaptability and generalization capabilities for specific scenarios; the large spatiotemporal sequence model can better capture the temporal evolution and spatial correlation characteristics of data, and the standardized access and quality verification, cross-domain entity association, fusion analysis and other layer-by-layer processing ensure the information quality of the input decision layer. The intelligent decision layer dynamically adjusts the intelligent agent parameters according to the application scenario requirements, and verifies the decision results through the post-verification module, which improves the pertinence and reliability of the decision.
[0061] Other beneficial effects of the present invention will be further described in the specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 It is a flowchart of the large model analysis method for multi-source heterogeneous data fusion in the present invention.
[0063] Figure 2 It is a structural block diagram of the large model analysis system for multi-source heterogeneous data fusion in the present invention. DETAILED DESCRIPTION
[0064] The present invention will be further explained below with reference to the accompanying drawings and specific embodiments.
[0065] Example 1:
[0066] like Figure 1 As shown, this embodiment provides a large model analysis method for multi-source heterogeneous data fusion, including the following steps:
[0067] S1: In a cloud data center with edge-cloud collaborative architecture, we use trusted data space technology and intelligent agent decision-making technology to build a large spatiotemporal sequence model.
[0068] The spatiotemporal sequence model includes a standardized access and quality verification layer, a cross-domain entity association layer, a fusion analysis layer, and an intelligent decision-making layer that are connected in sequence;
[0069] The standardized access and quality verification layer designs data interfaces, isolation mechanisms (such as virtualization and containerization), audit log formats, blockchain writing rules, and filtering strategies. By combining trusted data spaces and evidence auditing functions, the legitimacy of data sources and the integrity of content are ensured. Through blockchain-based data ownership verification, evidence auditing, and security requirements, a full-link trusted environment for data circulation is established, ensuring that data cannot be tampered with and its source is traceable, significantly improving the credibility of data sharing.
[0070] At the data flow level, the cross-domain entity association layer achieves automated cross-domain data association and full-process credibility verification based on the deep integration of dynamic semantic alignment and trusted data space.
[0071] The fusion analysis layer uses machine learning algorithms to mine deep information from multi-source heterogeneous data, achieve multi-source heterogeneous data fusion and analysis, and provide support for subsequent reinforcement learning decisions;
[0072] The intelligent decision-making layer combines reinforcement learning algorithms with trusted data space authentication and reputation assessment mechanisms, making a qualitative leap in the transparency and explainability of the decision-making process, and significantly enhancing its ability to adapt to dynamic environments.
[0073] The standardized access and quality verification layer includes a data isolation module, a data evidence audit module, a blockchain rights confirmation module, and a data source filtering module;
[0074] At the initial stage of data access, real-time, multi-source, heterogeneous data from different sources, in different formats, and potentially with unknown risks is physically or logically separated from the core processing system or database. The data isolation module creates a controlled, secure processing environment (similar to a "sandbox") for the data, within which preliminary format conversion, structured processing, virus scanning, and malicious code detection are performed. This prevents bad data from directly impacting the main system, avoiding potential contamination, attacks, or performance degradation, and effectively isolates security threats (such as viruses, malicious code, and data injection attacks) that may be posed by external data sources, protecting the core system.
[0075] The data evidence audit module records, marks, and audits the accessed data, generating data evidence information that cannot be tampered with or denied. This module typically records the data's source, access time, key features, processing operations, etc., and audits the data evidence. It also provides an audit log that records all access, processing, and modification operations on the data, facilitating traceability.
[0076] The blockchain ownership confirmation module uses the decentralized, tamper-proof, and traceable characteristics of blockchain technology to confirm and record the ownership, use rights, or production rights of data. It generates a unique blockchain "digital fingerprint" (such as a hash value) for each piece (or batch) of key data and writes it to the blockchain to ensure the authenticity of the data source and the integrity of its original state, preventing the data from being maliciously tampered with or forged.
[0077] The data source filtering module screens data from different data sources based on preset rules, data quality scores, and node reputation assessment mechanisms, identifying and eliminating data from low-reputation, unstable, or historically underperforming data sources. This includes preliminary checks on data format, integrity, timeliness, consistency, and other aspects, filtering out clearly unqualified data.
[0078] The cross-domain entity association layer includes domain knowledge graph and semantic mapping modules;
[0079] Domain knowledge graphs use a graph structure (node-edge-attribute) to store domain-specific or cross-domain structured knowledge, including entities (such as people, places, events, and devices), concepts, and their attributes and relationships. This provides rich contextual information and background knowledge for identified real-time cross-domain entities, and defines the entity's role, attributes, and standard relationships with other entities within the domain.
[0080] The semantic mapping module identifies and extracts entities with cross-domain association potential from standardized multi-source heterogeneous data (which may contain structured fields or unstructured text descriptions), and extracts key features of entities, such as name, type, context description, timestamp, spatial information, etc., for subsequent matching and association;
[0081] The fusion analysis layer includes multi-source homogeneous data fusion model and fusion data analysis model;
[0082] The multi-source homogeneous data fusion model is built based on the Spatio-Temporal Graph Neural Network (ST-GNN)-Attention-Multilayer Perceptron (MLP) algorithm, and includes a sequentially connected temporal graph feature extraction module built based on the ST-GNN algorithm, an attention weighting module built based on the Attention mechanism, and a multi-source homogeneous data fusion module built based on the MLP algorithm.
[0083] The time series graph feature extraction module extracts spatiotemporal features from the input correlated real-time multi-source homogeneous data. It organizes data into a graph structure (nodes represent entities or data points, and edges represent relationships or spatial proximity between entities) and uses a spatiotemporal graph neural network (ST-GNN) to learn the spatiotemporal dynamic representation of nodes (or edges). ST-GNN combines the ability of graph neural networks (GNNs) to capture spatial relationships with the ability of specific mechanisms (such as gating units and attention mechanisms) to capture temporal dependencies, and can learn complex interaction patterns of data in spatial and temporal dimensions. The attention weighting module receives spatiotemporal features extracted from the ST-GNN and uses the attention mechanism to dynamically calculate the importance weights of these features. It can learn which spatiotemporal features (from which nodes and which time steps) are more important for subsequent fusion and analysis tasks. The multi-source homogeneous data fusion module receives spatiotemporal features weighted by the attention mechanism and uses a multi-layer perceptron (MLP) to perform nonlinear transformation and integration on these features. It can learn complex interactions between features and fuse spatiotemporally extracted and attention-weighted features from different sources (although homogeneous, they may come from different sensors or systems) into a unified and compact representation (i.e., real-time fusion features).
[0084] The fusion data analysis model is constructed based on the Deep Belief Network (DBN)-Domain Adversarial Neural Network (DANN) algorithm, and includes a fusion data analysis module constructed based on the DBN algorithm and an adversarial training module constructed based on the DANN algorithm, which are connected in sequence.
[0085] The fusion data analysis module effectively processes the nonlinear relationships and complex dependencies in the input data, extracting information that cannot be captured by hand-crafted features. The core goal of the adversarial training module is to learn universal features that do not vary with the data source (domain). This is achieved by introducing a domain discriminator, which attempts to distinguish features from different domains, while the feature extraction module attempts to generate features that the discriminator cannot distinguish from their source. Through adversarial training, this module can effectively reduce or eliminate the impact of possible distribution differences (domain shift) between different data sources on the analysis results, forcing the model to focus on common information in the data that is not related to the domain.
[0086] The intelligent decision-making layer includes decision-making agents and post-verification modules;
[0087] The decision-making agent is constructed based on the Meta-Policy Optimization (MPO)-Multi-Objective Group Relative Policy Optimization (MOGRPO) algorithm, and includes a meta-policy optimization module constructed based on the MPO algorithm and an application scenario decision generation module constructed based on the MOGRPO algorithm, which are connected in sequence. The application scenario decision generation module includes an objective function set, an experience replay pool, an actor network, and an agent, and the agent is connected to the objective function set, the experience replay pool, and the actor network respectively.
[0088] The meta-strategy optimization module is used to generate the initial network parameters of the Actor network in the application scenario decision generation module so that these parameters can quickly adapt to new and unseen application scenario requirements and data analysis results, thereby improving the generalization ability of the model. Even under unseen inputs, the Actor network can be updated based on previous learning experience, thereby improving the adaptability of the decision-making agent. The objective function set of the application scenario decision generation module can handle multiple conflicting optimization goals, such as generation efficiency, decision cost, etc., and generate application scenario decisions that balance these goals. The agent learns historical experience through the experience replay pool and continuously optimizes its decision-making ability. The agent controls the Actor network based on the learned experience to generate more effective application scenario decisions. The design of the experience replay pool and the intelligent agent enables the decision-making intelligent agent to continuously learn and optimize, thereby improving the quality of decision generation. The application scenario decision generation module adopts a group exploration method, which can avoid falling into the local optimal solution to a certain extent. The Actor network outputs the distribution probability of actions under a given state. The goal is to learn an optimal strategy, that is, to maximize the long-term cumulative reward. In the continuous action space, the Actor network usually outputs a mean and an optional variance parameter to describe the probability distribution of the action. The experience replay pool is used to store historical experience for reuse during training. The application scenario decision generation module directly updates the Actor network through gradients, eliminating the Critic network in traditional reinforcement learning, making the algorithm structure simpler.
[0089] In a cloud data center with an edge-cloud collaborative architecture, we use trusted data space technology and intelligent agent decision-making technology to build a large spatiotemporal sequence model, including the following steps:
[0090] S1-1: Distributedly connect several edge devices with cloud data centers to build an edge-cloud collaborative architecture;
[0091] Multiple edge computing nodes (edge devices) deployed in different geographical locations and on different devices are connected to the central cloud data center through networks (such as 5G, Wi-Fi, and wired networks). Clear data flow, control flow, and computing task allocation rules are defined, allowing the edge devices to process real-time data locally and upload tasks that require centralized processing or model training to the cloud. At the same time, models and policies in the cloud can be distributed to the edge devices. The edge-cloud collaborative architecture aims to combine the low latency and high bandwidth of edge computing with the powerful computing power, storage capacity, and global optimization capabilities of the cloud center.
[0092] S1-2: In a cloud data center with edge-cloud collaborative architecture, we use trusted data space technology and intelligent agent decision-making technology to build a large spatiotemporal sequence model architecture.
[0093] Design and define the specific network structure, components at each layer and their interaction methods of the spatiotemporal sequence model;
[0094] S1-3: Use the privacy-enhanced fusion algorithm to improve the spatiotemporal sequence large model architecture and obtain an improved spatiotemporal sequence large model architecture;
[0095] The privacy-enhanced fusion algorithm uses an improved federated learning framework combined with an attention mechanism to weightedly aggregate the gradients of a large spatiotemporal sequence model architecture, and dynamically adjusts the noise intensity through adaptive differential privacy technology.
[0096] S1-4: Using the Stochastic Snow Goose Algorithm (SSGA), the initial large model parameters of the improved spatiotemporal sequence large model architecture are optimized to obtain the spatiotemporal sequence large model, including the following steps:
[0097] S1-4-1: Integrate the initial parameters of the standardized access and quality verification layer, cross-domain entity association layer, fusion analysis layer, and intelligent decision-making layer in the improved spatiotemporal sequence large model architecture to obtain the initial large model parameters;
[0098] S1-4-2: Encode the initial large model parameters into the individual vectors of SSGA individuals in the SSGA algorithm, and set the SSGA population parameters and the maximum number of iterations;
[0099] S1-4-3: Take minimizing the error value as the optimization goal, and set the fitness function of the SSGA algorithm according to the optimization goal;
[0100] The formula is:
[0101] Fit(P)=minMSN(P)
[0102] Where Fit(P) is the fitness function; MSN(P) is the mean square error function; P is the SSGA individual;
[0103] S1-4-4: Based on the individual vectors of the SSGA individuals and the SSGA population parameters, the Circle chaotic mapping sequence is used to generate initial solutions and obtain several initial solutions; each initial solution corresponds to an initial SSGA individual, and several initial SSGA individuals constitute the initial SSGA individual population;
[0104] The formula is:
[0105]
[0106] Where, P i The initial SSGA individual generated by the Circle chaotic map sequence, that is, the initial solution; P o is a randomly generated SSGA individual; i is the SSGA individual indicator; mod(*) is the remainder function;
[0107] Compared with the randomly distributed population, the initial position distribution of the improved SSGA population generated by the Circle chaotic mapping sequence is more uniform, which expands the search range of the SSGA population in space and increases the diversity of group positions. To a certain extent, it improves the defect that the algorithm is prone to falling into local extreme values, thereby improving the optimization efficiency of the algorithm.
[0108] S1-4-5: Use the fitness function to obtain the initial fitness value of each initial SSGA individual in the initial SSGA population, and take the initial SSGA individual with the lowest fitness value as the leader goose;
[0109] S1-4-6: Entering the exploration phase, the leader goose rotation mechanism, the calling guidance mechanism, and the dynamic reverse mechanism are introduced to iteratively update the initial SSGA population to obtain an updated SSGA population and retain the best individuals;
[0110] The leader goose rotation mechanism selects a new leader goose in each iteration based on the fitness value of the SSGA individuals. This mechanism can prevent the leader goose from falling into the local optimum too early and enhance the global search capability of the algorithm.
[0111] The formula is:
[0112]
[0113] Where, Be the leader of an update; is the initial SSGA individual with the third lowest fitness value in the initial SSGA population at the tth and t+1th iterations; is the fifth-to-last initial SSGA individual in the initial SSGA population with the highest fitness value at the tth iteration; t is the current iteration number; is the optimal individual; a is the first weight factor; rand is the random number generation function;
[0114] The calling guidance mechanism uses a sound wave propagation attenuation model to adjust the position update of individual SSGAs based on their distance from the leader goose. SSGAs that are closer are more influenced by the leader goose in their position updates, allowing them to quickly approach the optimal solution. SSGAs that are farther away are less influenced by the leader goose in their position updates, allowing them to maintain a certain level of exploration capability. This mechanism can prevent excessive aggregation or dispersion of the group and improve the local search accuracy of the algorithm.
[0115] The formula is:
[0116]
[0117] Where, It is an updated SSGA entity; is the initial SSGA individual of the tth iteration; is the sound intensity received by the initial SSGA individual; is the sound intensity parameter; L WA is the initial sound intensity; L low is the minimum acceptable sound intensity; a" is the convergence factor; is the initial SSGA individual with the farthest distance; r' is a random parameter; B(d) is the Brownian motion function; d is the Brownian motion parameter; ⊕ is the XOR processing symbol;
[0118]
[0119] Where a" is the convergence factor; tanh(.) is the hyperbolic tangent function; t is the current number of iterations; t max is the maximum number of iterations; a max 、a min are the maximum and minimum values of the convergence factor, respectively; λ is the decreasing rate parameter, k' is the decreasing period parameter, λ = -2π, k' = π;
[0120] Dynamic reversal mechanism, which dynamically reverses the initial SSGA individuals to improve the diversity of exploration directions and avoid falling into local optimality;
[0121] The formula is:
[0122]
[0123] Where, is the reverse SSGA individual updated once; γ is the decreasing inertia coefficient; L max , L min are the maximum and minimum values of the vector space respectively;
[0124] The leader goose of one update, several SSGA individuals of one update, and several reverse SSGA individuals of one update are integrated to obtain the SSGA population of one update, and the SSGA individual with the lowest fitness value is retained as the optimal individual;
[0125] S1-4-7: Entering the development phase, introducing the abnormal boundary strategy and Gaussian mutation mechanism, performing a second update on the once-updated SSGA population, obtaining a second-updated SSGA population, and retaining the optimal individual;
[0126] The abnormal boundary strategy calculates the difference between the fitness value of each updated SSGA individual and the group average fitness value. For SSGA individuals whose fitness value is much higher than the group average, their position update method will be adjusted, such as using Gaussian mutation mechanism, larger step size or smaller step size. This mechanism can help individuals avoid falling into local optimality and improve the convergence speed and accuracy of the algorithm;
[0127] The formula is:
[0128]
[0129] Where, is the second updated SSGA individual; is an updated SSGA individual; Fit(*) is the fitness function; Fit avg is the average fitness value of the group; is the SSGA individual with the highest fitness value; a' and e are the second and third weight factors; G(1,1) is the Gaussian mutation mechanism parameter;
[0130] S1-4-8: If the number of iterations is greater than or equal to the iteration threshold or the fitness value of the optimal individual is less than the fitness threshold, the optimal individual is output as the optimal solution;
[0131] S1-4-9: Decode the individual vector of the optimal individual to obtain the optimal initial large model parameters, and optimize the improved spatiotemporal sequence large model architecture based on the optimal initial large model parameters to obtain the initial spatiotemporal sequence large model;
[0132] S1-4-10: Input several preset training data sets, optimize and train the initial spatiotemporal sequence model, and obtain a trained spatiotemporal sequence model;
[0133] S2: At the edge of the edge-cloud collaborative architecture, lightweight deployment of the spatiotemporal sequence model is performed to obtain a lightweight spatiotemporal sequence model, including the following steps:
[0134] S2-1: Extract the large model metadata of the spatiotemporal sequence large model in the cloud data center, and use model pruning technology to structurally prune the large model metadata to obtain lightweight large model metadata, and send it to all edge terminals of the edge-cloud collaborative architecture;
[0135] S2-2: Based on the metadata of the lightweight large model, the spatiotemporal sequence large model is lightweight deployed at the edge of the edge-cloud collaborative architecture to obtain the initial lightweight spatiotemporal sequence large model;
[0136] S2-3: At each edge, collect the initial training dataset and use knowledge distillation and small sample learning mechanisms to dynamically optimize the initial training dataset to obtain the optimized training dataset;
[0137] The lightweight spatiotemporal sequence model combines knowledge distillation with small sample learning mechanisms to dynamically optimize the training dataset based on decision-making requirements, ensuring the model's ability to accurately analyze complex scenarios.
[0138] S2-4: Based on the optimized training data set, the initial lightweight spatiotemporal sequence large model is trained and optimized, the large model parameters of the optimized lightweight spatiotemporal sequence large model are extracted, and uploaded to the cloud data center;
[0139] S2-5: In the cloud data center, based on the big model parameters uploaded by all edge terminals of the edge-cloud collaborative architecture, a federated learning algorithm is used to adjust the spatiotemporal sequence big model. The adjusted big model parameters of the adjusted spatiotemporal sequence big model are extracted and sent to all edge terminals of the edge-cloud collaborative architecture.
[0140] S2-6: At each edge, the optimized lightweight spatiotemporal sequence large model is updated according to the adjusted large model parameters to obtain the final lightweight spatiotemporal sequence large model;
[0141] S3: Use a lightweight spatiotemporal sequence model to integrate and analyze real-time multi-source heterogeneous data collected at the edge to make real-time application scenario decisions. This includes the following steps:
[0142] S3-1: Use the edge to collect users' real-time multi-source heterogeneous data and preprocess it to obtain real-time multi-source homogeneous data;
[0143] Preprocessing includes data cleaning, missing value supplementation, error value removal, format conversion, and normalization to improve data quality and provide standard input for subsequent fusion analysis;
[0144] S3-2: Use the standardized access and quality verification layer of the lightweight spatiotemporal series large model to perform standardized access and quality verification on real-time multi-source homogeneous data to obtain standardized real-time multi-source homogeneous data, including the following steps:
[0145] S3-2-1: Standardized access and quality verification layer for inputting real-time multi-source homogeneous data into lightweight spatiotemporal series models;
[0146] The formula is:
[0147] D in =StandardizedAccessLayer(X pre )
[0148] IdentifyType(D in )∈{D struct ,D unstruct}
[0149] Where, X pre Real-time multi-source homogeneous data; D in is the real-time multi-source homogeneous data feature; StandardizedAccessLayer(*) is the feature engineering function of the standardized access and quality verification layer; IdentifyType(*) is the feature identification function; D struct ,D unstruct Input feature D for real-time multi-source homogeneous data in The structured and unstructured features of
[0150] S3-2-2: A lightweight isolation sandbox using a data isolation module at the standardized access and quality verification layer isolates incoming real-time multi-source homogeneous data.
[0151] The formula is:
[0152] D isolated =W clean [S isolated (D in )]
[0153] Where D isolated It is the isolated real-time multi-source homogeneous data feature; S isolated (*) is the isolation function of the lightweight isolation sandbox; W clean [*] is the incremental cleaning function;
[0154] S3-2-3: Use the data evidence audit module of the standardized access and quality verification layer to audit the real-time data evidence of real-time multi-source homogeneous data in the lightweight isolation sandbox;
[0155] The formula is:
[0156]
[0157] Where, It is the audit information of isolated real-time multi-source homogeneous data features; TDS(*) is the trusted data space evidence storage function; Verify(*) is the audit function;
[0158] S3-2-4: If the audit passes, the real-time quality verification result is output as qualified and the process proceeds to the next step. Otherwise, the real-time quality verification result is output as unqualified and the large model analysis process ends.
[0159] S3-2-5: Use the blockchain rights confirmation module of the standardized access and quality verification layer to verify the legitimacy of real-time multi-source homogeneous data;
[0160] The formula is:
[0161]
[0162] Where, is the legality information of isolated real-time multi-source homogeneous data features; LV(*) is the legality acquisition function; Verify'(*) is the legality verification function;
[0163] S3-2-6: If the legality verification passes, the real-time quality verification result is output as qualified and the process proceeds to the next step. Otherwise, the real-time quality verification result is output as unqualified and the large model analysis process ends.
[0164] S3-2-7: Use the data source filtering module of the standardized access and quality verification layer to filter low-quality data sources in the real-time multi-source homogeneous data based on the node reputation evaluation mechanism to obtain standardized real-time multi-source homogeneous data;
[0165] The formula is:
[0166] D std =Filter(D isolated ,cc)
[0167] X std =HY(D std )
[0168] Where D std is the standardized real-time multi-source homogeneous data feature; Filter(*) is the filter function of the node reputation evaluation mechanism; X std is the standardized real-time multi-source homogeneous data; HY(*) is the data restoration function; cc is the node reputation evaluation parameter;
[0169] S3-3: Use the cross-domain entity association layer of the lightweight spatiotemporal sequence large model to perform cross-domain entity association on the standardized real-time multi-source homogeneous data to obtain the associated real-time multi-source homogeneous data, including the following steps:
[0170] S3-3-1: Input standardized real-time multi-source homogeneous data into the cross-domain entity association layer of the lightweight spatiotemporal sequence model;
[0171] S3-3-2: Use the semantic mapping module of the cross-domain entity association layer to extract real-time cross-domain entities from standardized real-time multi-source homogeneous data;
[0172] S3-3-3: Based on real-time cross-domain entities, use the domain knowledge graph of the cross-domain entity association layer to perform cross-domain entity association on standardized real-time multi-source homogeneous data to obtain the associated real-time multi-source homogeneous data;
[0173] S3-4: Use the fusion analysis layer of the lightweight spatiotemporal sequence model to perform fusion analysis on the correlated real-time multi-source homogeneous data to obtain real-time data analysis results, including the following steps:
[0174] S3-4-1: Input the correlated real-time multi-source homogeneous data into the fusion analysis layer of the lightweight spatiotemporal sequence model;
[0175] S3-4-2: Use the multi-source homogeneous data fusion model of the fusion analysis layer to extract and fuse features from the associated real-time multi-source homogeneous data to obtain real-time fusion features, including the following steps:
[0176] S3-4-2-1: Use the time series graph feature extraction module of the multi-source homogeneous data fusion model of the fusion analysis layer to extract the real-time time series graph features of the correlated real-time multi-source homogeneous data;
[0177] S3-4-2-2: Based on the preset attention weights of the attention weighting module of the multi-source isomorphic data fusion model, the multi-source isomorphic data fusion module of the multi-source isomorphic data fusion model is used to perform weighted fusion on the real-time time series graph features to obtain real-time fusion features;
[0178] S3-4-3: Use the fusion data analysis model of the fusion analysis layer to analyze the real-time fusion features and obtain real-time data analysis results;
[0179] S3-5: Using the intelligent decision-making layer of the lightweight spatiotemporal sequence model, intelligent decision-making is generated based on real-time application scenario requirements and real-time data analysis results to obtain real-time application scenario decisions, including the following steps:
[0180] S3-5-1: Input the real-time data analysis results into the intelligent decision-making layer of the lightweight spatiotemporal sequence model;
[0181] S3-5-2: Based on the requirements of real-time application scenarios, the meta-strategy optimization module of the decision-making agent in the intelligent decision-making layer is used to adjust the action space parameters, policy network parameters, and reward function parameters of the decision-making agent to obtain the adjusted action space, adjusted policy network, and adjusted reward function.
[0182] For example, in smart transportation scenarios, action space parameters, policy network parameters, and reward function parameters should be adapted to the specific application scenario;
[0183] S3-5-3: Update the state space of the decision-making agent based on the real-time data analysis results to obtain the updated state space;
[0184] S3-5-4: Based on the adjusted policy network and the adjusted reward function, use the decision agent to make intelligent decisions based on the adjusted action space and the updated state space to obtain real-time application scenario decisions;
[0185] S3-5-5: Use the post-verification module of the intelligent decision-making layer to post-verify the initial real-time application scenario decision. If the post-verification passes, the final real-time application scenario decision is output; otherwise, the fusion analysis step is repeated;
[0186] Post-verification includes three-factor verification, which is as follows:
[0187] First level of verification: Based on the internal consistency check of the model, we ensure the logical rationality of the initial real-time application scenario decision and ensure that the initial real-time application scenario decision is logically self-consistent and complies with the preset rules and constraints.
[0188] Second level of verification: Introducing an external knowledge base or expert system to verify the rationality of the initial real-time application scenario decision, ensuring that the initial real-time application scenario decision is not only logically self-consistent but also consistent with real-world knowledge, experience, best practices, or the judgment of domain experts;
[0189] The third level of verification: Through the real-time feedback mechanism, the actual application effect of the initial real-time application scenario decision is verified to ensure that the initial real-time application scenario decision can achieve the expected effect in actual implementation, or at least is acceptable and effective.
[0190] Example 2:
[0191] like Figure 2 As shown, this embodiment provides a large model analysis system for multi-source heterogeneous data fusion, which is used to implement a large model analysis method. The system includes a cloud data center and several edge terminals. The cloud data center is equipped with a spatiotemporal sequence large model. Several edge terminals are communicated with the cloud data center, and the edge terminals are equipped with a lightweight spatiotemporal sequence large model.
[0192] The cloud data center is used to build a large spatiotemporal sequence model using trusted data space technology and intelligent agent decision-making technology; the large spatiotemporal sequence model is lightweight deployed at the edge of the edge-cloud collaborative architecture;
[0193] The edge is used to use a lightweight spatiotemporal sequence model to fuse and analyze real-time multi-source heterogeneous data collected at the edge to obtain real-time application scenario decisions.
[0194] The present invention discloses a large-scale model analysis method and system for multi-source heterogeneous data fusion, which provides security protection in data access, storage and sharing through trusted data space technology; combines data isolation, evidence auditing, blockchain confirmation and other mechanisms to effectively prevent data leakage and unauthorized access; the application of federated learning algorithm enables large-scale model parameters to be globally optimized without exposing the original data, further protecting the privacy of edge data; the edge-cloud collaborative architecture combines the powerful computing power of the cloud with the low latency and near real-time characteristics of the edge, deploys lightweight models on the edge, and can directly process local collected real-time data to meet the real-time requirements. Application scenarios with high requirements for accuracy; through technologies such as model pruning, knowledge distillation, small sample learning, and federated learning, not only the lightweight deployment of the model is achieved, but also local data can be used for fine-tuning at the edge, improving the model's adaptability and generalization capabilities for specific scenarios; the large spatiotemporal sequence model can better capture the temporal evolution and spatial correlation characteristics of data, and the standardized access and quality verification, cross-domain entity association, fusion analysis and other layer-by-layer processing ensure the information quality of the input decision layer. The intelligent decision layer dynamically adjusts the intelligent agent parameters according to the application scenario requirements, and verifies the decision results through the post-verification module, which improves the pertinence and reliability of the decision.
[0195] The present invention is not limited to the above optional embodiments. Anyone can derive various other forms of products based on the teachings of the present invention. The above specific embodiments should not be construed as limiting the scope of protection of the present invention. The scope of protection of the present invention shall be based on the scope defined in the claims, and the description can be used to interpret the claims.
Claims
1. A large-scale model analysis method for multi-source heterogeneous data fusion, characterized by: The steps include: In the cloud data center of the edge-cloud collaborative architecture, we use trusted data space technology and intelligent agent decision-making technology to build a large spatiotemporal sequence model; At the edge of the edge-cloud collaborative architecture, the spatiotemporal sequence large model is deployed in a lightweight manner to obtain a lightweight spatiotemporal sequence large model; Use a lightweight spatiotemporal sequence model to fuse and analyze real-time multi-source heterogeneous data collected at the edge to obtain real-time application scenario decisions.
2. The large-scale model analysis method for multi-source heterogeneous data fusion according to claim 1 is characterized by: The spatiotemporal sequence model includes a standardized access and quality verification layer, a cross-domain entity association layer, a fusion analysis layer, and an intelligent decision-making layer connected in sequence; The standardized access and quality verification layer includes a data isolation module, a data evidence audit module, a blockchain rights confirmation module, and a data source filtering module; The cross-domain entity association layer includes a domain knowledge graph and a semantic mapping module; The fusion analysis layer includes a multi-source homogeneous data fusion model and a fusion data analysis model; The intelligent decision-making layer includes a decision-making agent and a post-verification module.
3. The large-scale model analysis method for multi-source heterogeneous data fusion according to claim 2 is characterized by: In a cloud data center with an edge-cloud collaborative architecture, we use trusted data space technology and intelligent agent decision-making technology to build a large spatiotemporal sequence model, including the following steps: Distribute connections between several edge devices and cloud data centers to build an edge-cloud collaborative architecture; In the cloud data center of the edge-cloud collaborative architecture, we use trusted data space technology and intelligent agent decision-making technology to build a large-scale spatiotemporal sequence model architecture; Using a privacy-enhanced fusion algorithm, the spatiotemporal sequence large model architecture is improved to obtain an improved spatiotemporal sequence large model architecture; The optimization algorithm is used to optimize the initial large model parameters of the improved spatiotemporal sequence large model architecture to obtain the spatiotemporal sequence large model.
4. The large-scale model analysis method for multi-source heterogeneous data fusion according to claim 3 is characterized by: At the edge of the edge-cloud collaborative architecture, lightweight deployment of the spatiotemporal sequence model is performed to obtain a lightweight spatiotemporal sequence model, including the following steps: Extract the large model metadata of the spatiotemporal sequence large model in the cloud data center, and use model pruning technology to structurally prune the large model metadata to obtain lightweight large model metadata, and send it to all edge terminals of the edge-cloud collaborative architecture; Based on the metadata of the lightweight large model, the spatiotemporal sequence large model is lightweight deployed at the edge of the edge-cloud collaborative architecture to obtain the initial lightweight spatiotemporal sequence large model. At each edge, an initial training dataset is collected and dynamically optimized using knowledge distillation and small sample learning mechanisms to obtain an optimized training dataset. Based on the optimized training data set, the initial lightweight spatiotemporal sequence large model is trained and optimized, and the large model parameters of the optimized lightweight spatiotemporal sequence large model are extracted and uploaded to the cloud data center; In the cloud data center, based on the big model parameters uploaded by all edge terminals of the edge-cloud collaborative architecture, a federated learning algorithm is used to adjust the spatiotemporal sequence big model, extract the adjusted big model parameters of the adjusted spatiotemporal sequence big model, and send them to all edge terminals of the edge-cloud collaborative architecture; At each edge, the optimized lightweight spatiotemporal sequence large model is updated according to the adjusted large model parameters to obtain the final lightweight spatiotemporal sequence large model.
5. The large-scale model analysis method for multi-source heterogeneous data fusion according to claim 4 is characterized by: Using a lightweight spatiotemporal sequence model, we can integrate and analyze real-time multi-source heterogeneous data collected at the edge to make real-time application scenario decisions. This involves the following steps: Use the edge to collect users' real-time multi-source heterogeneous data and pre-process it to obtain real-time multi-source homogeneous data; Use the standardized access and quality verification layer of the lightweight spatiotemporal series large model to perform standardized access and quality verification on real-time multi-source homogeneous data to obtain standardized real-time multi-source homogeneous data; Use the cross-domain entity association layer of the lightweight spatiotemporal sequence large model to perform cross-domain entity association on standardized real-time multi-source homogeneous data to obtain real-time multi-source homogeneous data after association; Use the fusion analysis layer of the lightweight spatiotemporal sequence large model to perform fusion analysis on the real-time multi-source homogeneous data after correlation to obtain real-time data analysis results; Using the intelligent decision-making layer of the lightweight spatiotemporal sequence large model, intelligent decision-making is generated based on real-time application scenario requirements and real-time data analysis results to obtain real-time application scenario decisions.
6. The large-scale model analysis method for multi-source heterogeneous data fusion according to claim 5 is characterized by: Using the standardized access and quality verification layer of the lightweight spatiotemporal series large model, standardize access and quality verification of real-time multi-source homogeneous data to obtain standardized real-time multi-source homogeneous data, including the following steps: Input real-time multi-source homogeneous data into the standardized access and quality verification layer of the lightweight spatiotemporal series large model; A lightweight isolation sandbox using a data isolation module with standardized access and quality verification layers isolates incoming real-time multi-source homogeneous data. Use the data evidence audit module of the standardized access and quality verification layer to audit the real-time data evidence of real-time multi-source homogeneous data in the lightweight isolation sandbox; If the audit passes, the real-time quality verification result is output as qualified and the process proceeds to the next step. Otherwise, the real-time quality verification result is output as unqualified and the large model analysis process ends. Use the blockchain rights confirmation module of the standardized access and quality verification layer to verify the legitimacy of real-time multi-source homogeneous data; If the legality verification passes, the real-time quality verification result is output as qualified and the next step is entered; otherwise, the real-time quality verification result is output as unqualified and the large model analysis process ends; The data source filtering module of the standardized access and quality verification layer is used to filter low-quality data sources in real-time multi-source homogeneous data according to the node reputation evaluation mechanism to obtain standardized real-time multi-source homogeneous data.
7. The large-scale model analysis method for multi-source heterogeneous data fusion according to claim 6 is characterized by: Using the cross-domain entity association layer of the lightweight spatiotemporal sequence model, cross-domain entity association is performed on standardized real-time multi-source homogeneous data to obtain the associated real-time multi-source homogeneous data, including the following steps: Input standardized real-time multi-source homogeneous data into the cross-domain entity association layer of the lightweight spatiotemporal sequence model; Use the semantic mapping module of the cross-domain entity association layer to extract real-time cross-domain entities from standardized real-time multi-source homogeneous data; According to the real-time cross-domain entities, the domain knowledge graph of the cross-domain entity association layer is used to perform cross-domain entity association on the standardized real-time multi-source homogeneous data to obtain the associated real-time multi-source homogeneous data.
8. The large-scale model analysis method for multi-source heterogeneous data fusion according to claim 7 is characterized by: The fusion analysis layer of the lightweight spatiotemporal sequence model is used to perform fusion analysis on the real-time multi-source homogeneous data after correlation to obtain real-time data analysis results, including the following steps: Input the correlated real-time multi-source homogeneous data into the fusion analysis layer of the lightweight spatiotemporal sequence model; Use the multi-source homogeneous data fusion model of the fusion analysis layer to extract and fuse features from the real-time multi-source homogeneous data after association to obtain real-time fusion features; Use the fusion data analysis model of the fusion analysis layer to analyze the real-time fusion features and obtain real-time data analysis results.
9. The large-scale model analysis method for multi-source heterogeneous data fusion according to claim 8 is characterized by: The intelligent decision layer of the lightweight spatiotemporal sequence model generates intelligent decisions based on real-time application scenario requirements and real-time data analysis results, and obtains real-time application scenario decisions, including the following steps: Input the real-time data analysis results into the intelligent decision-making layer of the lightweight spatiotemporal sequence model; According to the requirements of real-time application scenarios, the action space, policy network, and reward function of the decision-making agent in the intelligent decision-making layer are adjusted to obtain the adjusted action space, adjusted policy network, and adjusted reward function; According to the real-time data analysis results, the state space of the decision-making agent is updated to obtain the updated state space; Based on the adjusted policy network and the adjusted reward function, a decision agent is used to generate intelligent decisions based on the adjusted action space and the updated state space to obtain the initial real-time application scenario decision. The post-verification module of the intelligent decision-making layer is used to post-verify the initial real-time application scenario decision. If the post-verification passes, the final real-time application scenario decision is output; otherwise, the fusion analysis step is repeated.
10. A large model analysis system for multi-source heterogeneous data fusion, used to implement the large model analysis method according to any one of claims 1 to 8, characterized in that: The system includes a cloud data center and several edge terminals. The cloud data center is provided with a large spatiotemporal sequence model. Several edge terminals are communicatively connected with the cloud data center, and the edge terminals are provided with a lightweight large spatiotemporal sequence model.
Citation Information
Cited By
Self-adaptive multi-source heterogeneous data cleaning method and system based on large model driving
CN121278248A