Heterogeneous data fusion method for optimizing inter-satellite switching of low earth orbit satellite network
Through the heterogeneous data fusion method, combined with the learnable projection matrix, dynamic graph structure and multi-head attention mechanism, the computing complexity and real-time requirements in inter-star switching decisions of low-orbit satellite networks are solved, and efficient and reliable inter-star switching decisions are achieved.
Patent Information
- Application Number
- CN202510467544.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-06-24
AI Technical Summary
The inter-star switching decisions of low-orbit satellite networks under high-speed motion, random terminal migration and complex environmental interference are faced with the contradiction between computing complexity and real-time requirements, which leads to the inability of traditional methods to meet the service tolerance threshold in scenarios such as emergency communications.
The heterogeneous data fusion method is adopted to collect multi-source heterogeneous data in real time, design a special encoder to extract features, use a learning projection matrix to align cross-modal features, build a dynamic graph structure, and combine learning sharing tokens and multi-head attention mechanisms for information interaction. Finally, the time-convolution network enhances timing dependency modeling and perform inter-star switching decisions.
This method effectively eliminates noise interference and semantic conflicts caused by direct splicing of heterogeneous features, reduces redundant computing overhead, improves the real-time and reliability of inter-satellite switching decisions, and meets the multi-objective coordination needs of low-orbit satellite networks in high dynamic scenarios.
Smart Images

Figure CN120200657A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence and deep learning, and specifically relates to a heterogeneous data fusion method for optimizing inter-satellite handover in a low-earth orbit satellite network. Background Art
[0002] As an important carrier for space-air-ground integrated communication, the inter-satellite handover mechanism of a low-earth orbit satellite network (LEO) directly determines the network continuity guarantee ability and service quality level. Under the multiple effects of high-speed satellite movement, random terminal migration, and complex environmental interference, the network topology exhibits dynamic characteristics. The traditional handover decision-making mode based on a single data domain faces fundamental challenges, and it is difficult to achieve robust decision-making relying solely on local physical layer indicators or fixed strategies. It is necessary to deeply integrate multi-domain heterogeneous data such as real-time physical states, service demand priorities, and environmental interference predictions, and construct a dynamic optimization model through cross-modal feature fusion to meet the multi-objective collaboration requirements in high-dynamic scenarios.
[0003] However, existing heterogeneous data fusion technologies have limitations in dealing with the dynamic nature of satellite networks. There are problems such as time asynchrony and space misalignment in the link state parameters (structured time-series data), service type labels (discrete semantic data), and meteorological grid images (spatial feature data) obtained through satellite-ground collaboration collected by on-board sensors, resulting in feature conflicts during cross-modal semantic alignment using traditional feature splicing methods.
[0004] Although existing dynamic fusion schemes based on graph attention networks (GAT) can adapt to topological changes through dynamic weight adjustment, their fully connected interaction mechanism will generate exponential computational complexity during large-scale spatio-temporal migration of satellite nodes. When the constellation scale expands to multiple satellite nodes, the calculation of attention weights for inter-satellite links by traditional GAT will trigger a large number of redundant operations, which can neither effectively utilize the predictability of satellite orbits to optimize the calculation path nor have a lightweight feature extraction mechanism for the spatio-temporal correlation of satellite-ground data streams. This contradiction between computational load and real-time requirements will directly lead to handover delays exceeding the service tolerance threshold in millisecond-level decision-making scenarios such as emergency communication, severely restricting the service reliability of satellite networks. Summary of the Invention
[0005] Aiming at the deficiencies of the prior art, the present invention provides a heterogeneous data fusion method for optimizing inter-satellite handover in a low-earth orbit satellite network, which includes:
[0006] Real-time collect multi-source heterogeneous data in the low-earth orbit satellite network, including physical domain data, service domain data, and environmental domain data, and perform feature extraction through different modal encoders respectively;
[0007] Map the features of each modality to a unified dimension through a learnable projection matrix to generate an aligned multi-source data feature matrix;
[0008] Construct a dynamic graph structure based on the aligned multi-source data feature matrix;
[0009] According to the dynamic graph structure, combine the learnable shared token and the multi-head attention mechanism to generate the shared token feature, and broadcast the shared token feature to all nodes and add it to the original node feature to obtain the fused node feature;
[0010] According to the fused node feature, use the temporal convolutional network to enhance the temporal dependence modeling and output the spatio-temporal fused feature matrix;
[0011] Make an inter-satellite handover decision according to the spatio-temporal fused feature matrix.
[0012] Preferably, the physical domain data includes satellite position, link state, and gateway station position; the service domain data includes service type labels and quality of service (QoS) requirements; the environmental domain data includes meteorological grid data and geographical occlusion area information.
[0013] Preferably, the modal encoder for feature extraction of the physical domain data is expressed as:
[0014]
[0015] where, represents the physical domain feature matrix at time t, represents the dynamic adjacency matrix of the physical domain at time t, represents the physical domain output feature matrix at time t, and ST-GCN represents the spatio-temporal graph convolutional network processing.
[0016] Preferably, use the Word2Vec pre-trained word vector model to perform semantic embedding on the service type labels, and dynamically fuse the semantic embedding and QoS parameters through a gating mechanism to generate the service domain output feature matrix.
[0017] Preferably, the modal encoder for feature extraction of the environmental domain data is expressed as:
[0018]
[0019] where A env represents the environmental domain adjacency matrix, represents the output environmental domain feature of the l-th layer of graph convolution, represents the initial input environmental domain feature, and GCN represents the hierarchical graph convolutional network processing.
[0020] Preferably, the process of generating the aligned multi-source data feature matrix includes:
[0021] Design a learnable projection matrix respectively according to the physical domain data, the service domain data, and the environmental domain data;
[0022] According to the learnable projection matrix, map the original feature dimension to the target dimension, and output the aligned multi-source data feature matrix.
[0023] Preferably, the dynamic graph construction process includes:
[0024] Set physical domain nodes, service domain nodes, and environment domain nodes according to the aligned multi-source data feature matrix;
[0025] Set physical edges according to satellite-terminal visibility and gateway-satellite link status, and set semantic edges according to the correlation weight between service requirements and gateway load and the attenuation influence coefficient of meteorological interference on the link;
[0026] Generate a dynamic weighted adjacency matrix based on the nodes and edges;
[0027] The dynamic graph structure is:
[0028] G(t) = (V, E, A(t), X)
[0029] where V represents the set of points in the dynamic graph structure, E represents the set of edges in the dynamic graph structure, A(t) represents the dynamic adjacency matrix, X represents the aligned multi-source data feature matrix, and G(t) represents the dynamic graph at time t.
[0030] Preferably, the process of generating shared token features includes:
[0031] Initialize S learnable token matrices, and splice them with each modality feature to form a shared token fusion input matrix;
[0032] Use a binary mask matrix to constrain the cross-modal interaction path, and process the shared token fusion input matrix through a masked multi-head attention mechanism to generate shared token features, and the process includes:
[0033] Calculate the key, value, and query matrices based on the shared token fusion input matrix, expressed as:
[0034] Q = X fusion ·W Q , K = X fusion ·W K , V = X fusion ·W V
[0035] where Q represents the query matrix, K represents the key matrix, V represents the value matrix, are the first, second, and third learnable parameter matrices, D represents the unified dimension, D h = D / H is the feature dimension of each attention head, and H is the number of attention heads;
[0036] According to the binary mask matrix, calculate the masked attention weights for each attention head independently, expressed as:
[0037]
[0038] where Q i represents the query matrix of the i-th attention head, K i represents the key matrix of the i-th attention head, V i represents the value matrix of the i-th attention head, head i represents the attention weight of the i-th attention head, ⊙ represents element-wise multiplication, and m represents the binary mask matrix;
[0039] Concatenate the outputs of each attention head and obtain the shared token features through a linear transformation, expressed as:
[0040] H shared = Concat(head1,…,head H )·W O
[0041] where H shared represents the shared token features, Concat is the concatenation operation, represents the fourth learnable parameter matrix, and head i represents the attention weight of the i-th attention head;
[0042] Broadcast the shared token features to all nodes and add them to the original features to obtain the fused node features:
[0043]
[0044] where N represents the total number of nodes, D represents the unified dimension, X represents the multi-source data feature matrix, and Repeat means replicating H shared N times.
[0045] Preferably, the temporal enhancement modeling is expressed as:
[0046]
[0047] where N represents the total number of nodes, D represents the unified dimension, TCN represents the temporal convolutional network processing, H temp represents the spatio-temporal fusion feature matrix, and H fused represents the fused node features.
[0048] The beneficial effects of the present invention are:
[0049] The present invention proposes a heterogeneous data fusion method for optimizing inter-satellite handover in a low-earth orbit satellite network. For multi-source heterogeneous data, dedicated encoders are designed respectively to achieve efficient extraction and unified alignment of cross-modal features. By introducing a learnable shared token mechanism and combining masked attention to constrain the cross-modal interaction path, multi-modal data is forced to interact within a unified semantic space, effectively eliminating noise interference and semantic conflicts caused by direct splicing of heterogeneous features, and reducing redundant computational overhead at the same time. By constructing a dynamic graph, the topology change of the satellite network is accurately reflected, and the temporal dependence relationship is enhanced by combining a temporal convolutional network to ensure that the data fusion result is synchronized with the dynamic environment. Through the innovative design of the technical architecture, the present invention realizes the full-process optimization of multi-source heterogeneous data from collection, encoding to fusion, thereby improving the real-time performance and reliability of inter-satellite handover decision-making, providing high-precision, low-overhead, and strong-real-time data support for inter-satellite handover in a low-earth orbit satellite network, and promoting the satellite communication system to move towards the direction of intelligence and adaptability. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following introduces the related technical solution drawings of the embodiments of the present invention. It should be understood that the drawings in the following introduction are only for conveniently and clearly expressing some embodiments of the technical solutions in the present invention. For those skilled in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0051] Figure 1 It is a step diagram of a heterogeneous data fusion method for optimizing inter-satellite handover in a low-earth orbit satellite network provided in an embodiment of the present invention;
[0052] Figure 2 It is a flowchart of fusing cross-modal shared tokens and node information provided in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0053] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention. For the step numbers in the following embodiments, they are only set for the convenience of explanation and illustration, and no limitation is made on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.
[0054] An embodiment of the present invention provides a heterogeneous data fusion method for optimizing inter-satellite handover in a low-earth orbit satellite network. The method steps are as Figure 1 shown.
[0055] In the embodiments of the present invention, multi-source heterogeneous data in a low-earth orbit satellite network is collected in real time, including three types of data: the physical domain, the service domain, and the environmental domain. The physical domain data includes structured time-series data such as satellite position, link status, and gateway station position; service type tags and quality of service (QoS) requirements, where the service type is a discrete semantic tag and the QoS requirement is a continuous numerical value; the environmental domain data includes image-type spatial data such as meteorological grid data and geographical occlusion area information. The above data is respectively subjected to cleaning, normalization, and format standardization processing. For example, the meteorological grid data is divided into uniform grid nodes, the service type tags are converted into vector form through one-hot encoding, and the link status parameters are aligned according to time windows.
[0056] In the embodiments of the present invention, dedicated encoders are designed to extract features according to the characteristics of different modal data, where:
[0057] The physical domain data is processed by a spatio-temporal graph convolutional network (ST-GCN). The numerical time-series data such as satellite position, link status, and gateway station position is input, and the input node feature matrix: and the dynamic adjacency matrix: are used. The spatio-temporal graph convolutional network (ST-GCN) is adopted to fuse the graph structure and time dependence relationship to generate a physical domain feature matrix, denoted as:
[0058]
[0059] where, represents the physical domain feature matrix at time t, represents the dynamic adjacency matrix of the physical domain at time t, represents the output feature matrix of the physical domain at time t, N p represents the number of physical domain nodes, d p represents the physical domain data dimension;
[0060] For the service domain data, a hierarchical coding strategy is adopted. The Word2Vec pre-trained word vector model is used to semantically embed the service type tags as: where d e is the embedding dimension; the continuous QoS parameters (bandwidth, delay, packet loss rate) are normalized as: where N s represents the number of service nodes. The learnable parameter matrix and the bias are introduced. The gating weight is generated through the Sigmoid function as: where G represents the influence weight of the corresponding service type tag on the QoS requirement; the semantic embedding and the QoS parameter are dynamically mixed according to the weight, denoted as:
[0061]
[0062] Among them, H ser represents the service domain feature matrix, E ser represents the business type semantic embedding vector, Q norm represents the normalized QoS metric, ⊙ represents element-wise multiplication, N s represents the number of service domain nodes, d s represents the service domain data dimension;
[0063] The environmental domain data divides the meteorological grid into N e nodes, and the node features include the grid numerical attributes and geographical location coding. Using a hierarchical graph convolutional network (GCN), multiple convolutional kernels are stacked to extract multi-scale features, which are expressed as:
[0064]
[0065] Among them, A env represents the environmental domain adjacency matrix, represents the output environmental domain features of the l-th layer of graph convolution, N e represents the number of environmental domain nodes, d e represents the environmental domain data dimension, represents the initial input environmental domain features.
[0066] In this embodiment, the features of each modality are mapped to a unified dimension through a learnable projection matrix to form an aligned multi-source data matrix, including:
[0067] A learnable projection matrix is designed for each modality, which is expressed as:
[0068]
[0069] Among them, W phy represents the learnable projection matrix of the physical domain, W ser represents the learnable projection matrix of the service domain, W env represents the learnable matrix of the environmental domain, d p represents the physical domain data dimension, d s represents the service domain data dimension, d e represents the environmental domain data dimension, D represents the unified feature dimension;
[0070] Mapping the original feature dimension to the target dimension, which is expressed as:
[0071]
[0072] Among them, represents the aligned physical domain feature matrix, represents the aligned service domain feature matrix, Denote the aligned environmental domain feature matrix, Denote the physical domain output feature matrix at time t, H ser Denote the service domain output feature matrix, Denote the environmental domain output feature matrix;
[0073] Output the aligned multi-source data feature matrix, denoted as:
[0074]
[0075] where X represents the aligned multi-source data feature matrix, N represents the total number of nodes, D represents the unified dimension, N p Denote the number of physical domain nodes, N s Denote the number of service domain nodes, N e Denote the number of environmental domain nodes.
[0076] Based on the aligned feature matrix, the embodiments of the present invention construct a dynamic graph structure. The node set includes physical domain nodes (satellites, terminals, gateway stations), service domain nodes (service types, QoS requirements), and environmental domain nodes (meteorological interference areas, geographical occlusion areas). The edge relationships are divided into physical edges and semantic edges. The weights of physical edges are dynamically adjusted according to the satellite-terminal visibility elevation angle model and link state parameters (such as delay, bandwidth), and the weights of semantic edges are calculated based on the cosine similarity between service requirements and gateway station load, meteorological attenuation coefficient, etc. The adjacency matrix adopts a generated dynamic weighted adjacency matrix, and the weights are updated in real time according to the link state. Finally, the output dynamic graph structure is:
[0077] G(t) = (V, E, A(t), X)
[0078] where V represents the set of nodes of the dynamic graph structure, E represents the set of edges of the dynamic graph structure, A(t) represents the dynamic adjacency matrix, X represents the aligned multi-source data feature matrix, and G(t) represents the dynamic graph at time t.
[0079] The embodiments of the present invention input the dynamic graph node feature matrix X, introduce cross-modal shared tokens as a bridge for information interaction between modalities, and combine the attention mechanism to dynamically control the information flow, so as to solve the problems of redundant calculation and semantic gap in traditional cross-modal fusion. Initialize S learnable initialization tokens T for the feature matrix of each modality, splice the initial tokens of each module with the feature of each modality, and finally combine the attention mechanism to extract the features of the shared tokens as the "transfer station" for information interaction between modalities. The process is as Figure 2 shown, including:
[0080] Initialize S learnable token matrices T, splice them with the aligned multi-source data feature matrix to form a fusion input matrix of shared tokens, and form the input of shared tokens, denoted as:
[0081]
[0082]
[0083] Among them, X fusion represents the fused input matrix, S represents the number of initial tokens, N represents the total number of nodes, D represents the unified dimension, T represents the learnable token matrix, and X represents the aligned multi-source data feature matrix;
[0084] Define a binary mask matrix to constrain the cross-domain interaction path, allowing only shared tokens to interact bidirectionally with all nodes, expressed as:
[0085]
[0086] Use the binary mask matrix to constrain the cross-modal interaction path, and process the shared token fused input matrix through the masked multi-head attention mechanism to generate shared token features. The process includes:
[0087] Calculate the key, value, and query matrices according to the shared token fused input matrix, expressed as:
[0088] Q = X fusion ·W Q , K = X fusion W K , V = X fusion ·W V
[0089] Among them, Q represents the query matrix, K represents the key matrix, and V represents the value matrix, are the first, second, and third learnable parameter matrices, D represents the unified dimension, and D h = D / H is the feature dimension of each attention head, and H is the number of attention heads;
[0090] Calculate the masked attention weights of each attention head independently according to the binary mask matrix, expressed as:
[0091]
[0092] Among them, Q i represents the query matrix of the i-th attention head, K i represents the key matrix of the i-th attention head, V i represents the value matrix of the i-th attention head, head i represents the attention weight of the i-th attention head, ⊙ represents element-wise multiplication, and M represents the binary mask matrix;
[0093] Concatenate the outputs of each attention head and obtain the shared token features through a linear transformation, expressed as:
[0094]
[0095] Among them, H share represents the shared token feature, Concat is the concatenation operation, represents the fourth learnable parameter matrix, head i represents the attention weight of the i-th attention head, and S represents the number of initial tokens;
[0096] Broadcast the shared token feature to all nodes and add it to the original feature to obtain the fused node feature, which is expressed as:
[0097]
[0098] Among them, N represents the total number of nodes, D represents the unified dimension, X represents the multi-source data feature matrix, and Repeat represents replicating H share N times.
[0099] In the embodiment of the present invention, the above steps prohibit the direct interaction between non-shared tokens through masking, only retain the output of the shared tokens as the cross-modal information carrier, broadcast the shared token feature back to each modal node, and splice it with the original feature to inject cross-modal common information, effectively reducing the redundant calculation in cross-modal heterogeneous data fusion and reducing the computational overhead.
[0100] In the embodiment of the present invention, the fused node feature H fused is input into the temporal convolutional network (TCN), combined with dilated causal convolution and residual connection, and the spatio-temporal fusion feature matrix is output, which is expressed as:
[0101]
[0102] Among them, N represents the total number of nodes, D represents the unified dimension, TCN represents the temporal convolutional network processing, H temp represents the spatio-temporal fusion feature matrix, and H fused represents the fused node feature.
[0103] In the embodiment of the present invention, according to the spatio-temporal fusion feature matrix H temp containing the collaborative information of the physical domain, service domain and environment domain, it is mapped to the decision scores of each node through a fully connected layer, the scoring weights are adjusted according to the service type, combined with the environment domain features, the candidate links affected by strong interference are eliminated, the satellite link with the highest score is selected as the handover target, and the handover protocol is triggered.
[0104] In the embodiments of the present invention, by fusing multi-source heterogeneous data, the inter-satellite handover decision mechanism of the low-Earth orbit satellite network is significantly optimized. Physical domain data (such as the real-time position information of satellites and terminals) accurately determines candidate satellites within the visible range of the terminal, providing basic topological support for handover; service domain data dynamically adjusts strategies according to the differential requirements of service types and service quality QoS requirement numerical indicators. For example, for eMBB services, high-bandwidth links are preferentially screened to meet large-volume transmission, while for delay-sensitive uRLLC services, low-delay paths are strictly selected to ensure real-time performance; environmental domain data combines meteorological grid attenuation coefficients and geographical occlusion models to avoid satellite links where signals are interfered by rainfall, clouds or terrain, and predicts and switches to stable channels in advance. Through the collaborative analysis of multi-source heterogeneous data, the system can synchronously weigh physical connection quality, service priority and environmental interference factors, realize refined decision-making for link handover in complex dynamic scenarios, break through the limitations of traditional single-data-driven, and comprehensively improve communication reliability and resource utilization rate.
[0105] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to the embodiments of the present invention without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A heterogeneous data fusion method for optimizing inter-satellite handover in a low-orbit satellite network, characterized in that: The following steps are involved: Real-time collection of multi-source heterogeneous data in low-orbit satellite networks, including physical domain data, service domain data, and environmental domain data, and feature extraction through different modal encoders; Mapping each modality feature to a unified dimension through a learnable projection matrix to generate an aligned multi-source data feature matrix; Based on the aligned multi-source data feature matrix, a dynamic graph structure is constructed; According to the dynamic graph structure, combined with the learnable shared token and multi-head attention mechanism, a shared token feature is generated, and the shared token feature is broadcasted to all nodes and added to the original node feature to obtain a fused node feature; According to the fusion node features, a temporal convolutional network is used to enhance temporal dependency modeling, and a spatiotemporal fusion feature matrix is output; Make inter-satellite switching decisions based on the space-time fusion feature matrix.
2. According to claim 1, a heterogeneous data fusion method for optimizing inter-satellite handover of a low-orbit satellite network is characterized in that: Physical domain data includes satellite position, link status and gateway location; service domain data includes service type labels and service quality QoS requirements; environmental domain data includes meteorological raster data and geographic obstruction area information.
3. The heterogeneous data fusion method for optimizing inter-satellite handover of a low-orbit satellite network according to claim 1 is characterized in that: The modal encoder for feature extraction of physical domain data is expressed as: in, represents the physical domain feature matrix at time t, represents the dynamic adjacency matrix of the physical domain at time t, represents the physical domain output feature matrix at time t, and ST-GCN represents the spatiotemporal graph convolutional network processing.
4. The heterogeneous data fusion method for optimizing inter-satellite handover of a low-orbit satellite network according to claim 1 is characterized in that: The process of extracting environmental domain data features includes: The Word2Vec pre-trained word vector model is used to semantically embed the service type label. The semantic embedding and QoS parameters are dynamically fused through the gating mechanism and learnable weights to generate the service domain output feature matrix.
5. The heterogeneous data fusion method for optimizing inter-satellite handover of a low-orbit satellite network according to claim 1 is characterized in that: The modal encoder for feature extraction of environmental domain data is expressed as: Among them, A env represents the environment domain adjacency matrix, represents the environmental domain features of the l-th layer graph convolution output, represents the initial input environment domain features, and GCN represents the layered graph convolutional network processing.
6. The heterogeneous data fusion method for optimizing inter-satellite handover of a low-orbit satellite network according to claim 1, characterized in that: The process of generating the aligned multi-source data feature matrix includes: Design a learnable projection matrix based on physical domain data, service domain data, and environment domain data respectively; According to the learnable projection matrix, the original feature dimension is mapped to the target dimension, and the aligned multi-source data feature matrix is output.
7. The heterogeneous data fusion method for optimizing inter-satellite handover of a low-orbit satellite network according to claim 1, characterized in that: The dynamic graph construction process includes: According to the aligned multi-source data feature matrix, physical domain nodes, service domain nodes and environment domain nodes are set; Physical edges are set based on satellite-terminal visibility and gateway-satellite link status, and semantic edges are set based on the association weights of business requirements and gateway loads and the attenuation effect coefficient of meteorological interference on the link. Generate a dynamic weighted adjacency matrix based on nodes and edges; The dynamic graph structure is: G(t)=(V,E,A(t),X) Among them, V represents the point set of the dynamic graph structure, E represents the edge set of the dynamic graph structure, A(t) represents the dynamic adjacency matrix, X represents the aligned multi-source data feature matrix, and G(t) represents the dynamic graph at time t.
8. The heterogeneous data fusion method for optimizing inter-satellite handover of a low-orbit satellite network according to claim 1, characterized in that: The process of generating shared token characteristics includes: Initialize S learnable token matrices and concatenate them with the features of each modality to form a shared token fusion input matrix; The binary mask matrix is used to constrain the cross-modal interaction path, and the shared token fusion input matrix is processed through the masked multi-head attention mechanism to generate shared token features. The process includes: The key, value, and query matrices are calculated by fusing the input matrix based on the shared tokens, expressed as: Q=X fusion ·W Q ,K=X fusion ·W K ,V=X fusion ·W V Among them, Q represents the query matrix, K represents the key matrix, V represents the value matrix, and W Q ,W K , are the first, second and third learnable parameter matrices, D represents the unified dimension, and D h =D / H is the feature dimension of each attention head, H is the number of attention heads; According to the binary mask matrix, the mask attention weight of each attention head is calculated independently, expressed as: Among them, Q i represents the query matrix of the i-th attention head, K i represents the key matrix of the i-th attention head, V i Represents the value matrix of the i-th attention head, head i represents the attention weight of the i-th attention head, ⊙ represents element-by-element multiplication, and m represents a binary mask matrix; The outputs of each attention head are concatenated and linearly transformed to obtain the shared token features, which can be expressed as: H share =Concat(head1,…,head H )·W O Among them, H share Indicates shared token features, Concat is a concatenation operation, represents the fourth learnable parameter matrix, head i represents the attention weight of the i-th attention head; Broadcast the shared token features to all nodes and add them to the original features to obtain the fused node features: Among them, N represents the total number of nodes, D represents the unified dimension, X represents the multi-source data feature matrix, and Repeat represents the H share Copy N times.
9. The heterogeneous data fusion method for optimizing inter-satellite handover of a low-orbit satellite network according to claim 1, characterized in that: The timing enhancement modeling is expressed as: Among them, N represents the total number of nodes, D represents the unified dimension, TCN represents the time convolution network processing, H temp represents the spatiotemporal fusion feature matrix, H fused Represents the fusion node feature.
Citation Information
Cited By
Mesh networking and satellite communication-based fusion system
CN121791926A