Feature extraction method and system based on self-attention and hypergraph neural network

By introducing self-attention and hypergraph neural networks into the feature extraction method, the problems of heterogeneity and weak correlation in multimodal data analysis are solved, and a more accurate and complete feature extraction effect is achieved.

CN120180070APending Publication Date: 2025-06-20SHENZHEN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510098304.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The prior art results are inaccurate and incomplete due to heterogeneity and weak correlation when integrating and analyzing multimodal data.

Method used

The feature extraction method based on self-attention and hypergraph neural network is adopted, and the weight of modal data is calculated through the self-attention mechanism, the hypergraph and hypergraph association matrix is ​​constructed, and multi-layer convolution is performed through the multi-modal hypergraph neural network to extract higher-order features.

Benefits of technology

The feature representation and analysis accuracy of multimodal data is improved, the characterization of complex high-order associations between multimodal data is enhanced, and a more comprehensive feature extraction of target objects is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180070A_ABST
    Figure CN120180070A_ABST
Patent Text Reader

Abstract

The invention discloses a feature extraction method and system based on self-attention and a hypergraph neural network, and the method comprises the steps: obtaining various modal data of a target object, constructing a feature processor, and calculating the weight of the modal data through a self-attention mechanism; constructing a corresponding hypergraph according to the modal data, and constructing a hypergraph incidence matrix according to the hypergraph; performing splicing processing on the hypergraph incidence matrix, and outputting a high-order incidence matrix; constructing a multi-modal hypergraph neural network, performing multi-layer convolution processing on the high-order incidence matrix, and constructing and outputting a high-order feature matrix; and classifying the high-order feature matrix according to the weight to obtain a feature extraction result of the target object. According to the method, the multi-modal hypergraph neural network is constructed, meanwhile, a self-attention mechanism is introduced to extract details specific to modals, feature representation of modal data is enhanced, a hypergraph is constructed for various modal data, and complex high-order association is combined to obtain a feature extraction result of a target object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data analysis, and in particular, to a feature extraction method, system, terminal and computer-readable storage medium based on self-attention and hypergraph neural network. Background Art

[0002] Multimodal data helps to comprehensively describe the same object or event, thus enabling a more comprehensive understanding of the problem.

[0003] However, due to the inherent heterogeneity and weak cross-source associations among multimodal data, existing methods have deficiencies in integrating and analyzing multimodal data, resulting in inaccurate data analysis results for the target object and incomplete feature extraction of the target object.

[0004] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention

[0005] The main purpose of the present invention is to provide a feature extraction method, system, terminal and computer-readable storage medium based on self-attention and hypergraph neural network, aiming to solve the problems of inaccuracy and incompleteness in the analysis results of multimodal data due to the heterogeneity and weak cross-source associations among multimodal data in the existing technology.

[0006] To achieve the above object, the present invention provides a feature extraction method based on self-attention and hypergraph neural network, and the feature extraction method based on self-attention and hypergraph neural network includes the following steps:

[0007] Obtain multiple modalities of data of multiple target objects, construct a feature processor, and use the self-attention mechanism to calculate the weight corresponding to each modality of data by using the feature processor;

[0008] Construct a corresponding hypergraph according to each modality of data, and construct a corresponding hypergraph incidence matrix according to the multiple hypergraphs;

[0009] Construct a corresponding hypergraph according to each modality of data, and construct a corresponding hypergraph incidence matrix according to the hypergraph;

[0010] Construct a multimodal hypergraph neural network, splice all the hypergraph incidence matrices to obtain a high-order incidence matrix, and input the high-order incidence matrix into the multimodal hypergraph neural network for multi-layer convolutional processing to output the high-order feature matrices of each convolutional layer;

[0011] Classify each high-order feature matrix according to all the weights to obtain the feature extraction results of all the target objects.

[0012] Optionally, for the feature extraction method based on self-attention and hypergraph neural network, where acquiring various modal data of multiple target objects, constructing a feature processor, and using the self-attention mechanism to calculate the weight corresponding to each piece of modal data by using the feature processor specifically includes:

[0013] Acquire various modal data of multiple target objects, where each target object has various modal data;

[0014] Construct an initial feature processor, and add the self-attention mechanism to the initial feature processor to obtain a feature processor;

[0015] Perform linear transformation on all the modal data to obtain a query vector, a key vector, and a value vector corresponding to each piece of modal data, and input each query vector, each key vector, and each value vector into the feature processor, and output the weight corresponding to each piece of modal data:

[0016]

[0017] where Attention(Q, K, V) represents the weight, Q represents the query vector, K represents the key vector, V represents the value vector, softmax represents the normalization layer, T represents the transpose, and d k represents the dimension of K.

[0018] Optionally, for the feature extraction method based on self-attention and hypergraph neural network, where constructing a corresponding hypergraph according to each piece of modal data and constructing a corresponding hypergraph incidence matrix according to the hypergraph specifically includes:

[0019] Define all the modal data in each piece of modal data as nodes, randomly fix a node, and select a preset number of neighboring nodes closest to the fixed node;

[0020] Connect the fixed node to all the neighboring nodes to obtain hyperedges until a preset number of hyperedges are constructed;

[0021] Construct a corresponding hypergraph for each piece of modal data according to the preset number of hyperedges, where each hypergraph includes a preset number of vertices;

[0022] Construct a corresponding hypergraph incidence matrix according to each hypergraph:

[0023]

[0024] where H represents the hypergraph incidence matrix, v represents the node, and e represents the hyperedge.

[0025] Optionally, in the feature extraction method based on self-attention and hypergraph neural network, when constructing the multi-modal hypergraph neural network, splicing all the hypergraph incidence matrices to obtain a high-order incidence matrix, and inputting the high-order incidence matrix into the multi-modal hypergraph neural network for multi-layer convolution processing to output the high-order feature matrix of each convolution layer, specifically including:

[0026] Construct a multi-modal hypergraph neural network, splice all the hypergraph incidence matrices to obtain a high-order incidence matrix;

[0027] Input the high-order incidence matrix into the multi-modal hypergraph neural network for multi-layer convolution processing, and the multi-modal hypergraph neural network constructs and outputs the high-order feature matrix of each convolution layer.

[0028] Optionally, in the feature extraction method based on self-attention and hypergraph neural network, when constructing the multi-modal hypergraph neural network, splicing all the hypergraph incidence matrices to obtain a high-order incidence matrix, and then further including:

[0029] Add the normalization layer to the multi-modal hypergraph neural network to update the multi-modal hypergraph neural network.

[0030] Optionally, in the feature extraction method based on self-attention and hypergraph neural network, when inputting the high-order incidence matrix into the multi-modal hypergraph neural network for multi-layer convolution processing, and the multi-modal hypergraph neural network constructs and outputs the high-order feature matrix of each convolution layer, specifically including:

[0031] Input the high-order incidence matrix into the updated multi-modal hypergraph neural network, and the multi-modal hypergraph neural network performs multi-layer convolution processing on the high-order incidence matrix to obtain the feature matrix of each convolution layer:

[0032]

[0033] where Y represents the feature matrix, D v represents the degree matrix of nodes, H represents the hypergraph incidence matrix, D e represents the degree matrix of hyperedges, W represents the weight matrix of hyperedges, X represents the input node feature matrix, and Θ represents the learnable parameters of the convolution layer;

[0034] The multi-modal hypergraph neural network uses a non-linear activation function to construct and output the high-order feature matrix of each convolution layer according to the feature matrix of each convolution layer:

[0035]

[0036] where X (l+1)denotes the high - order feature matrix with the convolutional layer number l, where l represents the current layer number of the convolutional layer, σ represents the non - linear activation function, and X (l) denotes the node feature matrix input to convolutional layer l, and Θ (l) denotes the learnable parameters of convolutional layer l.

[0037] Optionally, in the feature extraction method based on self - attention and hypergraph neural network, where classifying each of the high - order feature matrices according to all the weights to obtain the feature extraction results of all the target objects specifically includes:

[0038] Stack the high - order feature matrices of each convolutional layer to obtain a target high - order feature matrix;

[0039] Classify the target high - order feature matrix through the normalization layer and output the feature extraction results of each target object.

[0040] In addition, to achieve the above object, the present invention also provides a feature extraction system based on self - attention and hypergraph neural network, where the feature extraction system based on self - attention and hypergraph neural network includes:

[0041] A weight calculation module, configured to obtain multi - modal data of multiple target objects, construct a feature processor, and use the self - attention mechanism to calculate the weights corresponding to each of the modal data using the feature processor;

[0042] A hypergraph construction module, configured to construct a corresponding hypergraph according to each of the modal data, and construct a corresponding hypergraph incidence matrix according to the multiple hypergraphs;

[0043] A hypergraph splicing module, configured to construct a multi - modal hypergraph neural network, splice all the hypergraph incidence matrices to obtain a high - order incidence matrix, and input the high - order incidence matrix into the multi - modal hypergraph neural network for multi - layer convolutional processing, and output the high - order feature matrices of each convolutional layer;

[0044] A feature extraction module, configured to classify each of the high - order feature matrices according to all the weights to obtain the feature extraction results of all the target objects.

[0045] In addition, to achieve the above object, the present invention also provides a terminal, where the terminal includes: a memory, a processor, and a feature extraction program based on self - attention and hypergraph neural network stored on the memory and executable on the processor. When the feature extraction program based on self - attention and hypergraph neural network is executed by the processor, the steps of the above - described feature extraction method based on self - attention and hypergraph neural network are implemented.

[0046] In addition, to achieve the above object, the present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a feature extraction program based on self-attention and a hypergraph neural network, and when the feature extraction program based on self-attention and a hypergraph neural network is executed by a processor, the steps of the feature extraction method based on self-attention and a hypergraph neural network as described above are implemented.

[0047] In the present invention, multiple modalities of data of multiple target objects are acquired, a feature processor is constructed, and using a self-attention mechanism, the feature processor is used to calculate the weight corresponding to each modality of data; according to each modality of data, a corresponding hypergraph is constructed, and according to the hypergraph, a corresponding hypergraph adjacency matrix is constructed; a multi-modal hypergraph neural network is constructed, the hypergraph adjacency matrices are concatenated to obtain a high-order adjacency matrix, and the high-order adjacency matrix is input into the multi-modal hypergraph neural network for multi-layer convolutional processing to output the high-order feature matrices of each convolutional layer; according to all the weights, each high-order feature matrix is classified to obtain the feature extraction results of all the target objects. The present invention constructs a multi-modal hypergraph neural network and simultaneously introduces a self-attention mechanism to extract modality-specific details and enhance the feature representation of modality data, constructs hypergraphs for multiple modalities of data, and jointly complex high-order associations to obtain the feature extraction results of target objects. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 is a flowchart of a preferred embodiment of the feature extraction method based on self-attention and a hypergraph neural network of the present invention;

[0049] Figure 2 is a detailed flowchart of a preferred embodiment of the feature extraction method based on self-attention and a hypergraph neural network of the present invention;

[0050] Figure 3 is a structural diagram of a preferred embodiment of the feature extraction system based on self-attention and a hypergraph neural network of the present invention;

[0051] Figure 4 is a structural diagram of a preferred embodiment of the terminal of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] To make the object, technical solution and advantages of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0053] The feature extraction method based on self-attention and a hypergraph neural network according to a preferred embodiment of the present invention, as Figure 1As shown, the feature extraction method based on self-attention and hypergraph neural network includes the following steps:

[0054] Step S10: Obtain multi-modal data of multiple target objects, construct a feature processor, and use the self-attention mechanism to calculate the weight corresponding to each piece of the modal data using the feature processor.

[0055] Among them, as shown in Table 1, in this embodiment, the selection of target objects is mainly screened from the ADNI1 (ADNI, Alzheimer's Disease Neuroimaging Initiative, a global research project on Alzheimer's disease and its related dementias) (ADNI1, the first research stage of the Alzheimer's Disease Neuroimaging Initiative), ADNI2 (ADNI2, the second research stage of the Alzheimer's Disease Neuroimaging Initiative), and ADNIGO (a sub-project of ADNI) databases.

[0056] Table 1: Data source table

[0057]

[0058] Among them, AD represents Alzheimer's disease; NC represents Normal Control; EMCI represents Early Mild Cognitive Impairment; LMCI represents Late Mild Cognitive Impairment.

[0059] Among them, as shown in Table 2, the multi-modal data specifically includes structural magnetic resonance imaging (sMRI, Structural Magnetic Resonance Imaging), single nucleotide polymorphism (SNP, Single Nucleotide Polymorphism), cerebrospinal fluid (CSF, Cerebro-Spinal Fluid), and population phenotype data (P, Phenotype Data, including age, gender, and weight), and then it can be analyzed by the SA-MHGNN (Self-attention Multimodal Hypergraph Neural Networks, spatial attention mechanism) model to construct a corresponding hypergraph.

[0060] Table 2: Data modality table

[0061]

[0062] Among them, A-beta (Amyloid beta-protein) represents amyloid beta protein, P-tau (phosphorylated tau protein) represents phosphorylated tau protein, and T-tau (total tau protein) represents total tau protein.

[0063] Specifically, multiple modalities of data of multiple target objects are obtained, where each of the target objects has multiple modalities of data; an initial feature processor is constructed, and a self-attention mechanism is added to the initial feature processor to obtain a feature processor; linear transformation is performed on all the modality data to obtain a query vector, a key vector, and a value vector corresponding to each modality data, and each of the query vectors, each of the key vectors, and each of the value vectors are input into the feature processor to output the weight corresponding to each modality data:

[0064]

[0065] Among them, Attention(Q, K, V) represents the weight, Q represents the query vector, K represents the key vector, V represents the value vector, softmax represents the normalization layer, T represents the transpose, and d k represents the dimension of K.

[0066] Among them, for different modality data, a predefined feature extractor can be used for feature processing. For each modality, through the feature processor in the SA-MHGNN model, a self-attention mechanism is used to capture the long-range dependencies within the modality data.

[0067] Among them, the feature processor can calculate the dot product between the query vector, the key vector, and the value vector corresponding to each modality data to calculate the importance (i.e., weight) of each modality data, and the attention mechanism can dynamically adjust these weights. The attention mechanism allows the model to focus on the relevant parts of the input, helping to reduce noise and redundancy in the multi-modal data by emphasizing the most relevant features and relationships, thereby enhancing the relevant features in the modality data while suppressing the irrelevant features, so as to obtain modality-specific details and feature-enhanced representations, improving the data accuracy.

[0068] Step S20: According to each of the modality data, construct a corresponding hypergraph, and according to the hypergraph, construct a corresponding hypergraph incidence matrix.

[0069] Specifically, all the modal data in each type of the modal data are defined as nodes. Randomly fix one node and select a preset number of neighboring nodes that are closest to the fixed node. Connect the fixed node with all the neighboring nodes to obtain a hyperedge, and continue this process until a preset number of hyperedges are constructed. According to the preset number of hyperedges, a hypergraph corresponding to each type of the modal data is constructed, where each hypergraph includes a preset number of vertices. Based on each hypergraph, a corresponding hypergraph incidence matrix is constructed:

[0070]

[0071] where H represents the hypergraph incidence matrix, v represents a node, and e represents a hyperedge.

[0072] Among them, the method for constructing the hypergraph can be implemented by the KNN algorithm (K-Nearest Neighbors). First, calculate the Euclidean distance between every two nodes. Then, for each node, extract multiple nodes that are closest to it. These nodes can be used to construct a hyperedge. Repeat the above steps to obtain multiple vertices and corresponding multiple hyperedges, thereby constructing a corresponding hypergraph for each type of modal data. Then, a corresponding hypergraph incidence matrix can be constructed based on the hypergraph.

[0073] Among them, as Figure 2 shown, after the feature enhancement processing of the details of the modal data based on the above attention mechanism, the constructed hypergraph and its hypergraph incidence matrix are both representations after feature enhancement. According to this step, the hypergraph features and hyperedge features in each hypergraph can also be obtained, which helps to improve the accuracy of the subsequent model extraction results and helps the model to represent the complex high-order associations between different types of modal data.

[0074] Step S30: Construct a multi-modal hypergraph neural network, perform splicing processing on all the hypergraph incidence matrices to obtain a high-order association matrix, and input the high-order association matrix into the multi-modal hypergraph neural network for multi-layer convolution processing to output the high-order feature matrix of each convolutional layer.

[0075] Specifically, construct a multi-modal hypergraph neural network, perform splicing processing on all the hypergraph incidence matrices to obtain a high-order association matrix;

[0076] Input the high-order association matrix into the multi-modal hypergraph neural network for multi-layer convolution processing, and the multi-modal hypergraph neural network constructs and outputs the high-order feature matrix of each convolutional layer.

[0077] Among them, the multi-modal hypergraph neural network includes a multi-modal hypergraph and a hypergraph convolutional layer. After the multi-modal hypergraph completes the splicing of the hypergraph adjacency matrices of each modality data, a unified multi-modal hypergraph adjacency matrix (i.e., a high-order adjacency matrix) is formed, which can more efficiently capture the high-order correlations in the data and make the hypergraph superior to the traditional graph structure in modeling the complex relationships between features.

[0078] Further, the normalization layer is added to the multi-modal hypergraph neural network to update the multi-modal hypergraph neural network.

[0079] Among them, the normalization layer or the normalization exponential function (softmax function) can be constructed when calculating the feature importance of each modality data (i.e., the weight of each modality data mentioned above). This part can be used to classify the subsequent stacked hypergraph convolutions, so as to realize the feature extraction of the target object.

[0080] Further, use the hypergraph convolutional layer in the hypergraph neural network (i.e., the convolutional layer below) to perform multi-level convolutional processing on the high-order adjacency matrix, and each convolutional processing will generate a corresponding feature matrix.

[0081] Specifically, input the high-order adjacency matrix into the updated multi-modal hypergraph neural network, and the multi-modal hypergraph neural network performs multi-layer convolutional processing on the high-order adjacency matrix to obtain the feature matrix of each convolutional layer:

[0082]

[0083] Among them, Y represents the feature matrix, D v represents the degree matrix of nodes, H represents the hypergraph adjacency matrix, D e represents the degree matrix of hyperedges, W represents the weight matrix of hyperedges, X represents the input node feature matrix, and Θ represents the learnable parameters of the convolutional layer; the multi-modal hypergraph neural network uses a non-linear activation function to construct and output the high-order feature matrix of each convolutional layer according to the feature matrix of each convolutional layer:

[0084]

[0085] Among them, X (l+1) represents the high-order feature matrix with the convolutional layer number l, l represents the current layer number of the convolutional layer, σ represents the non-linear activation function, X (l) represents the node feature matrix input to the convolutional layer l, and Θ (l) represents the learnable parameters of the convolutional layer l.

[0086] Among them, for the high-order feature matrix obtained by each convolutional processing, a non-linear activation function is used for processing to construct the high-order feature matrix.

[0087] Step S40: Classify each of the high-order feature matrices according to all the weights to obtain the feature extraction results of all the target objects.

[0088] Specifically, stack the high-order feature matrices of each convolutional layer to obtain a target high-order feature matrix; classify the target high-order feature matrix through the normalization layer, and output the feature extraction results of each target object.

[0089] Among them, by stacking multiple layers of hypergraph convolutional layers, the high-order features of nodes and the complex relationships between multi-modalities in the modal data can be gradually captured. Finally, the stacked target high-order feature matrix is classified through the softmax function (i.e., the normalization layer) in the hypergraph neural network, and features are extracted by learning the global semantic representation of the nodes. Finally, the feature extraction of different target objects is realized, and the final feature extraction results are output.

[0090] Furthermore, the output feature extraction results can be recognized through the softmax function to obtain the relevant data for AD at different stages of the target object, and can help users identify risk factors in a timely manner. By combining the hypergraph and the self-attention mechanism, the performance of the model can be significantly improved, and the interpretability of the model output results can be increased.

[0091] The present invention constructs a multi-modal hypergraph neural network and simultaneously introduces a self-attention mechanism to extract modality-specific details and enhance the feature representation of modal data. Hypergraphs are constructed for multiple modal data, and complex high-order associations are combined to obtain the feature extraction results of target objects.

[0092] Furthermore, as Figure 3 shown, based on the above feature extraction method based on self-attention and hypergraph neural network, the present invention also correspondingly provides a feature extraction system based on self-attention and hypergraph neural network. Among them, the feature extraction system based on self-attention and hypergraph neural network includes:

[0093] A weight calculation module 51, configured to obtain multi-modal data of multiple target objects, construct a feature processor, and use the self-attention mechanism to calculate the weight corresponding to each modal data by using the feature processor;

[0094] A hypergraph construction module 52, configured to construct a corresponding hypergraph according to each modal data, and construct a corresponding hypergraph incidence matrix according to multiple hypergraphs;

[0095] The hypergraph splicing module 53 is used to construct a multi-modal hypergraph neural network, splice all the hypergraph correlation matrices to obtain a high-order correlation matrix, and input the high-order correlation matrix into the multi-modal hypergraph neural network for multi-layer convolution processing to output the high-order feature matrix of each convolution layer;

[0096] The feature extraction module 54 is used to classify each high-order feature matrix according to all the weights to obtain the feature extraction results of all the target objects.

[0097] Further, as Figure 4 shown, based on the above feature extraction method and system based on self-attention and hypergraph neural network, the present invention also correspondingly provides a terminal, and the terminal includes a processor 10, a memory 20, and a display 30. Figure 4 Only some components of the terminal are shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively.

[0098] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as the hard disk or memory of the terminal. In other embodiments, the memory 20 may also be an external storage device of the terminal, such as a plug-in hard disk equipped on the terminal, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 20 may also include both the internal storage unit and the external storage device of the terminal. The memory 20 is used to store application software installed on the terminal and various types of data, such as program codes for installing the terminal. The memory 20 may also be used to temporarily store data that has been output or will be output. In one embodiment, a feature extraction program 40 based on self-attention and hypergraph neural network is stored on the memory 20, and the feature extraction program 40 based on self-attention and hypergraph neural network can be executed by the processor 10, so as to implement the feature extraction method based on self-attention and hypergraph neural network in the present application.

[0099] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chips, and is used to run the program codes stored in the memory 20 or process data, such as executing the feature extraction method based on self-attention and hypergraph neural network, etc.

[0100] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch liquid crystal display, an OLED (Organic Light-Emitting Diode) touch device, etc. The display 30 is used to display information on the terminal and to display a visual user interface. Components of the terminal communicate with each other via a system bus.

[0101] In one embodiment, when the processor 10 executes the feature extraction program 40 based on self-attention and hypergraph neural network in the memory 20, the following steps are implemented:

[0102] Obtain multi-modal data of multiple target objects, construct a feature processor, and use the self-attention mechanism to calculate the weight corresponding to each piece of the modal data using the feature processor;

[0103] Construct a corresponding hypergraph according to each piece of the modal data, and construct a corresponding hypergraph incidence matrix according to the multiple hypergraphs;

[0104] Construct a corresponding hypergraph according to each piece of the modal data, and construct a corresponding hypergraph incidence matrix according to the hypergraph;

[0105] Construct a multi-modal hypergraph neural network, splice all the hypergraph incidence matrices to obtain a high-order incidence matrix, and input the high-order incidence matrix into the multi-modal hypergraph neural network for multi-layer convolution processing to output the high-order feature matrix of each convolutional layer;

[0106] Classify each high-order feature matrix according to all the weights to obtain the feature extraction results of all the target objects.

[0107] Among them, the obtaining multi-modal data of multiple target objects, constructing a feature processor, and using the self-attention mechanism to calculate the weight corresponding to each piece of the modal data using the feature processor specifically includes:

[0108] Obtain multi-modal data of multiple target objects, where each target object has multi-modal data;

[0109] Construct an initial feature processor, and add the self-attention mechanism to the initial feature processor to obtain a feature processor;

[0110] Perform a linear transformation on all the modal data to obtain a query vector, a key vector, and a value vector corresponding to each modal data, and input each query vector, each key vector, and each value vector into the feature processor to output the weight corresponding to each modal data:

[0111]

[0112] Among them, Attention(Q, K, V) represents the weight, Q represents the query vector, K represents the key vector, V represents the value vector, softmax represents the normalization layer, T represents the transpose, and d k represents the dimension of K.

[0113] Among them, constructing a corresponding hypergraph according to each type of the modality data, and constructing a corresponding hypergraph incidence matrix according to the hypergraph specifically includes:

[0114] Define all the modality data in each type of the modality data as nodes, randomly fix a node, and select a preset number of neighboring nodes closest to the fixed node;

[0115] Connect the fixed node with all the neighboring nodes to obtain hyperedges until a preset number of hyperedges are constructed;

[0116] Construct a corresponding hypergraph for each type of the modality data according to the preset number of hyperedges, where each hypergraph includes a preset number of vertices;

[0117] Construct a corresponding hypergraph incidence matrix according to each hypergraph:

[0118]

[0119] Among them, H represents the hypergraph incidence matrix, v represents the node, and e represents the hyperedge.

[0120] Among them, constructing a multi-modal hypergraph neural network, performing a splicing process on all the hypergraph incidence matrices to obtain a high-order incidence matrix, and inputting the high-order incidence matrix into the multi-modal hypergraph neural network for multi-layer convolution processing, and outputting the high-order feature matrices of each convolution layer specifically includes:

[0121] Construct a multi-modal hypergraph neural network, perform a splicing process on all the hypergraph incidence matrices to obtain a high-order incidence matrix;

[0122] Input the high-order incidence matrix into the multi-modal hypergraph neural network for multi-layer convolution processing, and the multi-modal hypergraph neural network constructs and outputs the high-order feature matrices of each convolution layer.

[0123] Among them, after constructing a multi-modal hypergraph neural network, performing a splicing process on all the hypergraph incidence matrices to obtain a high-order incidence matrix, it further includes:

[0124] Add the normalization layer to the multi-modal hypergraph neural network to update the multi-modal hypergraph neural network.

[0125] Among them, inputting the high-order correlation matrix into the multi-modal hypergraph neural network for multi-layer convolution processing, the multi-modal hypergraph neural network constructs and outputs the high-order feature matrix of each convolution layer, specifically including:

[0126] Input the high-order correlation matrix into the updated multi-modal hypergraph neural network, and the multi-modal hypergraph neural network performs multi-layer convolution processing on the high-order correlation matrix to obtain the feature matrix of each convolution layer:

[0127]

[0128] Among them, Y represents the feature matrix, D v represents the degree matrix of nodes, H represents the hypergraph correlation matrix, D e represents the degree matrix of hyperedges, W represents the weight matrix of hyperedges, X represents the input node feature matrix, and Θ represents the learnable parameters of the convolution layer;

[0129] The multi-modal hypergraph neural network uses a non-linear activation function to construct and output the high-order feature matrix of each convolution layer according to the feature matrix of each convolution layer:

[0130]

[0131] Among them, X (l+1) represents the high-order feature matrix with the convolution layer number l, l represents the current layer number of the convolution layer, σ represents the non-linear activation function, X (l) represents the node feature matrix input to convolution layer l, and Θ (l) represents the learnable parameters of convolution layer l.

[0132] Among them, classifying each of the high-order feature matrices according to all the weights to obtain the feature extraction results of all the target objects specifically includes:

[0133] Stack the high-order feature matrices of each convolution layer to obtain a target high-order feature matrix;

[0134] Classify the target high-order feature matrix through the normalization layer and output the feature extraction results of each target object.

[0135] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a feature extraction program based on self-attention and a hypergraph neural network. When the feature extraction program based on self-attention and a hypergraph neural network is executed by a processor, the steps of the feature extraction method based on self-attention and a hypergraph neural network as described above are implemented.

[0136] In summary, the present invention provides a feature extraction method and related device based on self-attention and hypergraph neural network. The method includes: obtaining multi-modal data of multiple target objects, constructing a feature processor, and using the self-attention mechanism to calculate the weight corresponding to each piece of the multi-modal data by using the feature processor; constructing a corresponding hypergraph according to each piece of the multi-modal data, and constructing a corresponding hypergraph incidence matrix according to the hypergraph; constructing a multi-modal hypergraph neural network, performing splicing processing on all the hypergraph incidence matrices to obtain a high-order incidence matrix, and inputting the high-order incidence matrix into the multi-modal hypergraph neural network for multi-layer convolution processing to output the high-order feature matrix of each convolutional layer; classifying each high-order feature matrix according to all the weights to obtain the feature extraction results of all the target objects. By constructing a multi-modal hypergraph neural network and introducing the self-attention mechanism at the same time, the present invention is used to extract modality-specific details and enhance the feature representation of multi-modal data. Hypergraphs are constructed for multiple modalities of data to jointly capture complex high-order associations, so as to obtain the feature extraction results of target objects.

[0137] It should be noted that in this article, the terms "including", "comprising" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or terminal. Without further limitation, the element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or terminal including the element.

[0138] Of course, those of ordinary skill in the art can understand that all or part of the processes of implementing the above method embodiments can be completed by instructing relevant hardware (such as a processor, a controller, etc.) through a computer program. The program can be stored in a computer-readable storage medium that can be read by a computer. When the program is executed, it can include the processes of the above method embodiments. The computer-readable storage medium can be a memory, a magnetic disk, an optical disk, etc.

[0139] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description. All such improvements and transformations should fall within the protection scope of the appended claims of the present invention.

Claims

1. A feature extraction method based on self-attention and hypergraph neural network, characterized in that: The feature extraction method based on self-attention and hypergraph neural network includes: Acquire multiple modal data of multiple target objects, construct a feature processor, and use the self-attention mechanism to use the feature processor to calculate the weight corresponding to each modal data; According to each of the modal data, a corresponding hypergraph is constructed, and according to the hypergraph, a corresponding hypergraph association matrix is ​​constructed; Constructing a multimodal hypergraph neural network, concatenating all the hypergraph association matrices to obtain a high-order association matrix, and inputting the high-order association matrix into the multimodal hypergraph neural network for multi-layer convolution processing, and outputting a high-order feature matrix of each convolution layer; According to all the weights, each of the high-order feature matrices is classified to obtain feature extraction results of all the target objects.

2. The feature extraction method based on self-attention and hypergraph neural network according to claim 1, characterized in that: The method of acquiring multiple modal data of multiple target objects, constructing a feature processor, and using the self-attention mechanism to calculate the weight corresponding to each modal data using the feature processor specifically includes: Acquire multiple modal data of multiple target objects, wherein each of the target objects has multiple modal data; Constructing an initial feature processor, and adding a self-attention mechanism to the initial feature processor to obtain a feature processor; Performing linear transformation on all the modal data to obtain a query vector, a key vector, and a value vector corresponding to each modal data, and inputting each query vector, each key vector, and each value vector into the feature processor, and outputting a weight corresponding to each modal data: Among them, Attention(Q,K,V) represents the weight, Q represents the query vector, K represents the key vector, V represents the value vector, softmax represents the normalization layer, T represents the transposition, and d k Represents the K dimension.

3. The feature extraction method based on self-attention and hypergraph neural network according to claim 1, characterized in that: The step of constructing a corresponding hypergraph according to each of the modal data, and constructing a corresponding hypergraph association matrix according to the hypergraph, specifically includes: defining all modal data in each of the modal data as nodes, randomly fixing a node, and selecting a preset number of neighboring nodes that are closest to the fixed node; Connecting the fixed node with all the neighboring nodes to obtain hyperedges until a preset number of hyperedges are constructed; According to a preset number of hyperedges, construct a hypergraph corresponding to each type of modal data, wherein each hypergraph includes a preset number of vertices; According to each of the hypergraphs, a corresponding hypergraph association matrix is ​​constructed: Among them, H represents the hypergraph association matrix, v represents the node, and e represents the hyperedge.

4. The feature extraction method based on self-attention and hypergraph neural network according to claim 2, characterized in that: The multimodal hypergraph neural network is constructed, all the hypergraph association matrices are concatenated to obtain a high-order association matrix, and the high-order association matrix is ​​input into the multimodal hypergraph neural network for multi-layer convolution processing, and a high-order feature matrix of each convolution layer is output, which specifically includes: Constructing a multimodal hypergraph neural network, concatenating all the hypergraph association matrices to obtain a high-order association matrix; The high-order association matrix is ​​input into the multimodal hypergraph neural network for multi-layer convolution processing, and the multimodal hypergraph neural network constructs and outputs a high-order feature matrix of each convolution layer.

5. The feature extraction method based on self-attention and hypergraph neural network according to claim 4 is characterized in that: The multimodal hypergraph neural network is constructed, and all the hypergraph association matrices are concatenated to obtain a high-order association matrix, and then the following steps are further included: The normalization layer is added to the multimodal hypergraph neural network, and the multimodal hypergraph neural network is updated.

6. The feature extraction method based on self-attention and hypergraph neural network according to claim 5, characterized in that: The step of inputting the high-order association matrix into the multimodal hypergraph neural network for multi-layer convolution processing, wherein the multimodal hypergraph neural network constructs and outputs a high-order feature matrix of each convolution layer, specifically includes: The high-order association matrix is ​​input into the updated multimodal hypergraph neural network, and the multimodal hypergraph neural network performs multi-layer convolution processing on the high-order association matrix to obtain a feature matrix of each convolution layer: Among them, Y represents the feature matrix, D v represents the degree matrix of the node, H represents the hypergraph association matrix, and D e represents the degree matrix of the hyperedge, W represents the weight matrix of the hyperedge, X represents the input node feature matrix, and Θ represents the learnable parameters of the convolutional layer; The multimodal hypergraph neural network uses a nonlinear activation function to construct and output a high-order feature matrix of each convolutional layer according to the feature matrix of each convolutional layer: Among them, X (l+1) represents the high-order feature matrix with a convolution layer number of l, l represents the current number of convolution layers, σ represents the nonlinear activation function, X ( l ) represents the node feature matrix of the input convolutional layer l, Θ (l) represents the learnable parameters of the convolutional layer l.

7. The feature extraction method based on self-attention and hypergraph neural network according to claim 5, characterized in that: The step of classifying each of the high-order feature matrices according to all the weights to obtain feature extraction results of all the target objects specifically includes: The high-order feature matrix of each convolutional layer is stacked to obtain a target high-order feature matrix; The target high-order feature matrix is ​​classified through the normalization layer, and the feature extraction result of each target object is output.

8. A feature extraction system based on self-attention and hypergraph neural network, characterized in that: The feature extraction system based on self-attention and hypergraph neural network includes: A weight calculation module is used to obtain multiple modal data of multiple target objects, construct a feature processor, and use the self-attention mechanism to use the feature processor to calculate the weight corresponding to each modal data; A hypergraph construction module, used to construct a corresponding hypergraph according to each of the modal data, and to construct a corresponding hypergraph association matrix according to the hypergraph; A hypergraph splicing module is used to construct a multimodal hypergraph neural network, splice all the hypergraph association matrices to obtain a high-order association matrix, and input the high-order association matrix into the multimodal hypergraph neural network for multi-layer convolution processing, and output a high-order feature matrix of each convolution layer; The feature extraction module is used to classify each of the high-order feature matrices according to all the weights to obtain feature extraction results of all the target objects.

9. A terminal, characterized in that: The terminal includes: a memory, a processor, and a feature extraction program based on self-attention and hypergraph neural network stored in the memory and executable on the processor. When the feature extraction program based on self-attention and hypergraph neural network is executed by the processor, the steps of the feature extraction method based on self-attention and hypergraph neural network as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a feature extraction program based on self-attention and hypergraph neural network, and when the feature extraction program based on self-attention and hypergraph neural network is executed by a processor, the steps of the feature extraction method based on self-attention and hypergraph neural network as described in any one of claims 1-7 are implemented.

Citation Information

Cited By

  • Unmanned vehicle cooperative control method based on gating-hypergraph neural network

    CN121069797A