Multi-modal data fusion method and device based on electric power scene and storage medium

By acquiring and encoding multimodal data in power scenarios, generating a complete mask matrix and performing Transformer encoding processing, the problem of poor multimodal data fusion effect in the prior art is solved, and more in-depth modal relationship mining and more efficient data fusion effect are achieved.

CN120145314APending Publication Date: 2025-06-13GUANGDONG POWER GRID CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510309382.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing multimodal data fusion methods rely on linear combination or feature splicing, and cannot deeply explore the potential relationships between different modes, resulting in poor multimodal data fusion effect.

Method used

By obtaining multimodal data of the power scene, encoding process is performed to obtain multimodal features and convert them into feature vectors of the same dimension. Then, the encoding information is added to generate the encoded feature vector, and a variety of mask matrices are generated based on these feature blocks, and combined into a complete mask matrix. The feature blocks in the encoded eigenvector are spliced ​​into a unified feature matrix, and the multimodal features are fused based on the complete mask matrix using the Transformer encoder.

Benefits of technology

It effectively improves the model's understanding and processing capabilities of multimodal data, deeply explores the potential relationships between different modes, and improves the accuracy and effectiveness of multimodal data fusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120145314A_ABST
    Figure CN120145314A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal data fusion method and device based on an electric power scene and a storage medium, and the method comprises the steps: obtaining multi-modal data of the electric power scene, carrying out the coding processing of the multi-modal data, obtaining multi-modal features, and converting the multi-modal features into feature vectors of the same dimension; adding coding information to the feature vector to obtain a coded feature vector; generating a plurality of mask matrixes based on feature blocks in the encoded feature vectors, and combining all the mask matrixes into a complete mask matrix; and feature blocks in the encoded feature vectors are spliced into a unified feature matrix, and the unified feature matrix is input into a Transform encoder, so that the Transform encoder carries out fusion processing on the multi-modal features in the unified feature matrix according to the complete mask matrix, and fusion data are obtained. According to the method, the potential relationship between different modals can be deeply mined, so that the multi-modal data fusion effect can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data fusion, and in particular to a multi-modal data fusion method, device and storage medium based on a power scenario. Background Art

[0002] In a power scenario, the diversity and complexity of data pose great challenges to information processing systems. The operation and maintenance and monitoring of power systems usually involve multi-modal data, such as text records, image monitoring, instrument readings, etc.

[0003] Existing multi-modal data fusion methods usually rely on linear combination or feature splicing, and cannot deeply explore the potential relationships between different modalities, resulting in poor multi-modal data fusion effects. Summary of the Invention

[0004] The present invention provides a multi-modal data fusion method, device and storage medium based on a power scenario to solve the technical problem that existing multi-modal data fusion methods usually rely on linear combination or feature splicing, and cannot deeply explore the potential relationships between different modalities, resulting in poor multi-modal data fusion effects.

[0005] The present invention provides a multi-modal data fusion method based on a power scenario, including:

[0006] Obtain multi-modal data of a power scenario, perform encoding processing on the multi-modal data to obtain multi-modal features, and convert the multi-modal features into feature vectors of the same dimension;

[0007] Add encoding information to the feature vectors to obtain encoded feature vectors;

[0008] Generate multiple mask matrices based on the feature blocks in the encoded feature vectors, and combine all the mask matrices into a complete mask matrix;

[0009] Concatenate the feature blocks in the encoded feature vectors into a unified feature matrix, and input the unified feature matrix into a Transformer encoder, so that the Transformer encoder performs fusion processing on the multi-modal features in the unified feature matrix according to the complete mask matrix to obtain fusion data.

[0010] Further, the encoding processing includes text modality encoding processing and image modality encoding processing. The performing encoding processing on the multi-modal data to obtain multi-modal features and converting the multi-modal features into feature vectors of the same dimension includes:

[0011] Perform text modality encoding processing and image modality encoding processing on the multi-modal data respectively to obtain text modality features and image modality features;

[0012] Based on a linear space transformation, the text modality features and the image modality features are converted into feature vectors of the same dimension.

[0013] Further, the encoding information includes position encoding, modality encoding, and an identifier. Adding the encoding information to the feature vector to obtain an encoded feature vector includes:

[0014] Determine the position encoding according to the feature position, dimension index, and feature dimension, and obtain a first encoding matrix according to the position encoding and the feature matrix of the feature vector; wherein, the feature matrix includes a text feature matrix and an image feature matrix;

[0015] Obtain a second encoding matrix according to the first encoding matrix and the embedding vector of the feature vector; wherein, the embedding vector includes a text modality embedding vector and an image modality embedding vector;

[0016] Insert the identifier into the second encoding matrix to obtain an encoded feature vector.

[0017] Further, the multiple mask matrices include a LookAhead mask matrix, an OnlyBeLook mask matrix, and a Pad mask matrix;

[0018] Among them, the LookAhead mask matrix is used to limit the observation to the feature blocks at and before the current time step in the feature blocks at each time step, and the feature blocks at future time steps cannot be observed;

[0019] The OnlyBeLook mask matrix is used for the feature blocks with unidirectional observation;

[0020] The Pad mask matrix is used to fill the incomplete feature blocks.

[0021] Further, the splicing of the feature blocks in the encoded feature vector into a unified feature matrix includes:

[0022] Set the shape parameters of each feature block, and the shape parameters include batch size, number of time steps, modality feature channels, sequence length, and feature dimension;

[0023] Merge the sequence length and the modality feature channels in each feature block into modality features to obtain updated shape parameters;

[0024] Flatten each feature block based on the updated shape parameters to obtain flattened feature blocks;

[0025] Splice each flattened feature block based on the modality features to obtain a spliced matrix, and the shape parameters of the spliced matrix include batch size, number of time steps, total sum of each modality feature, and feature dimension;

[0026] Flatten the number of time steps and the total sum of modal features of the concatenated matrix to obtain a unified feature matrix.

[0027] Furthermore, the fusion processing of the multi-modal features in the unified feature matrix according to the complete mask matrix to obtain fusion data includes:

[0028] Use a multi-layer processing structure to extract the multi-modal features in the unified feature matrix layer by layer, and perform fusion processing on the extracted multi-modal features to obtain fusion data, where the multi-modal features output by each layer of the processing structure are:

[0029] X out = EncoderLayer(X, mask)

[0030] where X out is the multi-modal feature output by each layer of the processing structure, and mask is the complete mask matrix.

[0031] The present invention provides a multi-modal data fusion device based on a power scenario, including:

[0032] A feature vector acquisition module, configured to acquire multi-modal data of a power scenario, perform encoding processing on the multi-modal data to obtain multi-modal features, and convert the multi-modal features into feature vectors of the same dimension;

[0033] A feature vector encoding module, configured to add encoding information to the feature vector to obtain an encoded feature vector;

[0034] A mask matrix generation module, configured to generate multiple mask matrices based on the feature blocks in the encoded feature vector, and combine all the mask matrices into a complete mask matrix;

[0035] A multi-modal feature fusion module, configured to splice the feature blocks in the encoded feature vector into a unified feature matrix, input the unified feature matrix into a Transformer encoder, and enable the Transformer encoder to perform fusion processing on the multi-modal features in the unified feature matrix according to the complete mask matrix to obtain fusion data.

[0036] Furthermore, the multi-modal feature fusion module is further configured to:

[0037] Set the shape parameters of each feature block, where the shape parameters include batch size, number of time steps, number of modal feature channels, sequence length, and feature dimension;

[0038] Merge the sequence length and the number of modal feature channels in each feature block into modal features to obtain updated shape parameters;

[0039] Flatten each of the feature blocks based on the updated shape parameters to obtain a flattened feature block;

[0040] Based on the modal features, each of the flattened feature blocks is spliced ​​to obtain a spliced ​​matrix, wherein shape parameters of the spliced ​​matrix include a batch size, a time step, a sum of each modal feature, and a feature dimension;

[0041] The time steps of the concatenated matrix and the sum of each modal feature are flattened to obtain a unified feature matrix.

[0042] The present invention also provides a terminal device, comprising: a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, and when the processor executes the computer program, the multimodal data fusion method based on the power scenario as described above is implemented.

[0043] The present invention also provides a computer-readable storage medium, which includes a stored computer program; wherein, when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the multimodal data fusion method based on the power scenario as described above.

[0044] The present invention adds coding information to the feature vector to obtain the encoded feature vector, generates multiple mask matrices based on the feature blocks in the encoded feature vector, and combines all the mask matrices into a complete mask matrix, which can effectively improve the model's understanding and processing capabilities of multimodal data. By combining all the mask matrices into a complete mask matrix, it can ensure that each feature block interacts in accordance with the rules, which can effectively improve the accuracy of multimodal data fusion.

[0045] Furthermore, the present invention concatenates the feature blocks in the encoded feature vector into a unified feature matrix, and uses a Transformer encoder to fuse the multimodal features in the unified feature matrix according to a complete mask matrix to obtain fused data, which can deeply explore the potential relationship between different modalities, thereby effectively improving the multimodal data fusion effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 is a flow chart of a multimodal data fusion method based on a power scenario provided by an embodiment of the present invention;

[0047] Figure 2 is another flow chart of a multimodal data fusion method based on a power scenario provided by an embodiment of the present invention;

[0048] Figure 3 It is a structural schematic diagram of a multimodal data fusion device based on an electric power scenario provided in an embodiment of the present invention. Detailed implementation manners

[0049] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0050] In the description of the present application, it should be understood that the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present application, unless otherwise specified, the meaning of "a plurality" is two or more.

[0051] In the description of the present application, it should be noted that unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific situations.

[0052] Please refer to Figure 1 , an embodiment of the present invention provides a multi-modal data fusion method based on a power scenario, including:

[0053] S1. Obtain multi-modal data of a power scenario, perform encoding processing on the multi-modal data to obtain multi-modal features, and convert the multi-modal features into feature vectors of the same dimension;

[0054] In the embodiment of the present invention, the multi-modal data includes text-modal data and image-modal data. The text-modal data may include numerical measurement data and text description data, and the image-modal data may include image or video data.

[0055] S2. Add encoding information to the feature vector to obtain an encoded feature vector;

[0056] In the embodiment of the present invention, the encoding information includes position encoding, modality encoding, and identifiers. By adding encoding information to the feature vector, the understanding and processing ability of the model for multi-modal data can be effectively improved, so that the order, modality source, and task information of the data can be fully considered during the feature fusion process.

[0057] S3, generating multiple mask matrices based on the feature blocks in the encoded feature vector, and combining all the mask matrices into a complete mask matrix;

[0058] In an embodiment of the present invention, the feature blocks include text task blocks, image task blocks, text observation blocks and image observation blocks, and the mask matrices include LookAhead mask matrices, OnlyBeLook mask matrices and Pad mask matrices. By combining all mask matrices into a complete mask matrix, it is possible to ensure that each feature block interacts in compliance with the rules, thereby effectively realizing effective collaborative processing of multimodal data.

[0059] S4. Concatenate the feature blocks in the encoded feature vector into a unified feature matrix, input the unified feature matrix into a Transformer encoder, and enable the Transformer encoder to fuse the multimodal features in the unified feature matrix according to the complete mask matrix to obtain fused data.

[0060] The embodiment of the present invention adds coding information to the feature vector to obtain a coded feature vector, generates multiple mask matrices based on the feature blocks in the coded feature vector, and combines all mask matrices into a complete mask matrix, which can effectively improve the model's understanding and processing capabilities of multimodal data. By combining all mask matrices into a complete mask matrix, it can ensure that each feature block interacts in compliance with the rules, which can effectively improve the accuracy of multimodal data fusion.

[0061] Furthermore, the embodiment of the present invention splices the feature blocks in the encoded feature vector into a unified feature matrix, and uses a Transformer encoder to fuse the multimodal features in the unified feature matrix according to the complete mask matrix to obtain fused data, which can deeply explore the potential relationship between different modalities, thereby effectively improving the multimodal data fusion effect.

[0062] In one embodiment, the encoding process includes text modality encoding process and image modality encoding process, and step S1, encoding the multimodal data to obtain multimodal features, and converting the multimodal features into feature vectors of the same dimension, includes:

[0063] S11, performing text modality encoding processing and image modality encoding processing on the multimodal data to obtain text modality features and image modality features respectively;

[0064] In the embodiment of the present invention, the text modality encoder may select the BERT (Bidirectional Encoder Representations from Transformers) model. As a bidirectional encoding model, BERT can fully capture the context relationship of text data and provide rich semantic feature representations for text data. The BERT model encodes the input sentence by splitting it into several sub-units (such as words or sub-words) and encodes each sub-unit, converting the text input into a feature vector to ensure the semantic integrity of text information in multi-modal fusion. After being encoded by BERT, the text modality data generates a feature matrix of a fixed length. Let the text sequence length be L and the hidden dimension be D, then the output text feature matrix is T ∈ R L×D . When the text sequence length is less than the maximum length, padding operations can be performed to make each feature matrix consistent, facilitating subsequent fusion with other modal features. Among them, the generation formula of the text feature matrix is T = BERT(X text ), X text is the text input vector, and T represents the generated feature matrix.

[0065] In the embodiment of the present invention, the image modality encoder may select the ViT (Vision Transformer) model. Based on the Transformer structure, the ViT model can divide an image into multiple small patches and encode these small patches to extract the spatial structure information of the image. The image modality data is first divided into patches of a fixed size. Assuming the size of the input image is H×W×C (height, width, number of channels), patches of size p×p are obtained, and N = (H×W) / p 2 patches are generated. The ViT model converts each image patch into a feature vector and finally generates an image feature matrix I ∈ R N ×D . The generation formula of this image feature matrix is I = ViT(X image ), where X image is the image input vector, and I is the output image feature matrix. The image data is first loaded through the PIL or OpenCV library and necessary preprocessing is performed to ensure that the image sizes are consistent for subsequent encoding operations.

[0066] S12. Based on linear space transformation, convert the text modality features and image modality features into feature vectors of the same dimension.

[0067] In an embodiment of the present invention, based on a linear space transformation, the text modality features and the image modality features are converted into feature vectors of the same dimension. Among them, the linear space transformation can be implemented through a fully connected layer (i.e., a linear layer) to project the input features onto the target feature dimension. The linear transformation formula can be:

[0068] y = W·x + b

[0069] where x represents the input feature vector, W is the weight matrix, b is the bias term, and y is the transformed feature of the output. After the linear space transformation, the dimensions of the text feature matrix T and the image feature matrix I are both set to T, I ∈ R N×D , where N represents the sequence length or the number of blocks, and D is the unified feature dimension.

[0070] In the embodiment of the present invention, by setting the input dimension and the output dimension to be the same, the spatial unification of features can be achieved through linear space transformation, which is beneficial to improving the accuracy and reliability of subsequent multi-modal feature fusion.

[0071] In one embodiment, the encoded information includes position encoding, modality encoding, and identifiers. Step S2, adding encoded information to the feature vector to obtain an encoded feature vector, includes:

[0072] S211. Determine the position encoding according to the feature position, dimension index, and feature dimension, and obtain the first encoding matrix according to the position encoding and the feature matrix of the feature vector; where the feature matrix includes a text feature matrix and an image feature matrix;

[0073] In the embodiment of the present invention, fixed sine and cosine functions can be used to generate position encoding, so that data at different positions have unique representations in the encoding space. The formula for position encoding is as follows:

[0074]

[0075] where pos is the encoding position, i is the dimension index, and d is the feature dimension. The position encoding uses sine and cosine functions, applying the sine function to even dimensions and the cosine function to odd dimensions to generate a unique encoding vector.

[0076] In actual use, the position encoding adds the position vector and the feature matrix element by element during forward propagation to obtain the first encoding matrix. For example, assuming that the text feature matrix is T′, and its dimension is L×D (L is the sequence length, D is the feature dimension), then after adding the position encoding, we get:

[0077] T pos = T′ + PE

[0078] where T posis the first encoding matrix.

[0079] In the embodiment of the present invention, the first encoding matrix is obtained by adding position encoding, which can effectively help the model understand the time series or spatial relationship of features, thereby enhancing the sensitivity of the model to sequential information.

[0080] S212. Obtain a second encoding matrix according to the first encoding matrix and the embedding vectors of the feature vectors; wherein, the embedding vectors include a text modality embedding vector and an image modality embedding vector;

[0081] In the embodiment of the present invention, the expression of the second encoding matrix is as follows:

[0082] T″ = T pos +ME Text

[0083] I″ = I pos +ME Image

[0084] wherein, the second encoding matrix includes a second text feature matrix T″ and a second image feature matrix I″, T pos 、I pos is the first encoding matrix, ME Text is the text modality embedding quantity, ME Image is the image modality embedding quantity.

[0085] In the embodiment of the present invention, fixed modality embedding vectors are generated through the embedding layer and added to the feature matrices of each modality, enabling the model to identify different modality information during the fusion process.

[0086] S213. Insert the identifier into the second encoding matrix to obtain the encoded feature vector.

[0087] In the embodiment of the present invention, the expression of the encoded feature vector is as follows:

[0088] S = [[TEXT], T″, [SEP], [IMAGE], I″]

[0089] wherein, S is the encoded feature vector, [SEP] is used to separate different tasks or feature blocks, [TEXT] is used to identify text modality features, enabling the model to recognize that the current sequence is text data, and [IMAGE] is used to identify image modality features, enabling the model to recognize that the current sequence is image data.

[0090] In the embodiment of the present invention, by inserting identifiers, the model can maintain a clear boundary in multi-modal feature processing and task differentiation, thereby effectively improving the accuracy of multi-modal data fusion.

[0091] In the embodiments of the present invention, by adding position encoding, modality encoding, and identifiers to the feature vectors, the understanding and processing capabilities of the model for multimodal data are effectively improved, so that the order, modality source, and task information of the data can be fully considered during the feature fusion process, and the dynamic collaborative fusion of multimodal data can be accurately achieved, which is beneficial to improving the accuracy of multimodal data fusion.

[0092] In one embodiment, the multiple mask matrices include a LookAhead mask matrix, an OnlyBeLook mask matrix, and a Pad mask matrix;

[0093] Among them, the LookAhead mask matrix is used to limit the observation to the feature blocks at and before the current time step in the feature blocks at each time step, and the feature blocks at future time steps cannot be observed;

[0094] The OnlyBeLook mask matrix is used for feature blocks with unidirectional observation;

[0095] The Pad mask matrix is used to fill in incomplete feature blocks.

[0096] In the embodiments of the present invention, the feature blocks in the encoded feature vectors in step S2 include text task blocks, image task blocks, text observation blocks, and image observation blocks, where:

[0097] The text task block is a text feature block for sending instructions to the model, which is obtained by encoding the text input and represents the task or expected text output.

[0098] The image task block is an image feature block for sending instructions to the model, which is obtained by encoding the image input and represents the task or expected image output.

[0099] The text observation block is an observation feature block received by the model from the text sensor, representing the real-time perception and input information of the model for the external text environment.

[0100] The image observation block is an observation feature block received by the model from the image sensor, representing the real-time perception and input information of the model for the external image environment.

[0101] In the embodiments of the present invention, each type of feature block has its specific interaction rules and mask requirements, which are described in detail as follows:

[0102] The LookAhead mask is a block upper triangular mask used for time series tasks. Its purpose is to limit the model to only observe the feature blocks at and before the current time step in the feature blocks at each time step, and the feature blocks at future time steps cannot be observed, so as to ensure the causal relationship in the time series and enable the model to make predictions only relying on past information.

[0103] The specific generation method of the LookAhead mask is as follows:

[0104] Let the total number of time steps be T, and the size of the feature block at each time step be L. Then the size of the block upper triangular mask matrix is (T×L)×(T×L). In the matrix, the feature blocks at each time step form a block upper triangular structure, and its representation is as follows:

[0105]

[0106] Among them, for each time step i, only the feature blocks at this time step and the previous time steps are allowed to perform attention interaction. When j > i, the feature blocks at time step j are masked, so that the model does not obtain the feature information of future time steps.

[0107] The LookAhead mask in the embodiments of the present invention can effectively prevent the model from accessing the feature blocks of future time step j at time step i, thereby maintaining the causal relationship. For example, for the case of T = 3 and L = 2, the generated mask is a block upper triangular matrix.

[0108] The OnlyBeLook mask is used for the feature blocks of one-way observation, that is, these feature blocks can only be observed by other modalities, and do not actively observe the information of other modalities. For example, for the "text task" and "image task" feature blocks, the task data can be obtained by the observation data ("text observation" and "image observation"), but will not actively obtain the information of other modalities.

[0109] The specific generation method of the OnlyBeLook mask is as follows:

[0110] Let the size of the feature matrix be N×N, and assume that the size of the data source block is M.

[0111] For the "text task block" or "image task block", set it in the mask matrix so that it does not perform attention operations on other modalities, but allows other modalities to observe.

[0112] The formula expression is as follows:

[0113]

[0114] In implementation, by setting the rows of this feature block to 1, a one-way attention mask is formed.

[0115] The Pad mask is used to fill the incomplete feature blocks to maintain the consistency of the feature blocks. For text data, if the input length is less than the maximum length L, it is filled at the end, and the corresponding area in the mask matrix is set to 1 to prevent the extra filling from affecting other feature blocks.

[0116] The formula expression of the Pad mask is:

[0117] Pad (i,j) = 1, for padding positions

[0118] The combination method of the complete mask matrix is as follows:

[0119] Final Mask = LookAh ead ∨ OnlyBeLook ∨ Pad

[0120] In the embodiment of the present invention, all mask matrices are combined into a complete mask matrix, and the masks are combined to form a final complete mask matrix, ensuring that each feature block interacts under the rules, so as to effectively realize the effective collaborative processing of multi-modal data.

[0121] In one embodiment, step S4, the splicing of the feature blocks in the encoded feature vectors into a unified feature matrix includes:

[0122] S41. Set the shape parameters of each feature block, where the shape parameters include batch size, number of time steps, number of modal feature channels, sequence length, and feature dimension;

[0123] In the embodiment of the present invention, the shape parameters of each feature block are set as (batch, time, channel, seq, dim), where batch represents the batch size, time represents the number of time steps, channel represents the number of channels of the modal features, seq represents the length of the sequence, and dim represents the feature dimension.

[0124] S42. Combine the sequence length and the number of modal feature channels in each feature block into modal features to obtain updated shape parameters;

[0125] In the embodiment of the present invention, the two dimensions of channel and seq are combined, so that the shape parameters of the feature blocks of each modality are adjusted to (batch, time, features, dim), where features represents the expanded modal feature dimension.

[0126] S43. Flatten each feature block based on the updated shape parameters to obtain a flattened feature block;

[0127] In the embodiment of the present invention, the expression of the flattened feature block can be:

[0128] M i = view(M i , batch, time, -1, dim)

[0129] where M i is the flattened feature block, and -1 represents the combination of channel and seq.

[0130] After the flattening in the embodiments of the present invention, the feature blocks retain the independence of the modality and the chronological information, providing a basis for the next splicing.

[0131] S44. Splice each of the flattened feature blocks based on the modality features to obtain a spliced matrix, where the shape parameters of the spliced matrix include batch size, time steps, the sum of each modality feature, and feature dimension;

[0132] In the embodiments of the present invention, after obtaining the flattened feature blocks, the feature data of different modalities are spliced together along the features dimension to generate a spliced matrix containing multi-modal information. The spliced matrix C is expressed as:

[0133] C = concat(M 1 , M 2 , …, M N , dim = 2)

[0134] where concat represents the splicing operation of tensors, and the shape of C is (batch, time, total_features, dim). At this time, total_features is the sum of the features of each modality.

[0135] S45. Flatten the time steps and the sum of each modality feature of the spliced matrix to obtain a unified feature matrix.

[0136] In the embodiments of the present invention, the time steps time and the feature sum total_features are further flattened to generate a unified feature matrix final_result with a shape of (batch, time × total_features, dim), so as to fully expand the time steps and modality features for the Transformer encoder to process.

[0137] Among them, the expression of the unified feature matrix is as follows:

[0138] final result = view(C, batch, time × total_features, dim).

[0139] In the embodiments of the present invention, through flattening and splicing processing, the feature blocks in the encoded feature vectors are spliced into a unified feature matrix, which includes information of each modality. Through appropriate flattening operations, the independent features of each modality can be retained, and the unified feature matrix allows information interaction and fusion between different modalities. The model can learn the correlation and dependence between different modality features, thereby realizing deeper information fusion and effectively improving the accuracy of multi-modal data fusion.

[0140] In one embodiment, fusing the multimodal features in the unified feature matrix according to the complete mask matrix to obtain fusion data includes:

[0141] Using a multi-layer processing structure to extract the multimodal features in the unified feature matrix layer by layer, and performing fusion processing on the extracted multimodal features to obtain fusion data, where the multimodal features output by each layer of the processing structure are:

[0142] X out = EncoderLayer(X, mask)

[0143] where X out is the multimodal feature output by each layer of the processing structure, and mask is the complete mask matrix.

[0144] In an embodiment of the present invention, the unified feature matrix is input into the Transformer encoder as input information X. The encoder consists of multiple Encoder layers, and each layer contains a multi-head self-attention mechanism and a feed-forward network, which can capture the complex relationships between multimodal features.

[0145] In an embodiment of the present invention, the multi-layer processing structure in the Transformer encoder extracts and fuses multimodal features layer by layer, and captures the long-range dependencies between different modal features through the multi-head self-attention mechanism. The calculation of the multi-head self-attention can be expressed as:

[0146]

[0147] where Q, K, and V are the representations of the query, key, and value of the feature matrix respectively, d k is the dimension of the key, which is used for scaling to ensure the numerical stability of the calculation. Mask is the complete mask matrix, making the attention scores close to zero at the masked positions.

[0148] Through multi-layer encoding of the Transformer encoder in the embodiments of the present invention, the potential relationships between multimodal features can be effectively captured, and the dynamic collaborative fusion of features can be realized. It can not only effectively improve the utilization efficiency of the model for multimodal features, but also effectively improve the accuracy and reliability of multimodal data fusion.

[0149] Please refer to Figure 2 , which is another flow schematic diagram of a multimodal data fusion method based on a power scenario provided by an embodiment of the present invention.

[0150] Implementing the embodiments of the present invention has the following beneficial effects:

[0151] In an embodiment of the present invention, by adding coding information to the feature vector, an encoded feature vector is obtained. Based on the feature blocks in the encoded feature vector, multiple mask matrices are generated, and all the mask matrices are combined into a complete mask matrix, which can effectively improve the model's understanding and processing ability of multimodal data. By combining all the mask matrices into a complete mask matrix, it can ensure that each feature block interacts under the rules, and can effectively improve the accuracy of multimodal data fusion.

[0152] Furthermore, in an embodiment of the present invention, by splicing the feature blocks in the encoded feature vector into a unified feature matrix, and using a Transformer encoder to fuse the multimodal features in the unified feature matrix according to the complete mask matrix to obtain fusion data, the potential relationships between different modalities can be deeply mined, thereby effectively improving the multimodal data fusion effect.

[0153] Please refer to Figure 3 , based on the same inventive concept as the above embodiment, the present invention provides a multimodal data fusion device based on a power scenario, including:

[0154] A feature vector acquisition module 10, configured to acquire multimodal data of a power scenario, perform encoding processing on the multimodal data to obtain multimodal features, and convert the multimodal features into feature vectors of the same dimension;

[0155] A feature vector encoding module 20, configured to add coding information to the feature vector to obtain an encoded feature vector;

[0156] A mask matrix generation module 30, configured to generate multiple mask matrices based on the feature blocks in the encoded feature vector, and combine all the mask matrices into a complete mask matrix;

[0157] A multimodal feature fusion module 40, configured to splice the feature blocks in the encoded feature vector into a unified feature matrix, input the unified feature matrix into a Transformer encoder, and enable the Transformer encoder to fuse the multimodal features in the unified feature matrix according to the complete mask matrix to obtain fusion data.

[0158] In one embodiment, the feature vector acquisition module 10 is further configured to:

[0159] Perform text modality encoding processing and image modality encoding processing on the multimodal data respectively to obtain text modality features and image modality features;

[0160] Based on linear space transformation, convert the text modality features and image modality features into feature vectors of the same dimension.

[0161] In one embodiment, the encoded information includes position encoding, modality encoding, and an identifier, and the feature vector encoding module 20 is further configured to:

[0162] Determine the position encoding according to the feature position, dimension index, and feature dimension, and obtain a first encoding matrix based on the position encoding and the feature matrix of the feature vector; wherein, the feature matrix includes a text feature matrix and an image feature matrix;

[0163] Obtain a second encoding matrix based on the first encoding matrix and the embedding vector of the feature vector; wherein, the embedding vector includes a text modality embedding vector and an image modality embedding vector;

[0164] Insert the identifier into the second encoding matrix to obtain the encoded feature vector.

[0165] In one embodiment, the multiple mask matrices include a LookAhead mask matrix, an OnlyBeLook mask matrix, and a Pad mask matrix;

[0166] Among them, the LookAhead mask matrix is used to limit the observation in the feature blocks at each time step to only the feature blocks at the current time step and those before it, and not to the feature blocks at future time steps;

[0167] The OnlyBeLook mask matrix is used for feature blocks with unidirectional observation;

[0168] The Pad mask matrix is used to fill in incomplete feature blocks.

[0169] In one embodiment, the multi-modal feature fusion module 40 is further configured to:

[0170] Set the shape parameters for each feature block, and the shape parameters include batch size, number of time steps, number of modality feature channels, sequence length, and feature dimension;

[0171] Combine the sequence length and the number of modality feature channels in each feature block into modality features to obtain updated shape parameters;

[0172] Flatten each of the feature blocks based on the updated shape parameters to obtain flattened feature blocks;

[0173] Concatenate each of the flattened feature blocks based on the modality features to obtain a concatenated matrix, and the shape parameters of the concatenated matrix include batch size, number of time steps, sum of each modality feature, and feature dimension;

[0174] Flatten the number of time steps and the sum of each modality feature in the concatenated matrix to obtain a unified feature matrix.

[0175] Correspondingly, an embodiment of the present invention further provides a terminal device, including: a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the multi-modal data fusion method based on the power scenario in any one of the above embodiments is implemented.

[0176] The terminal device of this embodiment includes: a processor, a memory, and a computer program and computer instructions stored in the memory and executable on the processor. When the processor executes the computer program, each step in the above Embodiment 1 is implemented, such as Figure 1 the steps S1 to S4 shown. Alternatively, when the processor executes the computer program, the functions of each module / unit in the above device embodiment are implemented, such as the multi-modal feature fusion module 40.

[0177] Exemplarily, the computer program can be divided into one or more modules / units. One or more modules / units are stored in the memory and executed by the processor to complete the present invention. One or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the terminal device. For example, the multi-modal feature fusion module 40 is used to splice the feature blocks in the encoded feature vector into a unified feature matrix, input the unified feature matrix into the Transformer encoder, and enable the Transformer encoder to perform fusion processing on the multi-modal features in the unified feature matrix according to the complete mask matrix to obtain fusion data.

[0178] The terminal device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The terminal device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that the schematic diagram is only an example of the terminal device and does not constitute a limitation on the terminal device. It may include more or fewer components than shown, or combine certain components, or different components. For example, the terminal device may further include input / output devices, network access devices, a bus, etc.

[0179] The so-called processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the terminal device, connecting all parts of the entire terminal device through various interfaces and circuits.

[0180] The memory can be used to store computer programs and / or modules. The processor realizes various functions of the terminal device by running or executing the computer programs and / or modules stored in the memory, and by calling the data stored in the memory. The memory may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function, etc.; the data storage area may store data created according to the use of the mobile terminal, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices.

[0181] Among them, if the modules / units integrated in the terminal device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of the present invention, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0182] Correspondingly, an embodiment of the present invention further provides a computer-readable storage medium. The computer-readable storage medium includes a stored computer program, wherein when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the multi-modal data fusion method based on the power scenario in any one of the above embodiments.

[0183] The above specific embodiments have further elaborated on the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. It is particularly pointed out that for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A multimodal data fusion method based on power scenarios, characterized in that: include: Acquire multimodal data of a power scene, encode the multimodal data to obtain multimodal features, and convert the multimodal features into feature vectors of the same dimension; Adding encoding information to the feature vector to obtain an encoded feature vector; Generating a plurality of mask matrices based on the feature blocks in the encoded feature vector, and combining all the mask matrices into a complete mask matrix; The feature blocks in the encoded feature vector are spliced ​​into a unified feature matrix, and the unified feature matrix is ​​input into a Transformer encoder, so that the Transformer encoder performs fusion processing on the multimodal features in the unified feature matrix according to the complete mask matrix to obtain fused data.

2. The multimodal data fusion method based on power scenarios according to claim 1 is characterized in that: The encoding process includes text modality encoding process and image modality encoding process, the encoding process is performed on the multimodal data to obtain multimodal features, and the multimodal features are converted into feature vectors of the same dimension, including: Performing text modality encoding processing and image modality encoding processing on the multimodal data to obtain text modality features and image modality features respectively; Based on linear space transformation, the text modality features and image modality features are converted into feature vectors of the same dimension.

3. The multimodal data fusion method based on power scenarios according to claim 1, characterized in that: The coding information includes a position code, a modality code and an identifier, and the adding of the coding information to the feature vector to obtain the encoded feature vector includes: Determine the position code according to the feature position, dimension index and feature dimension, and obtain a first coding matrix according to the position code and the feature matrix of the feature vector; wherein the feature matrix includes a text feature matrix and an image feature matrix; Obtaining a second encoding matrix according to the first encoding matrix and the embedding vector of the feature vector; wherein the embedding vector includes a text modality embedding vector and an image modality embedding vector; The identifier is inserted into the second encoding matrix to obtain an encoded feature vector.

4. The multimodal data fusion method based on power scenarios according to claim 1, characterized in that: The multiple mask matrices include a LookAhead mask matrix, an OnlyBeLook mask matrix and a Pad mask matrix; The LookAhead mask matrix is ​​used in the feature blocks of each time step to limit the observation of only the feature blocks of the time step and before, and the feature blocks of the future time step cannot be observed; The OnlyBeLook mask matrix is ​​used for feature blocks observed in one direction; The Pad mask matrix is ​​used to fill in the incomplete feature blocks.

5. The multimodal data fusion method based on power scenarios according to claim 1, characterized in that: The step of splicing the feature blocks in the encoded feature vector into a unified feature matrix includes: Setting shape parameters of each feature block, wherein the shape parameters include batch size, number of time steps, number of modal feature channels, sequence length, and feature dimension; The sequence length and the number of modal feature channels in each feature block are combined into modal features to obtain updated shape parameters; Flatten each of the feature blocks based on the updated shape parameters to obtain a flattened feature block; Based on the modal features, each of the flattened feature blocks is spliced ​​to obtain a spliced ​​matrix, wherein shape parameters of the spliced ​​matrix include a batch size, a time step, a sum of each modal feature, and a feature dimension; The time steps of the concatenated matrix and the sum of each modal feature are flattened to obtain a unified feature matrix.

6. The multimodal data fusion method based on power scenarios according to claim 1, characterized in that: The fusing the multimodal features in the unified feature matrix according to the complete mask matrix to obtain fused data includes: The multimodal features in the unified feature matrix are extracted layer by layer using a multi-layer processing structure, and the extracted multimodal features are fused to obtain fused data, wherein the multimodal features output by each layer of the processing structure are: X out =EncoderLayer(X,mask) Among them, X out The multimodal features output by each layer of the processing structure, mask is the complete mask matrix.

7. A multimodal data fusion device based on power scenarios, characterized in that: include: A feature vector acquisition module, used to acquire multimodal data of a power scenario, encode the multimodal data to obtain multimodal features, and convert the multimodal features into feature vectors of the same dimension; A feature vector encoding module, used for adding encoding information to the feature vector to obtain an encoded feature vector; A mask matrix generation module, used to generate multiple mask matrices based on the feature blocks in the encoded feature vector, and combine all mask matrices into a complete mask matrix; The multimodal feature fusion module is used to splice the feature blocks in the encoded feature vector into a unified feature matrix, input the unified feature matrix into a Transformer encoder, and enable the Transformer encoder to fuse the multimodal features in the unified feature matrix according to the complete mask matrix to obtain fused data.

8. The multimodal data fusion device based on electric power scenarios according to claim 7, characterized in that: The multimodal feature fusion module is also used for: Setting shape parameters of each feature block, wherein the shape parameters include batch size, number of time steps, number of modal feature channels, sequence length, and feature dimension; The sequence length and the number of modal feature channels in each feature block are combined into modal features to obtain updated shape parameters; Flatten each of the feature blocks based on the updated shape parameters to obtain a flattened feature block; Based on the modal features, each of the flattened feature blocks is spliced ​​to obtain a spliced ​​matrix, wherein shape parameters of the spliced ​​matrix include a batch size, a time step, a sum of each modal feature, and a feature dimension; The time steps of the concatenated matrix and the sum of each modal feature are flattened to obtain a unified feature matrix.

9. A terminal device, characterized in that: include: A processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, the multimodal data fusion method based on the power scenario as described in any one of claims 1 to 6 is implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored computer program; wherein, when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the multimodal data fusion method based on the power scenario as described in any one of claims 1-6.

Citation Information

Cited By

  • Multi-modal named entity recognition method and device based on image-text feature deep interaction

    CN121581045A

  • Multimodal named entity recognition method and device based on depth interaction of image-text features

    CN121581045B