A multi-modal data fusion processing method, device, equipment and storage medium

By combining feature mapping and neural network models with Pearson correlation coefficient, Taylor series fusion strategy and attention mechanism, the problem of insufficient information correlation in multimodal data fusion is solved, and the accuracy of data fusion is improved efficiently.

CN117216714BActive Publication Date: 2026-03-27INFORMATION & COMM BRANCH OF STATE GRID JIANGSU ELECTRIC POWER +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-04
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing technologies neglect the information correlation between data features in multimodal data fusion processing, resulting in the waste of useful information. Furthermore, feature fusion mostly remains at the low-level data level and fails to effectively utilize high-level information.

Method used

The original multimodal data is transformed into a multimodal augmented dataset by pre-setting feature mapping relationships. The feature fusion matrices within and between modalities are determined. The intermodal feature fusion is performed using a neural network model. Feature expansion and fusion are carried out by combining Pearson correlation coefficient, Taylor series fusion strategy and self and cross attention mechanism.

Benefits of technology

It improves the accuracy of multimodal data fusion processing, effectively captures relevant information within and between modes, avoids wasting useful information, and enhances data fusion performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117216714B_ABST
    Figure CN117216714B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a kind of multimodal data fusion processing method, device, equipment and storage medium.The method includes obtaining the original multimodal data set to be processed in power internet of things;Based on the preset feature mapping relationship, the original feature vector corresponding to each modality target data in each modality is mapped to obtain multimodal enhanced data set;Determine the relationship fusion matrix between the respective corresponding modal internal features of each enhanced modality in multimodal enhanced data set, and each enhanced modality and the respective corresponding relationship fusion matrix of each enhanced modality are fused to obtain multimodal feature fusion data set;Based on the initial modal feature of each fusion modality respectively corresponding in multimodal feature fusion data set and the feature fusion between modalities of each fusion modality by neural network model.The embodiments of the present application, by the above technical solution, solve the consistency and difference problem of intra-modal and inter-modal, improve the performance of multimodal data fusion processing and data fusion accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of electric power data processing, and in particular to a multi-modal data fusion processing method, device, equipment and storage medium. BACKGROUND

[0002] With the in-depth promotion of new power system construction, power Internet of Things sensing devices are increasing in a large amount, source, network, storage and load in the power system are deeply interacted, multi-element heterogeneous energy flow and information flow composed of data are deeply coupled, and its analysis must be changed from traditional isolated analysis mode to collaborative analysis of each link. On the one hand, a large number of various types of sensors arranged in the power system will produce a large amount of data related to power business, involving different sources such as power and non-power, and being represented in different forms such as numbers, images and texts. On the other hand, the development of big data, cloud platform and other technologies makes the data barriers between different fields and businesses broken down, and data sharing is more convenient. The data related to power business will be more extensive and diverse, and the multi-type, multi-structure and multi-source form provides a comprehensive and effective data basis for coping with various challenges of the power system. However, the multi-source, heterogeneity and massiveness of multi-modal data in the power Internet of Things make its interactive processing face many challenges for a long time. In order to realize data processing for differentiated business, high-speed and efficient processing of massive data is required in the transmission layer and processing layer. From the perspective of data utilization and fusion, the correlation and comprehensive analysis of multi-modal data are required to realize accurate, unified and data fusion of each link and each scene, which is an important foundation support for realizing full observation and full control of the power system.

[0003] At present, the information correlation between data features is often ignored in the research of multi-modal data fusion processing, resulting in waste of useful information, and feature fusion still stays at a low-order data level. However, high-order information of data also has superior information representation ability. Therefore, how to simultaneously utilize the high-order information of multi-modal data and the correlation information between features, and realize feature fusion processing within modes and data fusion processing between modes in turn, needs to design a multi-modal data fusion processing scheme in the power Internet of Things environment for research. SUMMARY

[0004] Therefore, the present application provides a multi-modal data fusion processing method, device, equipment and storage medium, which can focus on the consistency and difference between modes, try to capture the correlation information between modes, improve the data fusion accuracy on the basis of improving the multi-modal data fusion processing performance.

[0005] According to an aspect of the present application, an embodiment of the present application provides a multi-modal data fusion processing method, which comprises:

[0006] Obtaining a to-be-processed original multi-modal data set in a power Internet of Things;

[0007] For each modality in the to-be-processed original multi-modal data set, performing feature mapping on an original feature vector corresponding to each modality target data in each modality based on a preset feature mapping relationship, so as to convert the to-be-processed original multi-modal data set into a multi-modal enhanced data set;

[0008] Determining a relationship fusion matrix between internal features of each enhanced modality in the multi-modal enhanced data set, and fusing each enhanced modality and the relationship fusion matrix corresponding to each enhanced modality to obtain a multi-modal feature fusion data set;

[0009] Based on the initial modality features corresponding to each fusion modality in the multi-modal feature fusion data set and a pre-trained neural network model, performing feature fusion between modalities for each fusion modality to obtain the feature fusion data between modalities.

[0010] According to another aspect of the present application, the embodiments of the present application also provide a multi-modal data fusion processing device, the device comprising:

[0011] The acquisition module is configured to obtain a to-be-processed original multi-modal data set in a power Internet of Things;

[0012] The conversion module is configured to, for each modality in the to-be-processed original multi-modal data set, perform feature mapping on an original feature vector corresponding to each modality target data in each modality based on a preset feature mapping relationship, so as to convert the to-be-processed original multi-modal data set into a multi-modal enhanced data set;

[0013] The first fusion module is configured to determine a relationship fusion matrix between internal features of each enhanced modality in the multi-modal enhanced data set, and fuse each enhanced modality and the relationship fusion matrix corresponding to each enhanced modality to obtain a multi-modal feature fusion data set;

[0014] The second fusion module is configured to, based on the initial modality features corresponding to each fusion modality in the multi-modal feature fusion data set and a pre-trained neural network model, perform feature fusion between modalities for each fusion modality to obtain the feature fusion data between modalities.

[0015] According to another aspect of the present application, the embodiments of the present application also provide an electronic device, the electronic device comprising:

[0016] At least one processor; and

[0017] The memory is in communication connection with the at least one processor; wherein

[0018] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the multi-modal data fusion processing method according to any one of the embodiments of the application.

[0019] According to another aspect of the application, the embodiments of the application further provide a computer readable storage medium storing computer instructions for enabling a processor to implement the multi-modal data fusion processing method according to any one of the embodiments of the application when executed by the processor.

[0020] The technical scheme of the embodiments of the application can map the original feature vector corresponding to each modal target data in each modality in the to-be-processed original multi-modal data set by using the preset feature mapping relationship, so as to convert the to-be-processed original multi-modal data set into a multi-modal enhanced data set, can introduce the original feature vector into high dimension for feature expansion to expand the original multi-modal data representation capability; by determining the relationship fusion matrix between the internal features of each enhanced modality in the multi-modal enhanced data set, and fusing each enhanced modality and the relationship fusion matrix corresponding to each enhanced modality to obtain a multi-modal feature fusion data set, the correlation between different dimension features is considered, and the waste of useful information is effectively avoided; based on the initial modal features corresponding to each fusion modality in the multi-modal feature fusion data set and the neural network model, the feature fusion between the modalities is performed to obtain the inter-modal feature fusion data, which can focus on the consistency and difference between the intra-modal and inter-modal, try to capture the relevant information between the modalities, and improve the data fusion accuracy on the basis of improving the multi-modal data fusion processing performance.

[0021] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the application, nor is it used to limit the scope of the application. Other features of the application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the application, the following will briefly introduce the drawings needed in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor.

[0023] Figure 1 A flowchart of a multi-modal data fusion processing method provided by an embodiment of the application is shown in FIG. 1;

[0024] Figure 2 A flowchart of another multi-modal data fusion processing method provided by an embodiment of the application is shown in FIG. 2.

[0025] Figure 3 This is a structural block diagram of a multimodal data fusion processing device provided in an embodiment of the present invention;

[0026] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0027] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0028] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0029] In one embodiment, Figure 1 This is a flowchart of a multimodal data fusion processing method provided in an embodiment of the present invention. This embodiment can be applied to the case of multimodal data fusion processing in the power Internet of Things. The method can be executed by a multimodal data fusion processing device, which can be implemented in hardware and / or software and can be configured in an electronic device.

[0030] like Figure 1 As shown, the multimodal data fusion processing method in this embodiment includes the following specific steps:

[0031] S110. Obtain the raw multimodal dataset to be processed in the power Internet of Things.

[0032] The raw multimodal dataset to be processed refers to the raw multimodal data of the power system that is waiting to be fused. In this embodiment, each source or form of data information can be called a modality. The multimodal data of the power system can be understood as multimodal data related to power business, and its form of expression can include, but is not limited to, video, image, text data, etc.

[0033] In one embodiment, the original multimodal dataset to be processed includes M modalities, each corresponding to a corresponding original feature vector, and each original feature vector corresponding to an original feature label; the original feature vectors of the M modalities form an original feature space, which is represented as: Where, d m Let m be the original feature vector of the m-th modality; the original label space of the p classes corresponding to the M modalities is represented as follows: In some embodiments, the original multimodal dataset to be processed is expressed by the formula: D = {D m |1≤m≤M}, where D m Let m be the m-th mode, expressed by the formula: n is the total number of modal data in the m-th modality; Represented as modal data x i The original feature vector in the m-th modality, Let k be the original feature vector of the m-th mode, where k∈{1,2,...,d} m}, Represented as modal data x i The k-th original feature vector in the m-th mode, k∈{1,2,...,d} m}; Represented as modal data x i The corresponding original feature labels.

[0034] In this embodiment, the multimodal data of the power system can come from different fields and different stages, and have different forms and structures. Multimodal data related to the power system can be obtained from various systems such as business systems and user systems associated with the power Internet of Things (IoT) to form the original multimodal dataset to be processed. Of course, the multimodal data of the power IoT is multi-source, heterogeneous, and massive. Based on data structure, it can be divided into structured data, unstructured data, and semi-structured data. Structured data includes electrical quantities from the power system, economic data from social systems, and meteorological parameters from weather systems; unstructured data includes various images, videos, and experimental text data; and semi-structured data mainly consists of web page data. The aforementioned multimodal data involves different processing requirements, such as static and dynamic data, and originates from different times and spaces; this embodiment does not impose any limitations on these.

[0035] S120, for each modality in the to-be-processed original multi-modal data set, performing feature mapping on the original feature vector corresponding to each modality target data in the modality based on a preset feature mapping relationship, to convert the to-be-processed original multi-modal data set into a multi-modal enhanced data set.

[0036] The preset feature mapping relationship can be understood as a feature vector mapping relationship of the modality target data corresponding to each modality.

[0037] In this embodiment, the to-be-processed original multi-modal data set contains multiple modalities, each modality corresponding to multiple modality data, and each modality data corresponds to a corresponding original feature vector, which corresponds to a corresponding original label.

[0038] In this embodiment, for each modality in the to-be-processed original multi-modal data set, the original feature vector corresponding to each modality target data in each modality can be encoded accordingly to map the original feature vector corresponding to each modality target data to a high-dimensional feature space, thereby converting the original power internet of things to-be-processed original multi-modal data set into a multi-modal enhanced data set; in some embodiments, the multi-modal data can also be mapped to a specific space through a specific conversion algorithm, and then the data in different specific spaces can be mapped to the same space, thereby forming a multi-modal enhanced data set. Of course, the conversion algorithm can include but is not limited to a speech conversion algorithm, an image conversion algorithm, and a text conversion algorithm, so as to convert the speech data features through the speech conversion algorithm, convert the image data features through the image conversion algorithm, and convert the text data features through the text conversion algorithm. This embodiment does not limit this. In this embodiment, the multi-modal enhanced data set includes multiple enhanced enhanced modalities.

[0039] S130, determining a relationship fusion matrix between the internal features of each enhanced modality in the multi-modal enhanced data set, and fusing each enhanced modality and the relationship fusion matrix corresponding to each enhanced modality to obtain a multi-modal feature fusion data set.

[0040] The internal features of the modality can be understood as the features corresponding to each modality data in each enhanced modality. The relationship fusion matrix can be understood as a fusion matrix between the features corresponding to each modality data in each enhanced modality.

[0041] In the embodiment, for each enhanced modality in the multi-modal enhanced data set, a preset Pearson correlation coefficient can be used to determine the correlation degree between the internal features of each enhanced modality, and a relationship fusion matrix describing the correlation between the internal features of the corresponding modality can be determined according to the correlation degrees. On this basis, the correlation degree of each enhanced modality and its corresponding relationship fusion matrix fusion is determined, and each enhanced modality and its corresponding relationship fusion matrix are fused to obtain a corresponding feature fusion modality data subset according to the correlation degrees. In some embodiments, the n modal features corresponding to the target modality data in the enhanced modality can also be extracted, including image features, video features, speech features, and point cloud features; the correlation between the n modal features is determined; the n modal features are input into the corresponding modality network and feature-level fusion network respectively to obtain the corresponding modality decision and fusion decision, thereby forming a multi-modal feature fusion data set.

[0042] In S140, the initial modality features corresponding to each fusion modality in the multi-modal feature fusion data set and the pre-trained neural network model are used to perform feature fusion between modalities to obtain inter-modal feature fusion data.

[0043] The initial modality features can be understood as the initial features corresponding to the fusion modality, which can be extracted by a preset correlation model. In the embodiment, the preset correlation model can include but is not limited to extraction of image modality features, extraction of video modality features, and extraction of file features. Generally, the use of correlation models corresponding to different modalities is different. The preset correlation model is a related information extraction model in the prior art.

[0044] In the embodiment, the inter-modalities can be understood as the relationship between each fusion modality and the fusion modality in the multi-modal feature fusion data set. In the embodiment, the position information embedding containing time information can be obtained by the initial modality feature corresponding to each fusion modality in the multi-modal feature fusion data set and the first modality feature at the same latitude obtained by the neural network model, and the second modality feature is obtained by dimension-unified modality feature. According to the preset self-attention mechanism function, the weight distribution of the second modality sub-feature corresponding to each second modality feature in the second modality feature set is determined, and the new modality feature is obtained by using the weight distribution. On this basis, the cross-modal attention of the new modality feature corresponding modality is determined by using the preset cross-attention mechanism function, so as to obtain the inter-modal feature fusion information by performing cross-modal attention inter-modal feature fusion. In other embodiments, the probability matrix of the single modality data can also be calculated according to the pre-trained single modality multi-classification model and based on the fusion weight distribution of each single modality data, and the probability matrix of the single modality data is spliced into a fusion probability matrix to obtain the inter-modal fusion probability matrix. The embodiment is not limited in this regard.

[0045] The technical scheme of the embodiment of the application can map the original feature vector corresponding to each modality target data in each modality in the to-be-processed original multi-modal data set by a preset feature mapping relationship, so as to convert the to-be-processed original multi-modal data set into a multi-modal enhanced data set, and can introduce the original feature vector into high dimension for feature expansion to expand the representation ability of the original multi-modal data. By determining the relationship fusion matrix between the features corresponding to each enhanced modality in the multi-modal enhanced data set, and fusing each enhanced modality and the relationship fusion matrix corresponding to each enhanced modality, a multi-modal feature fusion data set is obtained, which effectively avoids the waste of useful information by considering the correlation between features of different dimensions. Based on the initial modality feature corresponding to each fusion modality in the multi-modal feature fusion data set and the neural network model, the inter-modal feature fusion of each fusion modality is performed to obtain inter-modal feature fusion data, which can focus on the consistency and difference between intra-modalities and inter-modalities, try to capture the relevant information between modalities, and improve the data fusion accuracy on the basis of improving the multi-modal data fusion processing performance.

[0046] In an embodiment, Figure 2The flowchart of another multi-modal data fusion processing method provided by an embodiment of the present application is based on the above-mentioned embodiments. In the embodiment, the original feature vectors corresponding to each modal target data in each modality are mapped based on a preset feature mapping relationship to convert the to-be-processed original multi-modal data set into a multi-modal enhanced data set. The relationship fusion matrix between the respective internal features of each enhanced modality in the multi-modal enhanced data set is determined, and each enhanced modality and the respective relationship fusion matrix corresponding to each enhanced modality are fused to obtain a multi-modal feature fusion data set. The feature fusion between modalities is further refined based on the initial modal features corresponding to each fusion modality in the multi-modal feature fusion data set and the pre-trained neural network model.

[0047] As shown in the figure, the multi-modal data fusion processing method in the embodiment can specifically include the following steps: Figure 2

[0048] S210, obtaining a to-be-processed original multi-modal data set in a power internet of things.

[0049] S220, for each modality in the to-be-processed original multi-modal data set, performing high-dimensional feature expansion on the original feature vectors corresponding to each modal target data in each modality according to a preset feature mapping relationship to obtain a high-dimensional feature space corresponding to each modality after expansion.

[0050] In the high-dimensional feature space, there are high-dimensional feature vectors corresponding to each modal target data.

[0051] In the embodiment, for each modality in the to-be-processed original multi-modal data set, the original feature vectors corresponding to each modal target data in each modality are expanded according to the preset feature mapping relationship to obtain a high-dimensional feature space corresponding to each modality after expansion. It can be understood that each original feature vector is encoded into the original feature space according to the preset feature mapping relationship, and the feature expansion is performed from the original d m dimensional input space to a new d m dimensional space to obtain a high-dimensional feature space corresponding to each modality after expansion.

[0052] In an embodiment, the preset feature mapping relationship is represented by a formula as follows: wherein, represents the high-dimensional feature space after high-dimensional feature expansion of the mth modality of the modal data x i . j∈{1,2,...,d m L},1≤k≤d m ,1≤l≤L,L defines a feature lifting rate, and L represents ​The maximum power value for high-dimensional feature expansion; since 1≤k≤d m Since 1 ≤ l ≤ L, therefore when k = d m When l = L, j = L(d m -1)+L, i.e., d m L.

[0053] In this embodiment, specifically It can be represented as:

[0054]

[0055] S230. Combine the various high-dimensional feature spaces into a multimodal augmented dataset.

[0056] In this embodiment, the high-dimensional feature space includes high-dimensional feature vectors corresponding to each modality target data, and the high-dimensional feature spaces corresponding to each modality are combined to form a multimodal augmentation dataset.

[0057] In one embodiment, the multimodal augmentation dataset includes at least two augmentation modalities, which are expressed by the formula: Among them, E m Let m be the m-th augmentation mode, and let j be the feature dimension corresponding to the m-th augmentation mode, which is expressed by the formula: L represents the feature enhancement rate, expressed as: Maximum power-law value that can be expanded; multimodal augmented dataset, expressed by the formula: E = {E} m |1≤m≤M}.

[0058] S240. For each augmented modality in the multimodal augmentation dataset, the correlation degree between the features within each augmented modality is determined using a preset Pearson correlation coefficient.

[0059] In this embodiment, for each augmentation modality in the multimodal augmentation dataset, a preset Pearson correlation coefficient can be used to determine the correlation degree between the features within each augmentation modality. In one embodiment, the preset Pearson correlation coefficient is expressed by the formula: Where a∈{1,2,...,d} m L},b∈{1,2,...,d m L}, This is represented as the dimension of the a-th augmented feature in the m-th modality. The corresponding modal data x i eigenvalues, This is represented as the b-th augmentation feature dimension in the m-th modality. The corresponding modal data x i eigenvalues, They are respectively represented as corresponding standard deviation, respectively denoted as corresponding mean

[0060] S250, according to each correlation degree, determine the relationship fusion matrix describing the association between the corresponding modal internal features.

[0061] In this embodiment, the relationship fusion matrix describing the association between the corresponding modal internal features can be determined according to the correlation degree between the corresponding modal internal features of each enhanced modal. In an embodiment, the relationship fusion matrix is represented by the formula: wherein, R m is the relationship fusion matrix corresponding to the mth modal, representing the correlation degree between the feature elements and in the mth modal; the matrix R m contains all pairs of feature relationships corresponding to the mth enhanced modal.

[0062] S260, determine the association degree of each enhanced modal and its corresponding relationship fusion matrix fusion based on the preset Taylor series fusion strategy.

[0063] In this embodiment, the association degree of each enhanced modal and its corresponding relationship fusion matrix fusion is determined according to the preset Taylor series fusion strategy. In this embodiment, the mth enhanced modal E m and the corresponding relationship fusion matrix R m fusion is represented as wherein, can be determined by the preset Taylor series fusion strategy, and in an embodiment, the preset Taylor series fusion strategy is represented by the formula: wherein, n is a positive integer and tends to infinity. By introducing infinitesimal o, the above formula can be simplified to Continue to simplify to get the formula: wherein, R (tL+n),j represents the feature element in the tL+n row and j column of the relationship fusion matrix; represents the feature value of the tL+n feature vector of the modal data x i under the mth modal; R k,j represents the feature element in the k row and j column of the relationship fusion matrix; represents the feature value of the kth feature vector of the modal data x i under the mth modal; w k ∈w, w represents a sequence with L length as the period, a total of d m periods, and a total length of d m L; R :jthe jth column of the relational fusion matrix, the Hadamard product of the relational fusion matrix.

[0064] In this embodiment, each enhanced modality and its corresponding relational fusion matrix are fused by the formula: wherein, is the modality data x i The feature vector after feature fusion.

[0065] S270, according to the correlation degree, each enhanced modality and its corresponding relational fusion matrix are fused to obtain a corresponding feature fusion modality data subset, and each feature fusion modality data subset is combined to form a multi-modal feature fusion data set.

[0066] In this embodiment, according to the correlation degree, each enhanced modality and its corresponding relational fusion matrix are fused to obtain a corresponding feature fusion modality data subset, and each feature fusion modality data subset is combined to form a multi-modal feature fusion data set. In an embodiment, the feature fusion modality data subset is represented by the formula: C m is the mth feature fusion modality data subset; the feature dimension of the mth enhanced modality in the multi-modal enhanced data set is the same as the feature dimension of the mth feature fusion modality data subset in the multi-modal feature fusion data set, which is represented by the formula: L represents the feature promotion rate, which is represented by the maximum power value that can be expanded; the multi-modal feature fusion data set is represented by the formula: C = {C m |1≤m≤M}.

[0067] S280, according to the preset correlation model, the initial modality feature corresponding to each fusion modality in the multi-modal feature fusion data set is extracted.

[0068] In this embodiment, according to the preset correlation model, the initial modality feature corresponding to each fusion modality in the multi-modal feature fusion data set is extracted. It should be noted that the preset correlation model can include but is not limited to the extraction of image modality features, the extraction of video modality features, and the extraction of file features. Generally speaking, the use of correlation models corresponding to different modalities is different, and the preset correlation model is a correlation information extraction model in the prior art.

[0069] S290, for each initial modality feature, each initial modality feature is input into a neural network model to obtain a first modality feature under the same latitude, and each first modality feature is combined to form a first modality feature set.

[0070] In this embodiment, for each initial modal feature, the neural network model is inputted with each initial modal feature to obtain a first modal feature at the same latitude, and each first modal feature is composed into a first modal feature set. Wherein, the first modal feature set is expressed by formula: Wherein, The first modal feature set with the feature dimension of each fusion modal corresponding initial modal feature unified to d is represented as X = [X1, X2,..., XM]. The first modal feature corresponding to the mth fusion modal is represented as Xm.

[0071] S2100, extract the time information and location information corresponding to each first modal feature in the first modal feature set respectively.

[0072] In this embodiment, the time information and location information corresponding to each first modal feature in the first modal feature set are extracted, which are the time information and labeled location information corresponding to the modal data respectively.

[0073] S2110, embed the time information and location information into the corresponding first modal feature to form a second modal feature, and compose the second modal feature into a second modal feature set.

[0074] In this embodiment, the time information and location information are embedded into the corresponding first modal feature to form a second modal feature, and the second modal feature is composed into a second modal feature set. Wherein, the second modal feature set is expressed by formula: X = [X1, X2,..., XM]. M ] Wherein, X m The second modal feature corresponding to the mth fusion modal is represented as Xm.

[0075] S2120, determine the weight distribution of each second modal feature in the second modal feature set according to the preset self-attention mechanism function, and weight average each second modal feature weight distribution to obtain the target modal feature corresponding to each second modal feature.

[0076] In this embodiment, the weight distribution of each second modal feature in the second modal feature set is determined according to the preset self-attention mechanism function, and each second modal feature weight distribution is weighted and averaged to obtain the target modal feature corresponding to each second modal feature.

[0077] S2130, compose each target modal feature into a new modal feature.

[0078] In the embodiment, each second modality feature corresponds to a target modality feature to form a new modality feature, and the new modality feature is expressed by a formula as follows: H = Attention (X), wherein H = [H1, H2,..., H M ] and H m represents the target modality feature of the mth modality; and Attention (·) represents a function of performing a self-attention mechanism.

[0079] In the embodiment, for the new modality feature, a preset cross-attention mechanism function is used to determine the cross-modality attention between the corresponding modalities of the new modality feature, and the cross-modality attention and the target modality feature are used to perform feature fusion between the modalities to obtain feature fusion information between the modalities.

[0080] In the embodiment, for the new modality feature, a preset cross-attention mechanism function is used to determine the cross-modality attention between the corresponding modalities of the new modality feature, and the cross-modality attention and the target modality feature are used to perform feature fusion between the modalities to obtain feature fusion information between the modalities. wherein f (·) represents a function of a cross-attention mechanism based on a feature vector, represents a feature fusion vector of the data sample i after the preset cross-attention mechanism function.

[0081] The technical scheme of the embodiment of the present application can further introduce the original feature vector into high dimension for feature expansion to expand the original multi-modal data representation capability by predefining a feature mapping relationship to perform high-dimensional feature expansion on the original feature vector corresponding to each modal target data in each modality to obtain a high-dimensional feature space corresponding to each modality after expansion, and combining the high-dimensional feature spaces to form a multi-modal enhanced data set. The correlation between the features in each enhanced modality in the multi-modal enhanced data set is determined by using a pre-defined Pearson correlation coefficient, a relationship fusion matrix describing the correlation between the features in the corresponding modality is determined according to the correlation, the correlation degree of the fusion of each enhanced modality and the relationship fusion matrix corresponding thereto is determined based on a pre-defined Taylor series fusion strategy, and each enhanced modality and the relationship fusion matrix corresponding thereto are fused according to the correlation degree to obtain a corresponding feature fusion modality data subset. The feature fusion modality data subsets are combined to form a multi-modal feature fusion data set, which can further consider the correlation between features of different dimensions and effectively avoid the waste of useful information. The initial modality features corresponding to each fusion modality in the multi-modal feature fusion data set are extracted according to a pre-defined correlation model. For each initial modality feature, the neural network model is inputted with the initial modality features to obtain first modality features at the same latitude. The time information and the position information corresponding to each first modality feature are embedded into the corresponding first modality feature to form second modality features. The weight distribution of each second modality feature corresponding to the second modality sub-feature is determined according to a pre-defined self-attention mechanism function, and the weight distribution of each second modality sub-feature is weighted and averaged to obtain a target modality feature corresponding to each second modality feature. The target modality features are combined to form new modality features. For the new modality features, a pre-defined cross-attention mechanism function is used to determine the cross-modal attention between the corresponding modalities of the new modality features, and the feature fusion information between the modalities is obtained by fusing the cross-modal attention and the target modality features. Further, the consistency and difference between the modalities can be focused on, the relevant information between the modalities can be captured as much as possible, and the data fusion accuracy can be improved on the basis of improving the multi-modal data fusion processing performance.

[0082] For example, three modalities in the to-be-processed original multi-modal data set D from the power Internet of Things, namely text, sound and image, are taken as examples for illustration.

[0083] First, three modalities in the to-be-processed original multi-modal data set D from the power Internet of Things, namely text, sound and image, are obtained. The text data set T, the sound data set V and the image data set G are generated, and multi-modal data classification fusion processing is performed.

[0084] Secondly, a feature enhancement process is performed on the multi-modal data set D, high-order information of sample features in the three modal data sets T, V and G in the power Internet of Things is modeled respectively, power of each feature value is encoded into the original space, and the original features of the three modes are mapped respectively to generate a new high-dimensional space. For the training sample x i The sample feature vector in each mode Based on the operation process of the above embodiment, the following is obtained:

[0085]

[0086] The original multi-modal data set D={D m |m∈{T,V,G}} is converted into an enhanced data set E={E m |m∈{T,V,G}} so as to achieve the purpose of expanding the data representation capability.

[0087] Then, the relationship fusion matrix between the internal features of the three different modes is generated, and the three modal enhanced data sets E={E m |m∈{T,V,G}} and the corresponding relationship fusion matrix are used to create a respective feature fusion data set C={C

[0088]

[0089] The multi-modal enhanced data set E={E m |m∈{T,V,G}} is converted into a feature fusion data set C={C m |m∈{T,V,G}} so as to avoid the waste of useful information due to the relevance between features of different dimensions.

[0090] Next, the convolutional neural network and the attention mechanism are used to perform data fusion between the text, sound and image modal data sets. The consistency and difference between the modes are focused on, and the relevant information between the current mode and other modes is captured as much as possible to improve the data fusion accuracy.

[0091] In an embodiment, Figure 3 A structural block diagram of a multi-modal data fusion processing device provided by an embodiment of the present application is provided, which is suitable for the case of multi-modal data fusion processing in the power Internet of Things, and the device can be realized by hardware / software. The device can be configured in an electronic device to realize the multi-modal data fusion processing method in the embodiment of the present application.

[0092] As shown in Figure 3 the device includes an acquisition module 310, a conversion module 320, a first fusion module 330 and a second fusion module 340;

[0093] The obtaining module 310 is configured to obtain a to-be-processed original multi-modal data set in a power internet of things.

[0094] The conversion module 320 is configured to, for each modality in the to-be-processed original multi-modal data set, perform feature mapping on an original feature vector corresponding to each modality target data in each modality based on a preset feature mapping relationship, so as to convert the to-be-processed original multi-modal data set into a multi-modal enhanced data set.

[0095] The first fusion module 330 is configured to determine a relationship fusion matrix between respective internal features of each enhanced modality in the multi-modal enhanced data set, and fuse each enhanced modality and the respective relationship fusion matrix corresponding to each enhanced modality to obtain a multi-modal feature fusion data set.

[0096] The second fusion module 340 is configured to perform inter-modality feature fusion on each fusion modality based on initial modality features respectively corresponding to each fusion modality in the multi-modal feature fusion data set and a pre-trained neural network model, to obtain inter-modality feature fusion data.

[0097] In the embodiment of the application, the conversion module performs feature mapping on an original feature vector corresponding to each modality target data in each modality in the to-be-processed original multi-modal data set based on a preset feature mapping relationship, so as to convert the to-be-processed original multi-modal data set into a multi-modal enhanced data set. The original feature vector can be introduced into high dimensions for feature expansion to expand the representation capability of the original multi-modal data set. The first fusion module determines a relationship fusion matrix between respective internal features of each enhanced modality in the multi-modal enhanced data set, and fuses each enhanced modality and the respective relationship fusion matrix corresponding to each enhanced modality to obtain a multi-modal feature fusion data set. The correlation between features of different dimensions is considered, and the waste of useful information is effectively avoided. The second fusion module performs inter-modality feature fusion on each fusion modality based on initial modality features respectively corresponding to each fusion modality in the multi-modal feature fusion data set and a neural network model, to obtain inter-modality feature fusion data. The consistency and difference between intra-modality and inter-modality are focused on, and the relevant information between modalities is captured as much as possible. On the basis of improving the multi-modal data fusion processing performance, the data fusion accuracy is improved.

[0098] In an embodiment, the to-be-processed original multi-modal data set includes M modalities, the M modalities respectively correspond to respective original feature vectors, and each original feature vector corresponds to an original feature label. The original feature vectors of the M modalities form an original feature space, which is represented as: wherein d mis the original feature vector of the mth modality; the p-class original label space corresponding to the M modalities is represented as:

[0099] The original multi-modal data set to be processed is represented by formula as D={D m |1≤m≤M} where D m is the mth modality, represented by formula as n is the total number of modal data under the mth modality; is the modal data x i is the original feature vector under the mth modality, is the kth original feature vector in the mth modality, k∈{1,2,...,d m}, is the modal data x i is the kth original feature vector under the mth modality, k∈{1,2,...,d m}; is the modal data x i corresponding original feature label.

[0100] In an embodiment, the conversion module 310 further comprises:

[0101] The expansion unit is configured to perform high-dimensional feature expansion on the original feature vector corresponding to each modality target data in each modality according to the preset feature mapping relationship, to obtain a high-dimensional feature space corresponding to each modality after expansion; wherein the high-dimensional feature space comprises high-dimensional feature vectors corresponding to each modality target data respectively.

[0102] The composition unit is configured to compose the high-dimensional feature spaces to obtain the multi-modal enhanced data set.

[0103] The multi-modal enhanced data set comprises at least two enhanced modalities, and the enhanced modality is represented by formula as: wherein E m is the mth enhanced modality, and the jth feature dimension corresponding to the mth enhanced modality is represented by formula as: L represents the feature promotion rate, represented by formula as the maximum power value that can be expanded;

[0104] The multi-modal enhanced data set is represented by formula as E={E m |1≤m≤M}.

[0105] In an embodiment, the preset feature mapping relationship is represented by formula as: wherein, is the modal data xi The m-th mode is processed through a high-dimensional feature space after high-dimensional feature expansion; j∈{1,2,...,d m L},1≤k≤d m 1≤l≤L, where L is defined as the feature enhancement rate, and L is expressed as The maximum power value for high-dimensional feature expansion; since 1≤k≤d m 1≤l≤L, therefore when k=d m When l = L, j = L(d m -1)+L, i.e., d m L.

[0106] In one embodiment, the first fusion module 330 further includes:

[0107] The correlation determination unit is used to determine the correlation between features within each augmentation mode for each augmentation mode in the multimodal augmentation dataset using a preset Pearson correlation coefficient.

[0108] The matrix determination unit is used to determine, based on the correlation degree, a fusion matrix describing the correlation between the internal features of the corresponding modality.

[0109] In one embodiment, the preset Pearson correlation coefficient is expressed by the formula: Where a∈{1,2,...,d} m L},b∈{1,2,...,d m L}, This is represented as the dimension of the a-th augmented feature in the m-th modality. The corresponding modal data x i eigenvalues, This is represented as the b-th augmentation feature dimension in the m-th modality. The corresponding modal data x i eigenvalues, They are respectively represented as The corresponding standard deviation They are respectively represented as The corresponding mean;

[0110] The relationship fusion matrix is ​​expressed by the formula: Among them, R m It is the relation fusion matrix corresponding to the m-th mode, represented by the feature elements in the m-th mode. and The degree of correlation between them; matrix R m It includes all paired feature relations corresponding to the m-th enhancement mode.

[0111] In an embodiment, the first fusion module 330 further comprises:

[0112] The association degree determination unit is configured to determine the association degree of each enhanced modality and its corresponding relational fusion matrix fusion based on the preset Taylor series fusion strategy.

[0113] The fusion data set determination unit is configured to fuse each enhanced modality and its corresponding relational fusion matrix according to the association degree to obtain a corresponding feature fusion modality data subset, and group each feature fusion modality data subset to form the multi-modal feature fusion data set.

[0114] In an embodiment, the preset Taylor series fusion strategy is expressed by a formula as follows: After simplification, the formula is obtained as follows: wherein, R (tL+n),j represents the feature element of the tL+n row and the j column in the relational fusion matrix; represents the modality data x i The eigenvalue of the tL+n feature vector under the mth modality; R k,j represents the feature element of the k row and the j column in the relational fusion matrix; represents the modality data x i The eigenvalue of the k feature vector under the mth modality; w k ∈w, w represents a sequence with L length as a period, a total of d m periods, and a total length of d m L; R :j represents the j column of the relational fusion matrix, represents the Hadamard product of the relational fusion matrix;

[0115] The fusion of each enhanced modality and its corresponding relational fusion matrix is expressed by a formula as follows: wherein, represents the modality data x i after feature fusion;

[0116] The feature fusion modality data subset is expressed by a formula as follows: C m represents the mth feature fusion modality data subset; the feature dimension of the mth enhanced modality in the multi-modal enhanced data set is the same as the feature dimension of the mth feature fusion modality data subset in the multi-modal feature fusion data set, which is expressed by a formula as follows: L represents the feature promotion rate, which is expressed as the maximum power value that can be expanded;

[0117] The multi-modal feature fusion dataset is expressed by a formula as: C={C m |1≤m≤M}.

[0118] In an embodiment, the second fusion module 340 further includes:

[0119] A first extraction unit is configured to extract initial modal features corresponding to each fusion modality in the multi-modal feature fusion dataset according to a preset correlation model;

[0120] A first determination module is configured to input each initial modal feature into the neural network model to obtain first modal features at the same latitude, and group the first modal features to form a first modal feature set, wherein the first modal feature set is expressed by a formula as: wherein, The first modal feature set with a unified feature dimension d of the initial modal features corresponding to each fusion modality is expressed as: Xm represents the first modal feature corresponding to the mth fusion modality, m∈(1,...,M);

[0121] A second extraction unit is configured to extract time information and location information corresponding to each first modal feature in the first modal feature set, respectively;

[0122] A second determination unit is configured to embed the time information and the location information into the corresponding first modal features to form second modal features, and group the second modal features to form a second modal feature set, wherein the second modal feature set is expressed by a formula as: X=[X1,X2,...,X M ], wherein X m Xm represents the second modal feature corresponding to the mth fusion modality;

[0123] A target feature determination unit is configured to determine a weight distribution of a second modal sub-feature corresponding to each second modal feature in the second modal feature set according to a preset self-attention mechanism function, and perform weighted average on the weight distribution of each second modal sub-feature to obtain a target modal feature corresponding to each second modal feature;

[0124] A grouping unit is configured to group the target modal features to form a new modal feature, wherein the new modal feature is expressed by a formula as: H=Attention(X), wherein H=[H1,H2,...,H M ], H m Hm represents the target modal feature of the mth modality; and Attention(·) represents a function of self-attention mechanism;

[0125] The fusion unit is configured to determine, for the new modality feature, cross-modal attention between the new modality feature and other modalities by using a preset cross-attention mechanism function, and perform feature fusion between the modalities according to the cross-modal attention and the target modality feature to obtain feature fusion information between the modalities. wherein f(·) represents a cross-attention mechanism-based function of a feature vector, represents a feature fusion vector of the data sample i after the preset cross-attention mechanism function.

[0126] The multi-modal data fusion processing device provided by the embodiments of the present application can perform the multi-modal data fusion processing method for a financial system provided by any of the embodiments of the present application, and has the corresponding function modules and beneficial effects of the execution method.

[0127] In an embodiment, Figure 4 A structural schematic diagram of an electronic device is provided for the embodiments of the present application. The electronic device 10 is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices (such as headsets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the present application described and / or claimed in this document.

[0128] As shown in Figure 4 The electronic device 10 includes at least one processor 11, and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11, wherein the memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0129] A plurality of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.

[0130] The processor 11 can be various general and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 performs various methods and processes described above, such as the multi-modal data fusion processing method.

[0131] In some embodiments, the multi-modal data fusion processing method can be implemented as a computer program tangibly embodied in a computer readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded onto the RAM 13 and executed by the processor 11, one or more steps of the multi-modal data fusion processing method described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to perform the multi-modal data fusion processing method by any other appropriate means, such as by means of firmware.

[0132] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (PLD), a computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0133] Computer programs used to implement the methods of the present application can be written in any combination of one or more programming languages. These computer programs can be implemented in a high level procedural or object oriented programming language, to communicate with a computer processing unit or processing units. Programs can be compiled or interpreted from readable program instructions for execution on a machine. Computer programs can be provided to a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the computer program

[0134] In the context of the present application, a computer readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. A computer readable storage medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. Alternatively, a computer readable storage medium can be a machine readable signal medium. More specific examples of a machine readable storage medium will include one or more lines of a program of instructions in a transitory signal form, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0135] To provide for interaction with a user, the systems and techniques described here can be implemented on an electronic device having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0136] The systems and techniques described herein can be implemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described herein, or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0137] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. A server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS service.

[0138] It should be understood that the various forms of flow shown above can be re-ordered, added to, or deleted from without departing from the scope of the present disclosure. For example, the steps recited in the present disclosure can be performed in parallel, in series, or in a different order, and the present disclosure is not limited in this regard.

[0139] The specific embodiments described above are not intended to be limiting, and persons skilled in the art will appreciate that various modifications, combinations, sub-combinations and alternatives can be made to the specific embodiments without departing from the spirit and principles of the present disclosure. Accordingly, the disclosure is not limited to the specific embodiments described above.

Claims

1. A multimodal data fusion processing method, characterized in that, include: Obtain the raw multimodal dataset to be processed in the power Internet of Things; For each modality in the original multimodal dataset to be processed, feature mapping is performed on the original feature vector corresponding to the target data of each modality in each modality based on a preset feature mapping relationship, so as to transform the original multimodal dataset to be processed into a multimodal augmented dataset. Determine the relationship fusion matrix between the features within each augmented modality in the multimodal augmentation dataset, and fuse each augmented modality and the relationship fusion matrix corresponding to each augmented modality to obtain a multimodal feature fusion dataset; Based on the initial modal features corresponding to each fusion modality in the multimodal feature fusion dataset and the pre-trained neural network model, feature fusion between the fusion modalities is performed to obtain the feature fusion data between the modalities; The step of determining the relationship fusion matrix between the internal features of each augmented modality in the multimodal augmentation dataset further includes: For each augmented modality in the multimodal augmentation dataset, a preset Pearson correlation coefficient is used to determine the correlation degree between the features within each augmented modality. Based on the correlation degree, a fusion matrix describing the correlation between the features within the corresponding modality is determined; The term "intermodal" refers to the interaction between different fusion modes in the multimodal feature fusion dataset.

2. The method according to claim 1, characterized in that, The original multimodal dataset to be processed includes M modalities, each corresponding to a corresponding original feature vector, and each original feature vector corresponding to an original feature label; the original feature vectors of the M modalities form an original feature space, which is represented as follows: ,in, Let be the original feature vector of the m-th modality; the original label space of the p classes corresponding to the M modalities is represented as: ; The original multimodal dataset to be processed is expressed by the formula: ,in, Let m be the m-th mode, expressed by the formula: n is the total number of modal data in the m-th modality; Represented as modal data The original feature vector in the m-th modality, This is represented as the k-th original feature vector in the m-th mode. , Represented as the modal data The k-th original feature vector in the m-th modality, ; , represented as the modal data The corresponding original feature labels.

3. The method according to claim 1, characterized in that, The step of performing feature mapping on the original feature vector corresponding to the target data of each modality in each modality based on a preset feature mapping relationship, so as to transform the original multimodal dataset to be processed into a multimodal augmented dataset, further includes: Based on the preset feature mapping relationship, the original feature vectors corresponding to the target data of each modality in each modality are extended in a high-dimensional feature space to obtain the extended high-dimensional feature space corresponding to each modality; wherein, the high-dimensional feature space includes the high-dimensional feature vectors corresponding to the target data of each modality; The high-dimensional feature spaces are combined to form the multimodal augmented dataset; The multimodal augmentation dataset includes at least two augmentation modalities, which are expressed by the following formula: ,in, Let m be the m-th augmentation mode, and let j be the feature dimension corresponding to the m-th augmentation mode, which is expressed by the formula: L represents the feature enhancement rate, expressed as The maximum power value that can be expanded; The multimodal augmented dataset is expressed by the formula: .

4. The method according to any one of claims 1 or 3, characterized in that, The preset feature mapping relationship is expressed by the formula: ,in, Represented as modal data The m-th mode is processed through a high-dimensional feature space after high-dimensional feature expansion; , , , L is defined as the feature enhancement rate, and L is expressed as The maximum power value for high-dimensional feature expansion; due to , , so when , hour, ,Right now .

5. The method according to claim 1, characterized in that, The preset Pearson correlation coefficient is expressed by the formula: ,in, , , This is represented as the dimension of the a-th augmented feature in the m-th modality. The corresponding modal data eigenvalues, This is represented as the b-th augmentation feature dimension in the m-th modality. The corresponding modal data eigenvalues, , They are respectively represented as , The corresponding standard deviation , They are respectively represented as , The corresponding mean; The relationship fusion matrix is ​​expressed by the formula: ,in, It is the relation fusion matrix corresponding to the m-th mode, represented by the feature elements in the m-th mode. and The degree of correlation between them; matrix It includes all paired feature relations corresponding to the m-th enhancement mode.

6. The method according to claim 1, characterized in that, The step of fusing each of the enhanced modalities and the relationship fusion matrix corresponding to each of the enhanced modalities to obtain a multimodal feature fusion dataset further includes: The degree of correlation between each enhanced mode and its corresponding fusion matrix is ​​determined based on a fusion strategy using a preset Taylor series. Based on the degree of correlation, each enhanced modality and its corresponding relationship fusion matrix are fused to obtain a corresponding feature fusion modality data subset, and the feature fusion modality data subsets are combined to form the multimodal feature fusion dataset.

7. The method according to claim 6, characterized in that, The fusion strategy based on the preset Taylor series is expressed by the following formula: After simplification, the formula is obtained: ;in, Represents the first in the relation fusion matrix The characteristic element of the row and the j-th column; Representing modal data In the m-th mode, the first The eigenvalues ​​of each eigenvector; This represents the feature element in the k-th row and j-th column of the relation fusion matrix; Represents the modal data The eigenvalue of the k-th eigenvector in the m-th mode; , , This represents a sequence with a period of length L, totaling... There are cycles, with a total length of . ; Let the j-th column of the relation fusion matrix be represented as " " denotes the Hadamard product of the relation fusion matrix; The fusion of each enhanced mode and its corresponding fusion matrix is ​​expressed by the following formula: ,in, Represented as modal data under the m-th mode. Feature vector after feature fusion; The feature fusion modality data subset is expressed by the following formula: , This is represented as the subset of the m-th feature fusion modality data; the feature dimension of the m-th enhanced modality in the multimodal augmentation dataset is the same as the feature dimension of the m-th feature fusion modality data subset in the multimodal feature fusion dataset, expressed by the formula: L represents the feature enhancement rate, expressed as The maximum power value that can be expanded; The multimodal feature fusion dataset is expressed by the formula: .

8. The method according to claim 1, characterized in that, The step of performing inter-modal feature fusion on each of the fusion modalities based on the initial modal features corresponding to each fusion modality in the multimodal feature fusion dataset and a pre-trained neural network model to obtain the inter-modal feature fusion data further includes: Based on a preset relevant model, the initial modal features corresponding to each fusion modality in the multimodal feature fusion dataset are extracted; For each initial modal feature, the initial modal features are input into the neural network model to obtain a first modal feature at the same dimension, and the first modal features are combined to form a first modal feature set; wherein, the first modal feature set is expressed by the formula: ,in, This represents the first modal feature set whose feature dimension is uniformly d for the initial modal features corresponding to each of the fused modalities; This is represented as the first modal feature corresponding to the m-th fusion mode. ; Extract the time information and location information corresponding to each first modality feature in the first modality feature set; The time information and the location information are embedded into the corresponding first modal features to form second modal features, and the second modal features are combined to form a second modal feature set; wherein, the second modal feature set is expressed by the formula: ,in, This is represented as the second modal feature corresponding to the m-th fusion mode; The weight distribution of the second modal sub-features corresponding to each second modal feature in the second modal feature set is determined according to the preset self-attention mechanism function, and the weight distribution of each second modal sub-feature is weighted and averaged to obtain the target modal feature corresponding to each second modal feature. The target modal features are combined to form a new modal feature; wherein the new modal feature is expressed by the formula: ,in, , Represents the target modal feature of the m-th mode; Represents a function that performs a self-attention mechanism; For the new modal feature, a preset cross-attention mechanism function is used to determine the cross-modal attention between the corresponding modalities of the new modal feature, and feature fusion between the modalities is performed based on the cross-modal attention and the target modal feature to obtain the feature fusion information between the modalities; wherein, the feature fusion information is expressed by the formula: ,in, A function representing the feature vector based on the cross-attention mechanism. This represents the feature fusion vector of data sample i after passing through the preset cross-attention mechanism function.

9. A multimodal data fusion processing device, characterized in that, include: The acquisition module is used to acquire the raw multimodal dataset to be processed in the power Internet of Things. The conversion module is used to perform feature mapping on the original feature vector corresponding to the target data of each modality in each modality in the original multimodal dataset to be processed, based on a preset feature mapping relationship, so as to convert the original multimodal dataset to be processed into a multimodal augmented dataset. The first fusion module is used to determine the relationship fusion matrix between the internal features of each augmented modality in the multimodal augmentation dataset, and to fuse each augmented modality and the relationship fusion matrix corresponding to each augmented modality to obtain a multimodal feature fusion dataset. The second fusion module is used to perform inter-modal feature fusion on each of the fusion modes based on the initial modal features corresponding to each fusion mode in the multimodal feature fusion dataset and the pre-trained neural network model to obtain the inter-modal feature fusion data; The first fusion module further includes: The correlation determination unit is used to determine the correlation between features within each augmentation mode for each augmentation mode in the multimodal augmentation dataset using a preset Pearson correlation coefficient. A matrix determination unit is used to determine, based on each of the correlation degrees, a fusion matrix describing the correlation between the internal features of the corresponding modality; The term "intermodal" refers to the interaction between different fusion modes in the multimodal feature fusion dataset.

10. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the multimodal data fusion processing method according to any one of claims 1-8.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the multimodal data fusion processing method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Cross-modal data processing method and device, storage medium and electronic device

    CN112199462A

  • Cross-modal retrieval method based on modal relation learning

    CN114817673A