Multi-modal data fusion method and device based on security protection and medium

By introducing security measures such as data encryption, deep neural network fusion, hash computing and blockchain storage into multimodal data fusion technology, the security problems of traditional multimodal data fusion technology are solved, and efficient and secure fusion of data is achieved.

CN120200833APending Publication Date: 2025-06-24SHANDONG ARTAPLAY INTELLIGENT TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510548368.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

Traditional multimodal data fusion technology has security problems in data transmission and storage, and is vulnerable to malicious attacks, data leakage and tampering, resulting in user privacy and system security being threatened.

Method used

A multimodal data fusion method based on security protection is adopted, including data preprocessing and alignment, data encryption, data fusion through deep neural networks, hash computing and digital signatures, and fusion features through smart contracts and blockchain storage are achieved to achieve secure fusion and storage of data.

Benefits of technology

Effectively resist malicious attacks and tampering, reduce the risk of data leakage, enhance data security, improve system credibility and integrity verification, compatible with structural differences in multimodal data, and ensure the security of the fusion process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120200833A_ABST
    Figure CN120200833A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal data fusion method and device based on safety protection and a medium, and the method comprises the steps: carrying out the data preprocessing of each collected modal data; performing data encryption on the modal data; performing data fusion on the encrypted data through a pre-trained deep neural network; and carrying out Hash calculation on the fusion feature to obtain a corresponding digital signature, and storing the fusion feature in the constructed block chain through the smart contract and the digital signature. Through data encryption and block chain storage, hostile attacks and tampering are effectively resisted, the data leakage risk of transmission and storage links is reduced, and the data security is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing, and particularly to a multi-modal data fusion method, device, and medium based on security protection. Background Art

[0002] In recent years, multi-modal data fusion technology has developed rapidly and is widely used in fields such as intelligent transportation, medical diagnosis, and security monitoring. Its characteristic is to integrate different types of data, such as images, sounds, videos, texts, and sensor data, to improve the accuracy of decision-making through comprehensive analysis.

[0003] However, in traditional solutions, during the application of multi-modal data fusion, due to the large differences in data structures between modalities, security issues have become increasingly prominent. During data transmission and storage, it is vulnerable to malicious attacks, data leakage, and tampering, which pose a serious threat to user privacy and system security. Summary of the Invention

[0004] To solve the above problems, this application proposes a multi-modal data fusion method based on security protection, including:

[0005] For each modality data collected, perform data preprocessing and align the modality data;

[0006] Perform data encryption on the modality data to obtain encrypted data;

[0007] Through a pre-trained deep neural network, fuse the encrypted data to obtain fusion features;

[0008] Perform a hash calculation on the fusion features to obtain corresponding digital signatures, and store the fusion features in the constructed blockchain through a smart contract and the digital signatures.

[0009] On the other hand, this application also proposes a multi-modal data fusion device based on security protection, including:

[0010] At least one processor; and,

[0011] A memory communicatively connected to the at least one processor; wherein,

[0012] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the multi-modal data fusion method based on security protection as described in any of the above examples.

[0013] On the other hand, the present application also proposes a non-volatile computer storage medium storing computer-executable instructions, and the computer-executable instructions are configured as: the multi-modal data fusion method based on security protection described in any of the above examples.

[0014] The multi-modal data fusion method based on security protection proposed by the present application can bring the following beneficial effects:

[0015] 1. Through data encryption and blockchain storage, malicious attacks and tampering can be effectively resisted, the risk of data leakage in the transmission and storage links can be reduced, and data security can be enhanced.

[0016] 2. The combination of hash calculation and digital signature with smart contracts realizes the immutability and traceability of fusion features, improves the system credibility, and ensures integrity verification.

[0017] 3. The data preprocessing and alignment links alleviate the impact of multi-modal data structure differences on secure fusion. At the same time, encrypted fusion avoids the security vulnerabilities of traditional plaintext processing and can be compatible with multi-modal differences.

[0018] 4. The deep neural network directly processes encrypted data to ensure the security of the fusion process. Moreover, the blockchain smart contract automatically performs verification, reduces the risk of human intervention, and realizes automated security protection. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation of the present application. In the drawings:

[0020] Figure 1 is a schematic flowchart of the multi-modal data fusion method based on security protection in an embodiment of the present application;

[0021] Figure 2 is a schematic diagram of the multi-modal data fusion device based on security protection in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] In order to make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0023] The technical solutions provided by the embodiments of the present application will be described in detail below in conjunction with the drawings.

[0024] AsFigure 1 As shown in Figure 1 , the embodiment of the present application provides a multi-modal data fusion method based on security protection, including:

[0025] S101: For each piece of modal data collected, perform data preprocessing and align the modal data.

[0026] For different application scenarios, the modal data to be collected and the collection devices used may also be different. For example, in the field of intelligent transportation, the collection devices include cameras, sensors, etc., and can collect image modal data such as vehicle images and pedestrian images, sensor modal data such as vehicle speed, traffic flow, pedestrian flow, temperature, etc., and audio modal data such as vehicle abnormal sounds and traffic noise.

[0027] At this time, for different modal data, corresponding data preprocessing methods are respectively adopted. For example, for text modal data, data preprocessing may include word segmentation, removing meaningless words such as stop words, and converting into digital vectors through the Word2Vec model; for image modal data or video modal data, data preprocessing may include image resizing, image normalization, removing image noise, etc.; for audio modal data, image preprocessing may include removing audio noise such as background sound, audio segmentation, and converting into spectrograms; for sensor data, preprocessing may include data standardization, deleting outliers, and supplementing missing values.

[0028] Furthermore, data alignment can also be performed to facilitate subsequent processing and analysis of the data. Data alignment can include multiple aspects: time alignment, space alignment, and feature dimension alignment.

[0029] Among them, for time alignment and space alignment, perhaps not all modal data needs to be aligned, while feature dimension alignment is to align all modal data.

[0030] For time alignment, it is usually applied to video modality, image modality, audio modality, sensor modality, etc. For example, the pictures taken by the camera, the sounds recorded by the microphone, and the sensor data detected by the sensor may be out of sync in time. Therefore, at least part of the modal data is time-aligned according to the timestamps corresponding to the modal data.

[0031] For space alignment, it is usually applied to video modality, image modality, sensor modality, etc. For example, the position of the object captured by the camera and the position detected by the radar may be inconsistent. At this time, at least part of the modal data is space-aligned according to the coordinate data of the processing target (usually the monitoring target) corresponding to the modal data.

[0032] For feature dimension alignment, the feature dimensions of different modal data may vary significantly. For example, a monitoring target is described by both text modality and image modality, where the feature dimension of the image modality is 1024 and the feature dimension of the text modality is 256. In this case, for each modal data, its corresponding data features are extracted through an encoder. For example, a pre-trained ResNet network is used to extract the spatial features of the image modality, a Transformer model or a BERT model is used to extract the semantic features of the text modality, and an LSTM model is used to extract the temporal features of sensor data and audio data, etc.

[0033] Construct a shared feature space, project the data features into the shared feature space through a multi-layer perceptron, and align the feature dimensions of the modal data. During the training process of the multi-layer perceptron, a contrastive loss can also be introduced to bring relevant samples closer and push away irrelevant samples. Finally, in the shared feature space, similar modal data will be closer in distance, facilitating subsequent processing of the modal data.

[0034] S102: Encrypt the modal data to obtain encrypted data.

[0035] For different data, different encryption methods can be adopted. For example, for ordinary data, the AES (Advanced Encryption Standard) encryption algorithm can be used, while for data with strict privacy requirements (such as medical imaging data, user privacy text, etc.), only the RSA encryption algorithm can be used.

[0036] S103: Through a pre-trained deep neural network, fuse the encrypted data to obtain fused features.

[0037] The deep neural network can be set up in the server. The data acquisition device collects the modal data and transmits it to the server, and the deep neural network performs data fusion to obtain fused features.

[0038] However, in this way, it is necessary to go through the processes of data acquisition - data encryption - data transmission - data decryption - data fusion in sequence. When there are many modal data and the data volume is large, it will bring great pressure to both the acquisition device and the server.

[0039] Based on this, the deep neural network includes a first sub-model set on the edge device and a second sub-model set on the server. Both the first sub-model and the second sub-model include multiple corresponding branches, and different branches are used to process the corresponding modal data.

[0040] Among them, different branches are used to process corresponding modal data. For example, the branches can be set as Convolutional Neural Networks (CNN) branches, Long Short-Term Memory (LSTM) branches, etc. The CNN branch is used to process image modal data and video modal data to extract spatial features such as edges and textures. The LSTM branch is used to process audio modal data, text modal data, and sensor modal data to capture time series relationships.

[0041] In addition, the edge device can be the acquisition device itself, or a device that is set in the same area as the acquisition device, can be directly connected to the acquisition device, and has corresponding data processing capabilities.

[0042] At this time, through each branch in the first sub-model, the encrypted data obtained through homomorphic encryption is processed for the modal data corresponding to this branch to obtain intermediate features.

[0043] If the encrypted data is encrypted using encryption algorithms such as AES or RSA in the above text, then when the edge device processes the data, it needs to decrypt the encrypted data first and then process it, which is not only time-consuming and laborious, but also prone to the risk of data exposure.

[0044] Based on this, when encrypting modal data, homomorphic encryption can be used. Homomorphic Encryption is an encryption technology that allows specific mathematical operations to be directly performed on encrypted data, and the operation result after decryption is the same as the result of performing the same operation on the original plaintext. It can enable ciphertext data to participate in calculations securely while protecting data privacy. For example, for sensor modal data, the lightweight homomorphic encryption algorithm Paillier algorithm can be selected. For image modal data, text modal data, etc., the fully homomorphic encryption algorithms BFV algorithm and CKKS algorithm can be selected. For some modal data with relatively high security requirements, the fully homomorphic encryption algorithm TFHE algorithm can be selected.

[0045] At this time, each branch in the first sub-model processes the modal data corresponding to this branch respectively to obtain intermediate features, and the intermediate features are still in an encrypted state at this time. Compared with the original modal data, they mainly reflect the core features of the modal data and contain less data volume. Sending the intermediate features to the corresponding branches in the second sub-model naturally places less pressure on the system's transmission.

[0046] In the second sub-model, through the corresponding branch, the intermediate features are decrypted to obtain the corresponding decrypted features. When the edge device performs homomorphic encryption, it has already pre-stored the corresponding encryption method and decryption method in the server. The server can decrypt according to the intermediate features to obtain the decrypted features.

[0047] At this time, the second sub-model further processes the decrypted features, and then through the cross-modal attention module in the second sub-model, determines the first weights corresponding to each branch, and performs weighted fusion on the decrypted features through the first weights to obtain the fused features.

[0048] In this way, the preliminary processing of data in the edge device can reduce the data transmission pressure, and moreover, using the homomorphic encryption algorithm, even in edge devices with weak protection capabilities, the risk of data security leakage can be reduced.

[0049] Furthermore, after the second sub-model obtains the decrypted features, further processing is still required.

[0050] The decrypted features are normalized to scale the features of different modalities to the same range. For example, Z-Score normalization is used.

[0051] The constructed noise reduction layer is determined, and the decrypted features are denoised through the noise reduction layer. The noise reduction layer includes an encoder and a decoder. The encoder includes a fully connected layer or a convolutional layer for compressing the decrypted features and gradually reducing the feature dimension. The decoder is symmetric with the encoder and is used to reconstruct the decrypted features according to the feature compression result and gradually restore the feature dimension.

[0052] During training, the loss function can adopt the mean-square error (MSE) to minimize the reconstruction error between the input noise features and the clean features.

[0053] Through the cross-modal attention module in the second sub-model, according to the reconstructed decrypted features, the dynamic weights corresponding to each branch are determined. The reconstructed decrypted features are weighted and fused through the dynamic weights to obtain the fused features. For example, when the current multi-modal data includes an image modality and a text modality, at this time, the image features are used as the Query, and the text features are used as the Key and Value. Q, K, and V are calculated through their respective corresponding feature values and learnable weight matrices. The corresponding attention scores are calculated through Q and K and used as the dynamic weights of the text modality, and so on to obtain the dynamic weights of the image modality, and finally the final fused features are obtained by weighted summation.

[0054] In addition, during model training, the first sub-model and the second sub-model can be jointly trained.

[0055] For each branch in the first sub-model and the second sub-model, pre-train it with the corresponding training sample set to initially optimize the feature extraction ability of each branch. Collect the corresponding data sets (including image data sets, text data sets, sensor data sets, etc.), pre-train the corresponding branches, and perform encryption adaptation to convert the pre-trained weights into a format that supports homomorphic encryption. For example, convert floating-point parameters into integer polynomials.

[0056] Input encrypted data to the first sub-model. Each branch in the first sub-model independently processes the encrypted data of its corresponding modal data and outputs intermediate features through forward propagation. For example, the convolutional neural network branch processes the encrypted data corresponding to the image modality, and the long short-term memory network branch processes the encrypted data corresponding to the text modality.

[0057] Input the intermediate features to the second sub-model. Each branch in the second sub-model independently processes the intermediate features of its corresponding modal data. After decrypting and processing the intermediate features, data fusion is performed through a cross-modal attention mechanism, and fused features are output.

[0058] Calculate the losses corresponding to each branch through backpropagation. For example, both the convolutional neural network branch and the long short-term memory network can use cross-entropy loss, or they can use different loss functions based on requirements.

[0059] Update the model parameters of the first sub-model and the second sub-model according to this loss.

[0060] In addition, when the gradient is passed from the backend to the frontend, due to the use of homomorphic encryption, mathematical operations are allowed on ciphertexts, but the exact derivative calculation of non-linear activation functions is not supported (for example, the derivative of ReLU is 1 in the positive region and 0 in the negative region, with piecewise constant characteristics). Directly processing such discontinuous and non-polynomial derivatives under ciphertext will make the operation infeasible. Therefore, the impact of decryption operations on the gradient needs to be considered, and a gradient proxy needs to be set (for example, approximating the derivative of ReLU with a polynomial) to make the ciphertext operation executable.

[0061] Of course, based on requirements, simulated encrypted noise can also be added to the intermediate features to improve the robustness of the second sub-model, or adversarial samples can be generated to enhance the model's anti-attack ability.

[0062] Fine-tune and verify the overall model. Freeze the first sub-model and only train the second sub-model to optimize the decision boundary. Perform end-to-end fine-tuning, slightly adjust all parameters, and improve the overall consistency.

[0063] S104: Perform a hash calculation on the fused features to obtain the corresponding digital signature, and store the fused features in the constructed blockchain through a smart contract and the digital signature.

[0064] Hashing uses an encryption hashing function (such as SHA-256, MD5, etc.) to process the fused features, generating a hash value of a fixed length, converting the high-dimensional fused features into a compact digital fingerprint for efficiently verifying data integrity.

[0065] Digital Signature is that the data owner encrypts the hash value using their own private key to generate a unique digital signature, mainly used for identity authentication and anti-tampering proof.

[0066] A Smart Contract is essentially an automated code script running on a blockchain, which pre-defines data storage rules and can be automatically executed without manual intervention.

[0067] 1. Through data encryption and blockchain storage, it can effectively resist malicious attacks and tampering, reduce the risk of data leakage in the transmission and storage links, and enhance data security.

[0068] 2. The combination of hashing calculation, digital signature and smart contract realizes the non-tamperability and traceability of the fused features, improves the system credibility, and ensures integrity verification.

[0069] 3. The data preprocessing and alignment link alleviate the impact of multi-modal data structure differences on secure fusion. At the same time, encrypted fusion avoids the security vulnerabilities of traditional plaintext processing and can be compatible with multi-modal differences.

[0070] 4. The deep neural network directly processes encrypted data to ensure the security of the fusion process. Moreover, the blockchain smart contract automatically performs verification, reducing the risk of human intervention and realizing automated security protection.

[0071] In one embodiment, the finally obtained fused features are often used for analyzing certain targets. For example, through image recognition analysis, the sparks in the scene are analyzed and recognized, or the vehicles, pedestrians, etc. in the traffic environment are analyzed and recognized. At this time, the multiple modal data collected can include corresponding image modalities (corresponding to scene images), sound modalities (on-site sounds collected by sound sensors), temperature sensor modalities (on-site temperatures collected by temperature sensors), etc.

[0072] However, the fused features obtained at this time often contain features of multiple dimensions, and the features contained are relatively messy and often not intuitive enough for users to quickly and intuitively obtain features that match their senses.

[0073] Based on this, according to the data analysis objective of the current scenario, through a large language model, determine the human sensory modality that can intuitively perceive the data analysis objective. The human sensory modality can include the tactile modality, the auditory modality, the olfactory modality, the heartbeat modality, etc. As long as it is related to the human senses, it can be used as the corresponding modality.

[0074] For different data analysis objectives in different scenarios, the human perception is also different. For example, for spark monitoring, the human sensory modality can include the visual modality, the tactile modality, etc., which are respectively used to observe the size of the spark and feel the on-site temperature. For traffic flow monitoring, the human sensory modality can include the visual modality, the auditory modality, etc., which are respectively used to observe the on-site traffic flow and feel the on-site noise, etc.

[0075] At this time, based on the human sensory modality, select the corresponding sensor modality data in the modality data as the specified modality data. Generally speaking, for the visual modality, the image modality data can be used as the specified modality data. However, for most data analysis processes, the image modality data is the commonly used modality data, and it usually occupies a relatively high weight in the fusion features. Therefore, it is not used as the specified modality data, and the specified modality data is only selected from the sensor modality data with a lower weight. For example, for the tactile modality, the temperature sensor modality data can be selected as the specified modality data, and for the auditory modality, the sound sensor modality data can be selected as the specified modality data.

[0076] At this time, based on the specified modality data, establish a sensory association with other modality data, and based on the sensory association, perform sensory correction on the specified modality data, so that in the finally obtained fusion features, the characteristic dimension of the modality data (that is, the specified modality data) that can directly reflect the human intuitive feeling can be further highlighted.

[0077] Furthermore, when establishing the sensory association, determine the main modality data corresponding to the current scenario in the other modality data. The other modality data refers to all the modality data in the input data except the specified modality data.

[0078] The main modality data can be preset or obtained through a large language model, and it is mainly used to reflect which data can best reflect the current on-site situation among all the modality data. For example, in spark monitoring and traffic flow monitoring, the image modality data is the main modality data, while in medical data monitoring, the feedback data of some medical devices is the main modality data. Only selecting the main modality data can reduce the corresponding computational consumption while ensuring the stability of the overall direction.

[0079] If multiple specified modal data are given, the large language model is used to determine the second weights corresponding to each specified modal data for the data analysis objective in the current scenario. For example, the large model prompt can be: the current scenario is A, the data analysis objective is B, and there are multiple modal data C and D here. According to the importance of each modal data for the data analysis objective in the analysis process, the weights of each modal data are given.

[0080] At this time, the sensory association between the specified modal data and the main modal data can be established according to the second weights. The sensory association includes the second weights corresponding to the main modal data and each specified modal data.

[0081] At this time, when performing sensory correction, all fusion features can be corrected, or only some fusion features can be selected for correction. For example, when it is determined that in the data analysis result corresponding to the fusion feature, the data analysis objective meets the preset conditions (for example, in spark monitoring, a spark has been detected, or in medical imaging, it is detected that the patient's signs meet some preset abnormal states), it is considered that the obtained fusion features need to be corrected again to obtain new corrected fusion features for subsequent data analysis.

[0082] At this time, through the sensory selection layer in the second sub-model, according to the second weights, the relevant channels corresponding to the specified modal data in the decrypted features are amplified, and the relevant channels are phase-synchronized with the main modal data. The amplification process can directly amplify the original feature channels using an amplification factor, and the value of the amplification factor can be set manually and can be dynamically adjusted during the model training process.

[0083] Phase synchronization is mainly used to ensure the physical synchronization of sensory signals of different modalities and prevent physical desynchronization caused by amplifying the processing channels. The method of phase synchronization can be to perform a Fourier transform on the decrypted features of the main modal data, extract the corresponding dominant frequency, and then adjust the vibration fundamental frequency of the feature generator of the specified modal data to this dominant frequency to complete the phase synchronization.

[0084] At this time, in the decrypted features, for the main modal data, its own weight is relatively high, so no adjustment is made; for other unimportant modal data, its own importance is low, so no adjustment is needed either; for some specified modal data that are relatively important and can intuitively reflect human feelings sensorially, but may not have a high weight themselves, their weights in the fusion features can be increased through sensory correction, so that the fusion features can more intuitively reflect human feelings. At this time, weighted fusion is performed on the decrypted features after sensory correction to obtain the fusion features.

[0085] Such as Figure 2As shown in the figure, the embodiment of the present application also proposes a multi-modal data fusion device based on security protection, including:

[0086] At least one processor; and,

[0087] A memory communicatively connected to the at least one processor; wherein,

[0088] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the multi-modal data fusion method based on security protection as described in any of the above embodiments.

[0089] The embodiment of the present application also proposes a non-volatile computer storage medium storing computer-executable instructions, and the computer-executable instructions are set to be the multi-modal data fusion method based on security protection as described in any of the above embodiments.

[0090] Each embodiment in the present application is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device and medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiments.

[0091] The device and medium provided by the embodiment of the present application correspond one by one to the method. Therefore, the device and medium also have beneficial technical effects similar to those of the corresponding method. Since the beneficial technical effects of the method have been described in detail above, the beneficial technical effects of the device and medium will not be elaborated here.

[0092] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.

[0093] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0094] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0095] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0096] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.

[0097] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. The memory is an example of computer-readable media.

[0098] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0099] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0100] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.

Claims

1. A multimodal data fusion method based on security protection, characterized in that: include: For each modal data collected, data preprocessing is performed, and the modal data is aligned; Performing data encryption on the modal data to obtain encrypted data; The encrypted data is fused by a pre-trained deep neural network to obtain fusion features; A hash calculation is performed on the fusion feature to obtain a corresponding digital signature, and the fusion feature is stored in the constructed blockchain through a smart contract and the digital signature.

2. The method according to claim 1, characterized in that: The modality data is aligned, specifically comprising: Time-aligning at least a portion of the modal data according to timestamps corresponding to the modal data; spatially aligning at least a portion of the modal data according to coordinate data of a processing target corresponding to the modal data; For each modal data, the corresponding data features are extracted through the encoder; A shared feature space is constructed, the data features are projected into the shared feature space through a multi-layer perceptron, and the feature dimensions of the modal data are aligned.

3. The method according to claim 1, characterized in that: The encrypted data is fused through a pre-trained deep neural network to obtain fusion features, including: Determining that the deep neural network includes a first sub-model set on an edge device and a second sub-model set on a server, and determining that the first sub-model and the second sub-model each include a plurality of corresponding branches, and different branches are used to process corresponding modal data; Through each branch in the first sub-model, for the modal data corresponding to the branch, the encrypted data obtained by homomorphic encryption is processed to obtain intermediate features; Sending the intermediate feature to a corresponding branch in the second sub-model, and decrypting the intermediate feature through the corresponding branch to obtain a corresponding decrypted feature; The first weight corresponding to each branch is determined through the cross-modal attention module in the second sub-model, and the decryption features are weightedly fused through the first weight to obtain a fused feature.

4. The method according to claim 3, characterized in that Determine the first weight corresponding to each branch through the cross-modal attention module in the second sub-model, and perform weighted fusion on the decryption features through the first weight to obtain fusion features, specifically including: Performing standardization on the decryption features; Determine a constructed denoising layer, and perform denoising on the decryption feature through the denoising layer; the denoising layer includes an encoder and a decoder, the encoder is used to perform feature compression on the decryption feature, and the decoder is used to reconstruct the decryption feature according to the feature compression result; Determine the dynamic weights corresponding to each branch according to the reconstructed decrypted features through the cross-modal attention module in the second sub-model; The reconstructed decryption features are weighted and fused using the dynamic weights to obtain fused features.

5. The method according to claim 3, characterized in that: The training process of the deep neural network includes: Jointly training the first sub-model and the second sub-model; For each branch in the first sub-model and the second sub-model, pre-training is performed using a corresponding training sample set; Inputting encrypted data to the first sub-model, independently processing the encrypted data of the modal data corresponding to itself through each branch in the first sub-model, and outputting intermediate features through forward propagation; Inputting intermediate features to the second sub-model, each branch in the second sub-model independently processes the intermediate features of the modal data corresponding to itself, decrypts and processes the intermediate features, performs data fusion through a cross-modal attention mechanism, and outputs fused features; The losses corresponding to each branch are calculated through back propagation, and the model parameters of the first sub-model and the second sub-model are updated according to the losses; wherein the first sub-model is updated by setting a gradient agent.

6. The method according to claim 3, characterized in that The method further comprises: Based on the data analysis target of the current scene, determine the human sensory modality that can intuitively perceive the data analysis target through a large language model; Based on the human sensory modality, selecting corresponding sensor modality data from the modality data as designated modality data; Based on the specified modal data, a sensory association is established with other modal data, and based on the sensory association, a sensory correction is performed on the specified modal data.

7. The method according to claim 6, characterized in that Based on the specified modal data, a sensory association is established with other modal data, specifically including: Determine, among other modal data, main modal data corresponding to the current scene; If there are multiple specified modal data, determining, by means of a large language model, a second weight corresponding to each specified modal data of the data analysis target in the current scene; A sensory association between the designated modality data and the main modality data is established according to the second weight.

8. The method according to claim 7, characterized in that Based on the sensory association, sensory correction is performed on the specified modal data, specifically including: Determining that in the data analysis result corresponding to the fusion feature, the data analysis target meets the preset situation; Through the sensory selection layer in the second sub-model, according to the second weight, amplifying the relevant channel corresponding to the specified modal data in the decrypted feature, and performing phase synchronization between the relevant channel and the main modal data; The decrypted features after sensory correction are weighted fused to obtain the fused features.

9. A multimodal data fusion device based on security protection, characterized in that: include: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the multimodal data fusion method based on security protection as described in any one of claims 1 to 8.

10. A non-volatile computer storage medium storing computer executable instructions, characterized in that: The computer executable instructions are configured as: the multimodal data fusion method based on security protection as described in any one of claims 1 to 8.

Citation Information

Cited By

  • Financial bill auditing and decision-making method and system, terminal and medium

    CN120975945A

  • Multi-modal feature signature generation method and device for intelligent agent and medium

    CN121387232A

  • Traffic data trusted sharing method, system and device and storage medium

    CN121414780A