Electromagnetic detection multi-mode data conversion method and device, electronic equipment and medium
By decomposing the parameterized functions into decoder and encoder, flexible conversion of electromagnetic detection of multimodal data is realized, the problem of difficulty in correlation and comprehensive analysis of multimodal data is solved, and data processing and analysis capabilities are improved.
Patent Information
- Application Number
- CN202510165574.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-06-06
AI Technical Summary
In the field of electromagnetic detection, multimodal data is difficult to directly correlate and analyze integratively, and traditional data processing methods cannot achieve flexible conversion between different mode data, which limits the in-depth understanding and effective utilization of electromagnetic detection information.
Modular processing of data conversion is realized by decomposing the parameterized functions into decoder and encoder. The source modal data is input to the encoder to obtain the joint representation data, reflecting the common characteristics between the source modal and the target modal; then the joint representation data is input to the decoder to obtain the target modal data.
It realizes flexible conversion between data in different modes, improves the ability to comprehensively process multimodal data, and provides more powerful tools to support data processing, analysis and application of electromagnetic detection.
Smart Images

Figure CN120105082A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure belongs to the field of data conversion technology, and more specifically, relates to an electromagnetic detection multi-modal data conversion method and device, electronic equipment, and medium. Background Art
[0002] In the field of electromagnetic detection, information data including language, vision, acoustics and other modalities are obtained. However, data of different modalities are often difficult to directly associate and comprehensively analyze, which greatly limits the in-depth understanding and effective use of electromagnetic detection information. Traditional data processing methods cannot achieve flexible conversion between different modal data, and have poor comprehensive processing capabilities for multimodal data, which makes them incapable of coping with complex electromagnetic detection scenarios. Summary of the invention
[0003] The purpose of the present disclosure is to provide a method and device for converting electromagnetic detection multimodal data, an electronic device, and a medium to enhance the ability to comprehensively process multimodal data.
[0004] In a first aspect of the disclosed embodiments, a method for converting electromagnetic detection multimodal data is provided, wherein a parameterized function is decomposed into a decoder and an encoder; Inputting source modality data into the encoder to obtain joint representation data, wherein the source modality data is any modality data among a plurality of modality data, and the joint representation data is common feature data representing the source modality data and the target modality data; The joint representation data is input into the decoder to obtain target modality data, where the target modality data is any modality data among the multiple modality data except the source modality data.
[0005] A second aspect of the embodiments of the present disclosure provides an electromagnetic detection multi-modal data conversion device, comprising: A decomposition module for decomposing the parameterized function into a decoder and an encoder; An encoding module, used for inputting source modality data into the encoder to obtain joint representation data, wherein the source modality data is any modality data among a plurality of modality data; A decoding module, used for inputting the joint representation data into the decoder to obtain target modality data, wherein the target modality data is any modality data among the multiple modality data except the source modality data; The joint representation data is data representing common features between the source modality data and the target modality data.
[0006] According to a third aspect of an embodiment of the present disclosure, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor implements the steps of the above-mentioned electromagnetic detection multimodal data conversion method when executing the computer program.
[0007] According to a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned electromagnetic detection multi-modal data conversion method are implemented.
[0008] The beneficial effects of the electromagnetic detection multimodal data conversion method and device, electronic device, and medium provided by the embodiments of the present disclosure are: the embodiments of the present disclosure realize modular processing of data conversion by decomposing the parameterized function into a decoder and an encoder, which is convenient for maintenance and optimization. Any modal data in the multimodal data can be input into the encoder as the source modality, and the obtained joint representation data reflects the common characteristics between the source modality and the target modality, which helps to explore the potential connection between different modalities. Finally, the joint representation data can be converted into the required target modal data through the decoder, realizing flexible conversion between different modal data, providing a more powerful tool for data processing, analysis, and application of electromagnetic detection, and improving the ability to comprehensively process multimodal data. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0010] Figure 1 A schematic diagram of a flow chart of a method for converting electromagnetic detection multi-modal data provided by an embodiment of the present disclosure; Figure 2 A structural block diagram of an electromagnetic detection multi-modal data conversion device provided in one embodiment of the present disclosure; Figure 3 A schematic block diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0011] In the following description, specific details such as specific system structures and technologies are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present disclosure. However, it should be clear to those skilled in the art that the present disclosure may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obstructing the description of the present disclosure with unnecessary details.
[0012] In order to make the purpose, technical solutions and advantages of the present disclosure more clear, specific embodiments will be described below in conjunction with the accompanying drawings.
[0013] Please refer to Figure 1 , Figure 1 A flow chart of a method for converting electromagnetic detection multimodal data provided by an embodiment of the present disclosure, the method comprising: S101: Decompose the parameterized function into a decoder and an encoder.
[0014] In this embodiment, the parameterized function refers to mapping data of different modes (such as electromagnetic data, seismic data, optical data, etc.) into a unified representation space through mathematical modeling, or realizing mutual conversion between data of different modes.
[0015] The purpose of decomposing the parameterized function into a decoder and an encoder is to modularize the complex process of multimodal data conversion. It helps to achieve data conversion from the source modality to the target modality, and facilitates precise control and optimization of the conversion process. The encoder can encode the input source modality data, and the decoder can decode the encoded features into the target modality data.
[0016] In this embodiment, the steps of decomposing the parameterized function into a decoder and an encoder are as follows: Step 1: Define the latent space: Determine the dimensionality of the latent space. The dimensionality of the latent space is lower than the input space to achieve dimensionality reduction or feature extraction.
[0017] Step 2: Design the Encoder: Design the encoder function to map the input variables to the latent space: The design of the encoder can be based on task requirements, for example: using linear transformations (such as PCA); using neural networks (such as convolutional neural networks CNN or fully connected networks); using kernel methods (such as kernel PCA).
[0018] Step 3: Design the Decoder: Design a decoder function to map the latent space back to the target space: The design of the decoder should be symmetrical or complementary to the encoder, for example: using an inverse linear transform (such as the inverse transform of PCA); using a neural network (such as a deconvolutional network or a generative adversarial network GAN); using an inverse kernel method.
[0019] Step 4: Joint Optimization: The parameters of the encoder and decoder are jointly trained by optimizing an objective function (e.g., minimizing the reconstruction error).
[0020] S102: Input source modal data into an encoder to obtain joint representation data, where the source modal data is any modal data among a plurality of modal data, and the joint representation data is common feature data representing the source modal data and the target modal data.
[0021] In this embodiment, the source modal data is any modal data selected from multiple modal data. Multiple modal data may include language modal data, visual modal data, acoustic modal data, etc. For example, when performing electromagnetic detection of underground space, the source modal data may be acoustic data obtained by detection, which contains rich acoustic information, but needs to be converted into other forms (such as vision or language) for better understanding and analysis. The source modal data has its own unique modal characteristics. In a multimodal data set, different source modal data represent different forms of information expression, and one of them can be selected as the source modality according to actual needs.
[0022] In this embodiment, when the source modality data is input into the encoder, the encoder may convert it into joint representation data. The joint representation data is the common feature data between the source modality data and the target modality data.
[0023] The encoder can map the features of the source modality data to a new feature space, so as to capture the common features between the source modality data and the target modality data. For example, if the source modality is acoustic data and the target modality is language data, the encoder can map the frequency, amplitude and other features of the acoustic data to a joint representation that can reflect both the acoustic features and the semantic features of the language data. The joint representation may be a feature vector that encodes the connection between the source modality and the target modality.
[0024] S103: Input the joint representation data into a decoder to obtain target modality data, where the target modality data is any modality data among the multiple modality data except the source modality data.
[0025] In this embodiment, the target modality data is the final result expected to be obtained through conversion. In the analysis and processing of multimodal data, data of different modalities have different application scenarios and advantages. By converting the source modality data into the target modality data, the fusion and conversion of different information can be achieved to meet different task requirements. For example, converting acoustic data into language data can facilitate the writing of text reports, and converting visual data into language data can assist in the description of image content.
[0026] In this embodiment, the decoder can remap the joint representation data to the feature space of the target modality. For example, if the joint representation data contains common features of acoustics and language, the decoder can decode it into language data based on these features, and convert the information in the joint representation into specific language expressions, such as mapping the features in the joint representation into language elements such as words, phrases or sentences.
[0027] It can be concluded from the above that this embodiment realizes modular processing of data conversion by decomposing the parameterized function into a decoder and an encoder, which is convenient for maintenance and optimization. Any modal data in the multimodal data can be input into the encoder as the source modality, and the obtained joint representation data reflects the common characteristics between the source modality and the target modality, which is helpful to explore the potential connection between different modalities. Finally, the joint representation data can be converted into the required target modal data through the decoder, realizing flexible conversion between different modal data, providing a more powerful tool for data processing, analysis and application of electromagnetic detection, and improving the ability of comprehensive processing of multimodal data.
[0028] In one embodiment of the present disclosure, it also includes: The original multimodal dataset is obtained based on the underground space detection data; The original multimodal dataset is preprocessed to obtain multiple modal data that are the same in time series.
[0029] In this embodiment, when electromagnetic detection and other methods are used to detect underground spaces, various sensors and detection equipment are used. These devices can collect different types of data. For example, acoustic sensors can collect acoustic data, cameras can obtain visual image data, and language descriptions and other information can also be collected. These different types of data constitute the original multimodal data set.
[0030] Due to the different data acquisition devices and methods of different modalities, the original data may not be synchronized in time, and the data length and format are inconsistent, making it difficult to directly perform joint representation learning and modality conversion operations. In this embodiment, the original multimodal data set can be timestamp aligned, data normalized, feature extracted, etc. to ensure that the multimodal data is consistent in the time dimension when performing data conversion, so that the encoder can be used to convert the source modality data into joint representation data, and then the decoder can be used to convert the joint representation data into the target modality data.
[0031] Through preprocessing operations, data from different modalities can be adjusted to be the same in time series. For example, for text information, the input can be aligned according to the boundaries of each word and zero-padded for each example. Similar time series synchronization operations can be performed for acoustic and visual data to have the same time length.
[0032] For example, a multimodal dataset is obtained for underground space monitoring, which consists of N labeled video clips, each of which contains a language ( )、Visual( ) and acoustics ( ) three kinds of modal information, denoted as ,in The corresponding tags for these fragments are . Synchronize multimodal data into time series data of the same length. For example, the language features of the i-th sample , Represents the lth word, L is the length of each example, and the sample also has a visual feature sequence Harmonic characteristics .
[0033] It can be concluded from the above that this embodiment obtains the original multimodal data set based on the underground space detection data, which provides a rich information source for subsequent processing. The original data set is preprocessed to make it the same in time series, which solves the problems of time asynchrony, length and format inconsistency of different modal data. It is helpful for the subsequent efficient processing of multimodal data, lays the foundation for the accuracy and reliability of electromagnetic detection multimodal data conversion, and improves the efficiency and effect of overall data processing.
[0034] In one embodiment of the present disclosure, the source modality data is input into an encoder to obtain joint representation data, including: The joint representation data is obtained based on the first formula, which is:
[0035] in, represents the joint representation data, RNN represents the recurrent network, Indicates Source modal data corresponding to time steps, Indicates The hidden state corresponding to the time step is Indicates the length of the source modality data.
[0036] In this embodiment, the source modality data It is a mode in the entire multimodal data, representing the data element of the source mode selected from the multimodal data set at the lth time step, which contains the information of this time step, such as the acoustic, visual or language information of this time step when processing underground space information.
[0037] In this embodiment, based on the recurrent network, for each time step l from 1 to L (L is the length of the source modality data), the recurrent network can be based on the hidden state of the previous time step and the source modal data of the current time step Perform calculations and gradually generate joint representation data.
[0038] Recurrent networks can use historical information to gradually integrate time series information by continuously updating hidden states to process sequence data. In this process, as time steps advance, the calculation results of each time step continue to accumulate, and the final joint representation data contains the information of the source modality data over the entire time series. The joint representation data is an abstraction and feature extraction of the source modality data. It integrates the information of different time steps to form a unified representation of the source modality data.
[0039] From the above, it can be concluded that this embodiment uses a recurrent network to process the source modal data and generate joint representation data, which lays the foundation for subsequent multimodal data conversion and analysis, makes full use of the characteristics of the recurrent network and time series information, and provides a guarantee for improving the overall performance of the multimodal data processing system.
[0040] In one embodiment of the present disclosure, the joint representation data is input into a decoder to obtain target modality data, including: The target modality data is obtained based on the probability distribution of the jointly represented data.
[0041] In one embodiment of the present disclosure, a probability distribution of the jointly represented data is obtained based on the second formula; The second formula is:
[0042] in, represents the probability distribution of the joint representation data of the target modality, represents the joint representation data of the target modality, represents the lth element of the target modality joint representation data, Indicates the length of the data represented by the union.
[0043] In this embodiment, in the process of converting the joint representation data into the target modality data, represents the lth element of the target modality joint representation data. In different modalities, can have different meanings. For text target modality, can be the lth word; for image target modality, It can be a local feature or a group of pixels in an image.
[0044] Each It is not generated independently, but can rely on some of the target modal data elements that have been generated before, which reflects the characteristics of sequence generation.
[0045] Probability distribution Reflects the data represented in a given joint and previously decoded parts of the target modal element In the case of The conditional probability of occurrence. This means that for each The generation of can estimate its most likely value based on the joint representation data and the decoded partial information.
[0046] In this embodiment, when the decoder generates the target modality data, it can be based on the joint representation data as well as The probability distribution of .
[0047] For each time step t, the decoder can consider the joint representation data The information provided is combined with the information of the target modal element that has been generated before. The probability distribution of .
[0048] From the above, it can be concluded that the present embodiment can make the generated target modal data an ordered sequence, and the generation of each element is based on the information of the previous element and the joint representation data, thereby gradually converting the joint representation data into complete target modal data. For example, in the process of converting information in the source modality of language into information in the target modality of image, the decoder can determine the next image element to be generated by calculating the probability distribution based on the information contained in the joint representation and the information of the generated partial image elements, and finally form a complete image representation. This generation method based on probability distribution makes the generation of target modal data more reasonable and accurate, while taking into account the previous and next associations and information dependencies in the generation process, ensuring the coherence and logic of data conversion.
[0049] In one embodiment of the present disclosure, it also includes: Inputting source modal data and target modal data into a target parameterization function to obtain target joint representation data; Input the target joint representation data into the recurrent neural network classifier to obtain the predicted label corresponding to the target joint representation data; The target parameterized function and the recurrent neural network classifier are parameterized based on the target joint representation data and the predicted label to obtain a parameterized function.
[0050] In this embodiment, the target parameterization function is calculated based on the input source modality data and target modality data to generate target joint representation data.
[0051] The recurrent neural network classifier receives the target joint representation data as input and outputs a predicted label. The predicted label can be a label or classification of the joint information of the current input source modality and target modality, which can be used in the subsequent evaluation and optimization process. For example, in underground space monitoring, the predicted label can represent the current state of the underground space (such as safe, dangerous, etc.) or the category of objects contained in it. The recurrent neural network classifier processes the target joint representation data, extracts the information that can be used for classification or labeling, and converts it into specific labels, which helps to further understand and analyze the data.
[0052] In this embodiment, the target joint representation data and the predicted label are used as feedback information to adjust the parameters of the target parameterization function and the recurrent neural network classifier. The difference between the predicted label and the true label is evaluated and the parameters are adjusted according to the difference. If the prediction result is inaccurate, the parameters of the target parameterization function and the recurrent neural network classifier need to be adjusted to optimize their performance.
[0053] By continuously adjusting the parameters, the target parameterized function and the recurrent neural network classifier can better generate joint representations and predict labels for the source and target modal data. In this process, optimization algorithms such as gradient descent can be used to continuously update the parameters of the target parameterized function and the recurrent neural network classifier, and ultimately obtain a parameterized function with better performance.
[0054] For example, The target parameterization function can map the source modal data and the target modal data into a joint representation space. Assume that the source modal data is X s , the target modal data is X t , the target parameterization function can be expressed as: z=fθ(X s ,X t ) Where: X s ∈R ds : Source modal data, dimension is ds.
[0055] xt∈R dt : Target modal data, dimension is dt.
[0056] fθ represents the target parameterized function, which is a neural network with parameter θ.
[0057] z∈R dz : The target joint represents the data with dimension dz.
[0058] Specific calculation: Assuming fθ is a multi-layer perceptron, its calculation process can be expressed as: h1 =σ(W 1 [X s ;X t ]+b 1 ) h 2 =σ(W 2 h1+b 2 ) … z=σ(W k h k −1+b k ) in: [xs;xt] represents the concatenation of source modality data and target modality data.
[0059] W k and b k are the weight matrix and bias vector of the kth layer respectively.
[0060] σ is the activation function (such as ReLU, Sigmoid, etc.).
[0061] After obtaining the target joint representation data z, it is input into a recurrent neural network classifier to obtain the predicted label. Assuming that the task is a classification task, the recurrent neural network classifier can be expressed as: y=gϕ(z) in: gϕ: Recurrent neural network classifier with parameter ϕ.
[0062] y∈R C : The probability distribution of the predicted label, C is the number of categories.
[0063] Specific calculation: Assuming gϕ is a recurrent neural network, its calculation process is: 1. Initialize hidden state: h0=0 where h0∈R dh is the initial hidden state and dh is the dimension of the hidden state.
[0064] 2. Recurrent neural network time step calculation: For each time step t, the calculation formula of the recurrent neural network is: h t =σ(W h h t −1+W z z+b h ) in: h t ∈R dh : represents the hidden state at time step t.
[0065] Wh∈Rdh×dh: weight matrix of hidden state.
[0066] Wz∈Rdh×dz: represents the weight matrix of the target joint representation data z.
[0067] bh∈Rdh: Bias vector representing the hidden state.
[0068] σ is the activation function (such as Tanh, ReLU, etc.).
[0069] 3. Output layer calculation: At the last time step T, the hidden state h T Input to the output layer and get the predicted label: o=W o h T +b o y=Softmax(o) in: W o ∈RC×dh: represents the weight matrix of the output layer.
[0070] b o ∈RC: represents the bias vector of the output layer.
[0071] o∈R C : Represents the raw output (logits) of the output layer.
[0072] The Softmax function converts logits into a probability distribution:
[0073] Complete Process 1. Input: source modal data xs and target modal data xt.
[0074] 2. Joint representation: Calculate the target joint representation data z through the target parameterization function fθ: 4.z=fθ(X s ,X t ) 5. Recurrent Neural Network Classifier: Input the target joint representation data z into the recurrent neural network classifier gϕ to obtain the predicted label y: 6.y=gϕ(z) 7. Output: y is the probability distribution of the predicted label.
[0075] It can be concluded from the above that this embodiment generates target joint representation data through source modality data and target modality data, and then obtains predicted labels through a recurrent neural network classifier, and then adjusts the parameters of related functions according to the predicted results and the joint representation data to optimize the function performance. Through continuous feedback and adjustment, the target parameterized function and the recurrent neural network classifier can better process multimodal data, achieve more accurate joint representation and label prediction, and finally obtain an optimized parameterized function to provide better support for the processing of multimodal data.
[0076] In one embodiment of the present disclosure, it also includes: Input the target modality data into the encoder to obtain the inverse joint representation data; Input the inverse joint representation data into the decoder to obtain the target source modality data; The parameterized function is adjusted based on the difference between the target source modal data and the source modal data.
[0077] In this embodiment, in order to ensure that the model learns a joint representation that can retain the maximum information of all modalities, a cycle consistency loss is introduced during the modality conversion. T Convert back to source modality data X again S .
[0078] In this embodiment, the target modal data X T Input encoder In this process, the encoder The target modal data X can be T Perform feature extraction and encoding operations. Encoder Based on its internal parameters and structure, the target modal data X T Convert to reverse union representation data ,Right now . Reverse union represents data It contains an abstract representation of the target modality data, which can capture the common features or related information that may exist with the source modality from the perspective of the target modality. For example, when processing multimodal data for underground space monitoring, if the source modality is acoustic data and the target modality is visual data, the visual data is input into the encoder, and the generated reverse joint representation data can reflect the features in the visual data that may be related to the acoustic data, providing an intermediate information representation for subsequent conversion.
[0079] Reverse union represents data Input decoder .Decoder The reverse joint representation data can be decoded according to its internal parameters and logic. Reverse union represents data Convert to target source modality data , the formula is The purpose is to restore the feature representation encoded from the target modality to the form of the source modality, thereby achieving the reverse conversion from the target modality data to the source modality data. The reverse conversion is to check the information integrity and accuracy of the model during the modality conversion process, such as remapping the feature information previously converted from visual data into a representation similar to acoustic data for comparison with the original source modality acoustic data.
[0080] The target source modal data obtained With the original source modality data X S Compare them and evaluate the differences between them. You can quantify the target source modality data by calculating the Euclidean distance, mean square error, or more complex similarity metrics between them. With the original source modality data X S The degree of difference between them.
[0081] Based on the evaluated differences, the parameterized function If the target source modal data and the source modal data are very different, it means that more information is lost or distorted in the process of converting from the target mode back to the source mode, which means that the parameterized function There are deficiencies in modality transfer and joint representation learning.
[0082] From the above, it can be concluded that this embodiment optimizes the multimodal data conversion process by reversely converting the target modal data to obtain the target source modal data, compares the difference with the original source modal data, and then uses the cycle consistency loss to adjust the parameterized function, so as to ensure that the model's conversion between different modes can retain information more accurately and robustly, thereby improving the performance of the entire system when processing multimodal data.
[0083] Corresponding to the electromagnetic detection multi-modal data conversion method of the above embodiment, Figure 2 This is a structural block diagram of an electromagnetic detection multi-modal data conversion device provided by an embodiment of the present disclosure. For ease of description, only the parts related to the embodiment of the present disclosure are shown. Figure 2 The electromagnetic detection multi-modal data conversion device 20 includes: a decomposition module 21, an encoding module 22, and a decoding module 23.
[0084] Wherein, the decomposition module 21 is used to decompose the parameterized function into a decoder and an encoder; The encoding module 22 is used to input the source modality data into the encoder to obtain the joint representation data, wherein the source modality data is any modality data among the multiple modality data; A decoding module 23, used for inputting the joint representation data into a decoder to obtain target modality data, where the target modality data is any modality data among the multiple modality data except the source modality data; The joint representation data is the common feature data between the source modality data and the target modality data.
[0085] In one embodiment of the present disclosure, the electromagnetic detection multi-modal data conversion device 20 further includes: a preprocessing module; the preprocessing module is specifically used for: The original multimodal dataset is obtained based on the underground space detection data; The original multimodal dataset is preprocessed to obtain multiple modal data that are the same in time series.
[0086] In one embodiment of the present disclosure, the encoding module 22 is specifically used to: The joint representation data is obtained based on the first formula, which is:
[0087] in, represents the joint representation data, RNN represents the recurrent network, Indicates Source modal data corresponding to time steps, Indicates The hidden state corresponding to the time step is Indicates the length of the source modality data.
[0088] In one embodiment of the present disclosure, the decoding module 23 is specifically used for: The target modality data is obtained based on the probability distribution of the jointly represented data.
[0089] In one embodiment of the present disclosure, the decoding module 23 is further configured to: Based on the second formula, the probability distribution of the jointly represented data is obtained; The second formula is:
[0090] in, represents the probability distribution of the joint representation data of the target modality, represents the joint representation data of the target modality, represents the lth element of the target modality joint representation data, Indicates the length of the data represented by the union.
[0091] In one embodiment of the present disclosure, the electromagnetic detection multi-modal data conversion device 20 further includes: a parameter optimization module; the parameter optimization module is specifically used to: Inputting source modal data and target modal data into a target parameterization function to obtain target joint representation data; Input the target joint representation data into the recurrent neural network classifier to obtain the predicted label corresponding to the target joint representation data; The target parameterized function and the recurrent neural network classifier are parameterized based on the target joint representation data and the predicted label to obtain a parameterized function.
[0092] In one embodiment of the present disclosure, the electromagnetic detection multi-modal data conversion device 20 further includes: a model optimization module; the model optimization module is specifically used to: Input the target modality data into the encoder to obtain the inverse joint representation data; Input the inverse joint representation data into the decoder to obtain the target source modality data; The parameterized function is adjusted based on the difference between the target source modal data and the source modal data.
[0093] See also Figure 3 , Figure 3 A schematic block diagram of an electronic device provided by an embodiment of the present disclosure. Figure 3 The electronic device 300 in the embodiment shown may include: one or more processors 301, one or more input devices 302, one or more output devices 303 and one or more memories 304. The processors 301, input devices 302, output devices 303 and memories 304 communicate with each other via a communication bus 305. The memory 304 is used to store computer programs, which include program instructions. The processor 301 is used to execute the program instructions stored in the memory 304. The processor 301 is configured to call the program instructions to execute the functions of the modules in the above-mentioned device embodiments, such as Figure 2 The functions of modules 21 to 23 are shown.
[0094] It should be understood that in the embodiment of the present disclosure, the processor 301 may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0095] The input device 302 may include a touch panel, a fingerprint collection sensor (for collecting the user's fingerprint information and fingerprint direction information), a microphone, etc., and the output device 303 may include a display (LCD, etc.), a speaker, etc.
[0096] The memory 304 may include a read-only memory and a random access memory, and provide instructions and data to the processor 301. A portion of the memory 304 may also include a non-volatile random access memory. For example, the memory 304 may also store information about the device type.
[0097] In a specific implementation, the processor 301, input device 302, and output device 303 described in the embodiments of the present disclosure can execute the implementation methods described in the first and second embodiments of the electromagnetic detection multimodal data conversion method provided in the embodiments of the present disclosure, and can also execute the implementation methods of the electronic device described in the embodiments of the present disclosure, which will not be repeated here.
[0098] In another embodiment of the present disclosure, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by the processor, all or part of the processes in the above-mentioned embodiment method are implemented, and the computer program can also be completed by instructing the relevant hardware through the computer program. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by the processor, the steps of each of the above-mentioned method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal, telecommunication signal and software distribution medium, etc.
[0099] The computer-readable storage medium may be an internal storage unit of the electronic device of any of the aforementioned embodiments, such as a hard disk or memory of the electronic device. The computer-readable storage medium may also be an external storage device of the electronic device, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (SecureDigital, SD) card, a flash card (Flash Card), etc. equipped on the electronic device. Furthermore, the computer-readable storage medium may also include both an internal storage unit of the electronic device and an external storage device. The computer-readable storage medium is used to store computer programs and other programs and data required by the electronic device. The computer-readable storage medium may also be used to temporarily store data that has been output or is to be output.
[0100] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this disclosure.
[0101] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the electronic devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0102] In the several embodiments provided in the present application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces or units, or it can be an electrical, mechanical or other form of connection.
[0103] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the embodiments of the present disclosure.
[0104] In addition, each functional unit in each embodiment of the present disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0105] The above are only specific embodiments of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present disclosure, and these modifications or replacements should be included in the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be based on the protection scope of the claims.
Claims
1. A method for converting electromagnetic detection multimodal data, characterized in that: include: Decompose the parameterized function into a decoder and an encoder; Inputting source modality data into the encoder to obtain joint representation data, wherein the source modality data is any modality data among a plurality of modality data, and the joint representation data is common feature data representing the source modality data and the target modality data; The joint representation data is input into the decoder to obtain target modality data, where the target modality data is any modality data among the multiple modality data except the source modality data.
2. The electromagnetic detection multi-modal data conversion method according to claim 1, characterized in that: Also includes: The original multimodal dataset is obtained based on the underground space detection data; The original multimodal data set is preprocessed to obtain a plurality of modal data that are identical in time series.
3. The electromagnetic detection multi-modal data conversion method according to claim 1, characterized in that: The step of inputting the source modality data into the encoder to obtain the joint representation data comprises: The joint representation data is obtained based on the first formula, wherein the first formula is: in, represents the joint representation data, RNN represents the recurrent network, Indicates Source modal data corresponding to time steps, Indicates The hidden state corresponding to the time step is Indicates the length of the source modality data.
4. The electromagnetic detection multi-modal data conversion method according to claim 1, characterized in that: The step of inputting the joint representation data into the decoder to obtain target modality data comprises: The target modality data is obtained based on the probability distribution of the jointly represented data.
5. The electromagnetic detection multi-modal data conversion method according to claim 4, characterized in that: Obtaining a probability distribution of the joint representation data based on a second formula; The second formula is: in, represents the probability distribution of the joint representation data of the target modality, represents the joint representation data of the target modality, represents the lth element of the target modality joint representation data, Indicates the length of the data represented by the union.
6. The electromagnetic detection multi-modal data conversion method according to claim 1, characterized in that: Also includes: Inputting the source modal data and the target modal data into a target parameterization function to obtain target joint representation data; Inputting the target joint representation data into a recurrent neural network classifier to obtain a prediction label corresponding to the target joint representation data; The target parameterized function and the recurrent neural network classifier are parameterized based on the target joint representation data and the predicted label to obtain the parameterized function.
7. The electromagnetic detection multi-modal data conversion method according to claim 1, characterized in that: Also includes: Inputting the target modality data into the encoder to obtain inverse joint representation data; Inputting the inverse joint representation data into the decoder to obtain target source modality data; The parameterized function is adjusted based on a difference between the target source modality data and the source modality data.
8. An electromagnetic detection multi-modal data conversion device, characterized in that: include: A decomposition module for decomposing the parameterized function into a decoder and an encoder; An encoding module, used for inputting source modality data into the encoder to obtain joint representation data, wherein the source modality data is any modality data among a plurality of modality data; A decoding module, used for inputting the joint representation data into the decoder to obtain target modality data, wherein the target modality data is any modality data among the multiple modality data except the source modality data; The joint representation data is data representing common features between the source modality data and the target modality data.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.