A data processing method and system
Patent Information
- Application Number
- CN202210784636.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-05
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2042-07-05
AI Technical Summary
[0004]但是,现有技术无法对多模态数据进行有效处理,导致无法有效保障任务处理的准确度
[0042] The data processing method and system proposed in this invention include: a feature fusion network and a feature extraction network group. The feature extraction network group includes at least two feature extraction networks, each corresponding to different single-modal data. The feature fusion network includes at least one feature fusion layer, and at least one feature extraction network includes at least two feature extraction layers. This invention can obtain multiple matching preprocessed modal data. Each preprocessed modal data is input to the first feature extraction layer of each feature extraction network. Before the last feature extraction layer in each feature extraction network, the modal feature data extracted by any feature extraction layer is input to the next feature extraction layer in the same feature extraction network, and also to the feature fusion layer in the same layer. The feature fusion data generated by the fusion processing of the received different modal feature data by any feature fusion layer is input to the next feature extraction layer in each feature extraction network. The modal feature data extracted by the last feature extraction layer in each feature extraction network is input to the target task network to obtain the task result output by the target task network. The modal data input into the target task network by this invention can largely retain the features of the original single modality, and can also add additional feature information of other extracted modal data, thus effectively improving the task processing accuracy of the target task network.
Smart Images

Figure CN115204282B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a data processing method and system. Background Technology
[0002] With the development of computer science and technology, the application of artificial intelligence technology is constantly improving.
[0003] Currently, to effectively ensure the accuracy of tasks such as detection, classification, and sequence output of neural networks or models, most artificial intelligence applications focus on the integrated processing of multimodal data. Multimodal data refers to multiple modalities of data used as input to a neural network or model to process the same task. For example, fingerprint data, voice data, and facial data are used as input to a classification neural network for person recognition.
[0004] However, existing technologies cannot effectively process multimodal data, which makes it impossible to effectively guarantee the accuracy of task processing. Summary of the Invention
[0005] In view of the above problems, the present invention provides a data processing method and system that overcomes or at least partially solves the above problems, the technical solution of which is as follows:
[0006] A data processing method is applied to a data processing system, the data processing system comprising: a feature fusion network and a feature extraction network group, wherein the feature extraction network group includes at least two feature extraction networks, each feature extraction network corresponding to different single-modal data; wherein the feature fusion network includes at least one feature fusion layer, and at least one of the feature extraction networks includes at least two feature extraction layers; the data processing method includes:
[0007] Obtain multiple matching preprocessed modal data;
[0008] Each of the preprocessed modal data is respectively input into the first feature extraction layer of each of the feature extraction networks;
[0009] Before the last feature extraction layer in each of the aforementioned feature extraction networks, the modal feature data extracted by any feature extraction layer is input to the next feature extraction layer in the same feature extraction network, and also input to the feature fusion layer in the same layer.
[0010] The feature fusion data generated by fusing the received feature data of different modalities by any feature fusion layer is input into the feature extraction layer of the next layer in each feature extraction network.
[0011] The modal feature data extracted from the last feature extraction layer in each of the aforementioned feature extraction networks are input into the target task network to obtain the task result output by the target task network.
[0012] Optionally, each of the aforementioned feature extraction networks has the same number of feature extraction layers, and each has at least two layers.
[0013] Optionally, the feature fusion network has N layers, and each feature extraction network has M layers, where M is the value 1 obtained by subtracting N from M.
[0014] Optionally, the first feature extraction layer in each of the feature extraction networks extracts modal feature data from the corresponding preprocessed modal data.
[0015] Optionally, each feature extraction layer in each feature extraction network after the first layer, after obtaining the modal feature data input from the previous feature extraction layer in the same feature extraction network and the feature fusion data input from the previous feature fusion layer, concatenates the input modal feature data and feature fusion data to obtain concatenated data, and extracts modal feature data from the concatenated data.
[0016] Optionally, the feature extraction network group includes a first feature extraction network and a second feature extraction network, and the feature fusion network is a network based on a multi-head attention mechanism.
[0017] Optionally, before the last feature extraction layer, each feature extraction layer performs a block-based operation on the modal feature data after extracting the modal feature data, obtaining block-based feature data, and inputting the block-based feature data into the feature fusion layer of the same layer; wherein, the block-based feature data includes multiple block-based feature data.
[0018] Optionally, after obtaining the segmented feature data input from the feature extraction layers of the two feature extraction networks, any feature fusion layer, based on the segmented feature data input from the feature extraction layer of the first feature extraction network and paying attention to the segmented feature data input from the feature extraction layer of the second feature extraction network, obtains first attention data and sends the first attention data to the next feature extraction layer in the first feature extraction network; based on the segmented feature data input from the feature extraction layer of the second feature extraction network and paying attention to the segmented feature data input from the feature extraction layer of the first feature extraction network, obtains second attention data and sends the second attention data to the next feature extraction layer in the second feature extraction network.
[0019] Optionally, the target task network includes a pooling layer and two fully connected layers.
[0020] Optionally, obtaining multiple matching preprocessed modal data includes:
[0021] Obtain multiple single-modal data output by the multimodal sensor;
[0022] According to the predefined preprocessing method, each obtained single-modal data is preprocessed to obtain the matching preprocessed modal data.
[0023] A data processing system includes: a feature fusion network and a feature extraction network group, wherein the feature extraction network group includes at least two feature extraction networks, each feature extraction network corresponding to different single-modal data; wherein the feature fusion network includes at least one feature fusion layer, and at least one feature extraction network includes at least two feature extraction layers; the data processing system further includes: a first acquisition unit, a first input unit, a second input unit, a third input unit, a fourth input unit, and a second acquisition unit; wherein:
[0024] The first obtaining unit is used to obtain multiple matching preprocessed modal data;
[0025] The first input unit is used to input each of the preprocessed modal data into the first feature extraction layer of each feature extraction network respectively;
[0026] The second input unit is used to input the modal feature data extracted by any feature extraction layer into the next feature extraction layer in the same feature extraction network before the last feature extraction layer in each of the feature extraction networks, and to input it into the feature fusion layer of the same layer.
[0027] The third input unit is used to input the feature fusion data generated by fusing the received different modal feature data by any feature fusion layer into the next feature extraction layer in each of the feature extraction networks.
[0028] The fourth input unit is used to input the modal feature data extracted by the last feature extraction layer in each feature extraction network into the target task network;
[0029] The second obtaining unit is used to obtain the task results output by the target task network.
[0030] Optionally, each of the aforementioned feature extraction networks has the same number of feature extraction layers, and each has at least two layers.
[0031] Optionally, the feature fusion network has N layers, and each feature extraction network has M layers, where M is the value 1 obtained by subtracting N from M.
[0032] Optionally, the first feature extraction layer in each of the feature extraction networks extracts modal feature data from the corresponding preprocessed modal data.
[0033] Optionally, each feature extraction layer in each feature extraction network after the first layer, after obtaining the modal feature data input from the previous feature extraction layer in the same feature extraction network and the feature fusion data input from the previous feature fusion layer, concatenates the input modal feature data and feature fusion data to obtain concatenated data, and extracts modal feature data from the concatenated data.
[0034] Optionally, the feature extraction network group includes a first feature extraction network and a second feature extraction network, and the feature fusion network is a network based on a multi-head attention mechanism.
[0035] Optionally, before the last feature extraction layer, each feature extraction layer performs a block-based operation on the modal feature data after extracting the modal feature data, obtaining block-based feature data, and inputting the block-based feature data into the feature fusion layer of the same layer; wherein, the block-based feature data includes multiple block-based feature data.
[0036] Optionally, after obtaining the segmented feature data input from the feature extraction layers of the two feature extraction networks, any feature fusion layer, based on the segmented feature data input from the feature extraction layer of the first feature extraction network and paying attention to the segmented feature data input from the feature extraction layer of the second feature extraction network, obtains first attention data and sends the first attention data to the next feature extraction layer in the first feature extraction network; based on the segmented feature data input from the feature extraction layer of the second feature extraction network and paying attention to the segmented feature data input from the feature extraction layer of the first feature extraction network, obtains second attention data and sends the second attention data to the next feature extraction layer in the second feature extraction network.
[0037] Optionally, the target task network includes a pooling layer and two fully connected layers.
[0038] Optionally, the first obtaining unit includes: a third obtaining unit, a preprocessing unit, and a fourth obtaining unit;
[0039] The third obtaining unit is used to obtain multiple single-mode data output by the multimodal sensor;
[0040] The preprocessing unit is used to preprocess each obtained single-modal data according to a predefined preprocessing method.
[0041] The fourth obtaining unit is used to obtain the matching preprocessed modal data.
[0042] The data processing method and system proposed in this invention include: a feature fusion network and a feature extraction network group. The feature extraction network group includes at least two feature extraction networks, each corresponding to different single-modal data. The feature fusion network includes at least one feature fusion layer, and at least one feature extraction network includes at least two feature extraction layers. This invention can obtain multiple matching preprocessed modal data. Each preprocessed modal data is input to the first feature extraction layer of each feature extraction network. Before the last feature extraction layer in each feature extraction network, the modal feature data extracted by any feature extraction layer is input to the next feature extraction layer in the same feature extraction network, and also to the feature fusion layer in the same layer. The feature fusion data generated by the fusion processing of the received different modal feature data by any feature fusion layer is input to the next feature extraction layer in each feature extraction network. The modal feature data extracted by the last feature extraction layer in each feature extraction network is input to the target task network to obtain the task result output by the target task network. The modal data input into the target task network by this invention can largely retain the features of the original single modality, and can also add additional feature information of other extracted modal data, thus effectively improving the task processing accuracy of the target task network.
[0043] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention, it can be implemented according to the contents of the specification. In order to make the above and other objects, features and advantages of the present invention more obvious and understandable, specific embodiments of the present invention are described below. Attached Figure Description
[0044] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0045] Figure 1 A flowchart of the first data processing method provided by an embodiment of the present invention is shown;
[0046] Figure 2 A schematic diagram of the structure of a data processing system provided in an embodiment of the present invention is shown;
[0047] Figure 3 A schematic diagram of another data processing system provided by an embodiment of the present invention is shown;
[0048] Figure 4 A schematic diagram of another data processing system provided by an embodiment of the present invention is shown;
[0049] Figure 5 A schematic diagram of the structure of a data processing system provided in an embodiment of the present invention is shown. Detailed Implementation
[0050] Exemplary embodiments of the invention will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the invention are shown in the drawings, it should be understood that the invention may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this invention will be thorough and complete, and will fully convey the scope of the invention to those skilled in the art.
[0051] like Figure 1 As shown, this embodiment proposes a first data processing method, which can be applied to a data processing system. The data processing system includes a feature fusion network and a feature extraction network group. The feature extraction network group includes at least two feature extraction networks, each corresponding to different single-modal data. The feature fusion network includes at least one feature fusion layer, and at least one feature extraction network includes at least two feature extraction layers. The method may include the following steps:
[0052] S101. Obtain multiple preprocessed modal data that match the given data;
[0053] Among them, the feature fusion network is a neural network that can fuse different data according to a predefined data fusion processing method. Specifically, the feature fusion layer can send the data obtained after fusion processing to the next feature extraction layer in different feature extraction networks.
[0054] In this context, a feature extraction network is a neural network that can extract data from input data according to a predefined data extraction method. Specifically, any feature extraction layer can input the extracted data into the next feature extraction layer in the same feature extraction network, as well as into a feature fusion layer of the same number of layers.
[0055] Optionally, the feature extraction network can be a convolutional neural network or a Transformer network. Specifically, to effectively reduce the computational load during data fusion and ensure the convenience and effectiveness of data fusion, the different feature extraction networks in this invention can be of the same type. It should be noted that both the feature fusion network and the feature extraction network in this embodiment can be trained and pre-trained neural networks.
[0056] Optionally, the present invention can annotate a set of multimodal data according to task requirements. For example, target annotation outside the vehicle or lip reading annotation inside the vehicle;
[0057] Optionally, this invention can design and train a feature extraction network structure for feature extraction of a specific unimodal data. In this case, the invention can perform some transformation processing on the modal data according to the input format requirements of different feature extraction networks. For example, when processing RGB image modal data, a convolutional neural network may need to pre-size the image to a specified input, while a trAnsformer network may need to divide the image into 16 non-overlapping small image blocks. In this case, the invention can use the corresponding feature layers in different unimodal models for the next fusion step, requiring the feature output scale to be the same. Optionally, the trained unimodal model weights can be used as pre-trained weights in the multimodal training stage.
[0058] One of the preprocessed modal data can be the data obtained after preprocessing the single modal data.
[0059] Optionally, in other data processing methods proposed in this embodiment, step S101 may include:
[0060] Obtain multiple single-modal data output by the multimodal sensor;
[0061] According to the predefined preprocessing method, each obtained single-modal data is preprocessed to obtain matching preprocessed modal data.
[0062] Among them, single-modal data can be one type of data, such as fingerprint data, voice data, and facial data used for person recognition.
[0063] It should be noted that in this embodiment, multiple single-modal data can be corresponding. For example, in the cockpit of an intelligent car, video data and voice data of the same driver within the same time period; or, during the driving process of an intelligent car, camera images and LiDAR point cloud data of the same external scene.
[0064] Specifically, this invention can acquire multiple single-modal data through multiple modal sensors. For example, in applications where intelligent vehicles need to acquire multimodal data for person recognition, this invention can acquire video data through a camera and audio data through a microphone. The camera and microphone belong to different modal sensors, while the video data and audio data belong to different single-modal data. It should be noted that single-modal data can also be point cloud data from LiDAR; this invention does not limit the specific data types of modal sensors and single-modal data.
[0065] Specifically, this invention can utilize preprocessing methods, including time alignment and slicing, to ensure that the preprocessed modal data is suitable as input data for corresponding feature extraction networks. It also ensures that the preprocessed modal data are matched, guaranteeing that the subsequent set of data fed into the neural network represents different modal descriptions of the target scene within the same time and space—that is, temporal and spatial overlap. For example, for video and audio data obtained within the same time period inside a smart car, this invention can first convert the audio data into a spectrogram, transforming a continuous audio waveform into a single frame spectrogram, and then uniformly extract frames from the video data to obtain discrete frame image data.
[0066] Specifically, the present invention can determine the preprocessing method based on specific modal characteristics and the structure of the feature extraction layer in the feature extraction network. The present invention does not limit the specific processing method of the preprocessing method.
[0067] It should be noted that the present invention can utilize multiple feature extraction layers in the feature extraction network to perform multiple rounds of feature extraction on multimodal data, and perform multiple data interactions between the feature extraction network and the feature fusion network, so that the feature fusion network can perform effective data fusion, thereby enabling the task network to improve the accuracy of task processing based on the effectively fused data.
[0068] Optionally, each feature extraction network has the same number of feature extraction layers, and each has at least two layers.
[0069] Optionally, the number of feature fusion layers in the feature fusion network is N, and the number of feature extraction layers in each feature extraction network is M, where M minus N equals 1.
[0070] To better illustrate the specific processing procedure of the data processing method and the composition structure of the data processing system of the present invention, this embodiment proposes and combines the following for the application scenario of human recognition in intelligent vehicles. Figure 2 The data processing system shown is explained.
[0071] Specifically, in the application scenario of person recognition in intelligent vehicles, multimodal data can include two unimodal data, namely video data and audio data. In this case, the data processing system designed for this application scenario according to the present invention can be as follows: Figure 2As shown, the data processing system may include a feature fusion network and a feature extraction network group. The feature extraction network group includes a feature extraction network C1 for extracting feature data from video data A and a feature extraction network C2 for extracting feature data from audio data B. C1 and C2 are of the same network type and each includes M feature extraction layers. The feature fusion network includes N feature fusion layers, where M is one greater than N. Within the same feature extraction network, upper and lower feature extraction networks can interact with each other; each feature fusion layer can interact with feature extraction layers of the same number.
[0072] S102. Input each preprocessed modal data into the first feature extraction layer of each feature extraction network respectively;
[0073] Specifically, after obtaining the preprocessed modal data, this invention can input each preprocessed modal data into the first feature extraction layer of the corresponding feature extraction network. For example... Figure 2 As shown, in Figure 2 In the data processing system shown, after obtaining video data A and audio data B, the present invention can input A into the first feature extraction layer in the corresponding C1 and input B into the first feature extraction layer in the corresponding C2.
[0074] S103. Before the last feature extraction layer in each feature extraction network, input the modal feature data extracted by any feature extraction layer into the next feature extraction layer in the same feature extraction network, and input it into the feature fusion layer in the same layer.
[0075] S104. The feature fusion data generated by fusing the received feature data of different modalities by any feature fusion layer is input into the feature extraction layer of the next layer in each feature extraction network.
[0076] S105. Input the modal feature data extracted from the last feature extraction layer in each feature extraction network into the target task network.
[0077] Optionally, the first feature extraction layer in each feature extraction network extracts modal feature data from the corresponding preprocessed modal data.
[0078] Specifically, in the same feature extraction network, the first feature extraction layer can extract the corresponding feature data from the single-modal data input to the feature extraction network; then, the first feature extraction layer can output the extracted feature data to the second feature extraction layer in the same feature extraction network, as well as the feature fusion layer, which is also the first layer in the feature fusion network.
[0079] Understandably, the first-layer feature fusion layer can receive feature data output from the first-layer feature extraction layers of each feature extraction network. Then, the first-layer feature fusion layer can fuse the received feature data output from the first-layer feature extraction layers of each feature extraction network to generate corresponding fused feature data, which is then input into the second-layer feature extraction layers of each feature extraction network.
[0080] Specifically, the second feature extraction layer can receive the feature data output by the first feature extraction layer in the same feature extraction network, as well as the feature fusion data output by the first feature fusion layer. Then, the second feature extraction layer can concatenate the received feature data and the feature fusion data to obtain concatenated data. After that, it can extract the corresponding feature data from the concatenated data and output them to the third feature extraction network in the same feature extraction network, as well as the feature fusion layer, which is also the second layer in the feature fusion network.
[0081] It is understandable that the second layer and subsequent feature fusion layers, like the first layer, can receive feature data output from the same layer of the feature extraction network, perform fusion processing on the received feature data to obtain feature fusion data, and then send it to the next feature extraction layer in each feature extraction network; for example... Figure 2 As shown, the second feature fusion layer can receive feature data output from the feature extraction layers of C1 and C2, which are also the second layer. It performs fusion processing on the received feature data to obtain feature fusion data, and sends it to the feature extraction layers of C1 and C2 respectively.
[0082] Optionally, each feature extraction layer after the first layer in each feature extraction network, after obtaining the modal feature data input from the previous feature extraction layer in the same feature extraction network and the feature fusion data input from the previous feature fusion layer, concatenates the input modal feature data and feature fusion data to obtain the concatenated data, and extracts the modal feature data from the concatenated data.
[0083] Specifically, any feature extraction layer from the third to the penultimate layer in each feature extraction network can receive feature data output from the previous feature extraction layer in the same feature extraction network, as well as feature fusion data output from the previous feature fusion layer. The received data is concatenated, features are extracted from the concatenated data, and the extracted feature data is sent to the next feature extraction layer in the same feature extraction network, and to the feature fusion layer of the same number of layers; for example... Figure 2As shown, the third feature extraction layer in C1 can receive the feature data output by the second feature extraction layer in C1, as well as the feature fusion data output by the second feature fusion layer. It concatenates the received data, extracts features from the concatenated data, and sends the extracted feature data to the fourth feature extraction layer and the third feature fusion layer in C1, respectively.
[0084] Specifically, the last feature extraction layer in each feature extraction network can receive feature data output from the previous feature extraction layer in the same network, as well as feature fusion data output from the previous feature fusion layer. It then concatenates the received data, extracts features from the concatenated data, and sends the extracted feature data to the target task network; for example... Figure 2 As shown, the feature extraction layer of layer M in C1 can receive the feature data output by the feature extraction layer of layer M-1 in C1, as well as the feature fusion data output by the feature fusion layer of layer M-1, i.e., layer N. It concatenates the received data, extracts features from the concatenated data, and sends the extracted feature data to the target task network.
[0085] Optionally, in other data processing methods proposed in this embodiment, the number of feature extraction layers in each feature extraction network can be different; for example... Figure 3 The data processing system shown can have three feature extraction layers in feature extraction network C3 and two feature extraction layers in feature extraction network C4. It should be noted that even if the number of feature extraction layers differs in each feature extraction network, the specific data processing and interaction methods of each feature extraction layer and feature fusion layer can all follow the steps S102, S103, S104, and S105 described above. For example, Figure 3 The second feature extraction layer in C3 shown can be combined with... Figure 2 The second feature extraction layer in C1 is the same as that in C3. It can also receive the data output from the first feature extraction layer above the first feature extraction layer in the same feature extraction network, as well as the data output from the first feature fusion layer above the first feature extraction layer. The received data is concatenated, and features are extracted from the concatenated data. The extracted feature data is then sent to the third feature extraction layer below the first feature extraction layer in C3, as well as the second feature fusion layer of the same layer.
[0086] Optionally, in other data processing methods proposed in this embodiment, the value of M minus N can be greater than 1. In this case, the data processing and interaction methods of each feature extraction layer in the feature extraction network and each feature extraction layer in the feature fusion network can also be performed according to the above steps S102, S103, S104, and S105. For example... Figure 4As shown, each feature extraction network has 3 feature extraction layers, and the feature fusion network has 1 feature fusion layer. This feature fusion layer can receive feature data output from the first-layer feature extraction layers in each feature extraction network. Then, the first-layer feature fusion layer can fuse the received feature data output from the first-layer feature extraction layers in each feature extraction network to generate corresponding fused feature data, which is then input into the second-layer feature extraction layers in each feature extraction network.
[0087] Optionally, in other data processing methods proposed in this embodiment, the data processing system may include more than two feature extraction networks. In this case, the data processing and interaction methods of each feature extraction layer in each feature extraction network and each feature extraction layer in the feature fusion network may also be performed according to the above steps S102, S103, S104 and S105.
[0088] It should be noted that the present invention can perform fusion operations on multiple feature layers, and then concatenate the fused features with their respective single-modal features before continuing the calculation, thus preserving the original single-modal features to a large extent, and additionally adding feature information of other modalities after extraction.
[0089] S106. Obtain the task results output by the target task network.
[0090] The target task network can be a neural network designed to perform a specific task, such as lip reading, lip movement detection, person recognition, object detection, or object recognition. Optionally, in the field of intelligent vehicles, the target task network can be used to achieve road environment perception.
[0091] Specifically, this invention can first train the target task network to obtain a trained target task network, and then use the target task network to process the corresponding task, thus ensuring the accuracy of the target task network.
[0092] Specifically, this invention can construct a network for a specific task based on the fused features, according to the specific task requirements, such as multimodal detection, classification, and sequence output.
[0093] Optionally, the target task network includes one pooling layer and two fully connected layers.
[0094] Specifically, the target task network can integrate features through pooling operations, and then perform a two-layer fully connected network to complete classification and regression tasks, obtaining the category for the classification task or the numerical value for the regression task.
[0095] It should be noted that this invention can process multimodal data including images, videos, audio, and text, is applicable to a wide variety of data types, and is suitable for most modalities and tasks, exhibiting full decoupling and greater flexibility. Furthermore, through its modular network design, this invention reduces the coupling between modules and can be effectively integrated with various unimodal solutions and task networks.
[0096] It should also be noted that, since the modal data input into the target task network can largely retain the features of the original single modality, and additional feature information from other extracted modal data can be added, the task processing accuracy of the target task network can be effectively improved. Furthermore, this invention can improve the overall recognition effect after modal fusion without adding too much additional computational burden.
[0097] The data processing method proposed in this embodiment can be applied to a data processing system, which includes a feature fusion network and a feature extraction network group. The feature extraction network group includes at least two feature extraction networks, each corresponding to different single-modal data. The feature fusion network includes at least one feature fusion layer, and at least one feature extraction network includes at least two feature extraction layers. This invention can obtain multiple matching preprocessed modal data. Each preprocessed modal data is input to the first feature extraction layer of each feature extraction network. Before the last feature extraction layer in each feature extraction network, the modal feature data extracted by any feature extraction layer is input to the next feature extraction layer in the same feature extraction network, and also to the feature fusion layer in the same layer. The feature fusion data generated by the fusion processing of the received different modal feature data by any feature fusion layer is input to the next feature extraction layer in each feature extraction network. The modal feature data extracted by the last feature extraction layer in each feature extraction network is input to the target task network to obtain the task result output by the target task network. The modal data input into the target task network by this invention can largely retain the features of the original single modality, and can also add additional feature information of other extracted modal data, thus effectively improving the task processing accuracy of the target task network.
[0098] based on Figure 1 This embodiment proposes a second data processing method. In this method, the feature extraction network group includes a first feature extraction network and a second feature extraction network, and the feature fusion network is a network based on a multi-head attention mechanism.
[0099] Before the last feature extraction layer, each feature extraction layer performs a block-based operation on the modal feature data after extracting the modal feature data, obtaining block-based feature data, and inputting the block-based feature data into the feature fusion layer of the same layer; wherein, the block-based feature data includes multiple block-based feature data.
[0100] After obtaining the segmented feature data input from the feature extraction layers of the two feature extraction networks, each feature fusion layer, based on the segmented feature data input from the feature extraction layer of the first feature extraction network, pays attention to the segmented feature data input from the feature extraction layer of the second feature extraction network to obtain first attention data, and sends the first attention data to the next feature extraction layer in the first feature extraction network; based on the segmented feature data input from the feature extraction layer of the second feature extraction network, it pays attention to the segmented feature data input from the feature extraction layer of the first feature extraction network to obtain second attention data, and sends the second attention data to the next feature extraction layer in the second feature extraction network.
[0101] Specifically, when a feature extraction layer sends its extracted feature data to the feature fusion layer, it can first perform a block-based operation on the feature data to obtain block-based feature data, and then input the block-based feature data into the feature fusion layer. It should be noted that when a feature extraction layer sends feature data to the next feature extraction layer in the same feature extraction network, it does not need to perform a block-based operation; it can directly output the extracted feature data to the next feature extraction layer in the same feature extraction network.
[0102] To better illustrate the block-based operation performed by the feature extraction layer, this invention can use A and B to represent two types of unimodal data for example. Specifically, after obtaining the two types of unimodal data, A and B, this invention can input these two types of unimodal data into two separate feature extraction networks. Then, the l-th feature extraction layer can extract the output of the l-th layer feature vector or feature map, which is defined as the output after the block-based operation. and Where, X∈R batch×tokens×dims `batch` represents the number of data points in a batch during training, `tokens` represents the number of blocks, and `dims` represents the dimension of each block; subsequently, the l-th feature extraction layer can... and The input is fed into the corresponding feature fusion layer.
[0103] It should be noted that the pooling layer in the target task network can be average pooling or max pooling along the dim dimension.
[0104] Specifically, the feature fusion layer can receive segmented feature data sent by the feature extraction layers in each feature extraction network. Then, the feature fusion layer can perform mutual attention operations to fuse the segmented feature data and obtain the corresponding fused feature data.
[0105] Specifically, before performing feature fusion at the feature fusion layer, this invention can predefine computation methods such as self-attention, mutual attention, and concatenation. Among them:
[0106] Self-attention computation can be performed as follows:
[0107]
[0108] Mutual attention can be calculated in the following ways:
[0109] CrossAttention(X A ,X B =Attention(X) A W Q ,X B W K ,X B W V ).
[0110] Specifically, the feature fusion layer can perform feature fusion processing using the predefined computational methods described above after receiving the data output from each feature extraction layer. For example, after receiving the data from the aforementioned feature extraction layers, the feature fusion layer can... and Then, the above calculation method can be used to... and After processing, the following results were obtained:
[0111]
[0112] in, The feature fusion layer can be based on Notice The obtained feature fusion data, The feature fusion layer can be based on Notice The obtained feature fusion data;
[0113] Specifically, the feature fusion layer can obtain feature fusion data. and Afterwards, The data is sent to the next feature extraction layer in the feature extraction network corresponding to modality A. It is sent to the next feature extraction layer in the feature extraction network corresponding to modality data B;
[0114] Specifically, after receiving the feature fusion data from the feature fusion layer, each feature extraction layer can concatenate the feature fusion data with its respective single-modal features according to a predefined concatenation operation to obtain concatenated data. The concatenation operation can be: This represents concatenating multiple inputs to obtain a new output, and its dimension can be represented as batch × (tokens). A +tokens B )×dims. For example, when the feature extraction layer receives the above... After that, you can and The splicing is performed; while the feature extraction layer receives the above... After that, you can and The components are then spliced together. It's understandable that modes A and B each pass through X... AB and X BA The exchange and integration of information has been completed.
[0115] It should be noted that this invention can train each neural network and the target task network through an end-to-end training method. Specifically, this invention can evaluate the trained model and adjust the hyperparameters of the multimodal model from the previous step to optimize it and obtain the optimal model results.
[0116] Specifically, in the application scenario of intelligent vehicles, this invention can deploy a trained and optimized data processing system onto an onboard computing platform. Specifically, this invention can perform necessary post-processing on the output results of the target task network, such as filtering operations, to obtain stable perception results. Subsequently, this invention can output the results to other downstream modules of the intelligent vehicle, such as the decision-making and planning control module, and the cockpit human-machine interaction module.
[0117] It should also be noted that this invention, through the mutual attention mechanism, can complete the information interaction of multimodal data, realize the effective screening and fusion of modal information, avoid information redundancy caused by simple calculation and splicing, effectively reduce the amount of computation, and ensure the training and processing effects of each network.
[0118] The data processing method proposed in this embodiment can achieve information interaction of multimodal data through mutual attention mechanism, realize effective screening and fusion of modal information, avoid information redundancy caused by simple calculation and splicing, effectively reduce the amount of computation and ensure the training effect and processing effect of each network.
[0119] and Figure 1 The corresponding method is as follows, such as Figure 5As shown, this embodiment proposes a data processing system. The data processing system includes: a feature fusion network and a feature extraction network group. The feature extraction network group includes at least two feature extraction networks, each corresponding to different single-modal data. The feature fusion network includes at least one feature fusion layer, and at least one feature extraction network includes at least two feature extraction layers. The data processing system also includes: a first obtaining unit 101, a first input unit 102, a second input unit 103, a third input unit 104, a fourth input unit 105, and a second obtaining unit 106.
[0120] The first obtaining unit 101 is used to obtain multiple preprocessed modal data that match each other;
[0121] The first input unit 102 is used to input each preprocessed modal data into the first feature extraction layer of each feature extraction network respectively;
[0122] The second input unit 103 is used to input the modal feature data extracted by any feature extraction layer into the next feature extraction layer in the same feature extraction network before the last feature extraction layer in each feature extraction network, and to input it into the feature fusion layer in the same layer.
[0123] The third input unit 104 is used to input the feature fusion data generated by fusing the received different modal feature data by any feature fusion layer into the next feature extraction layer in each feature extraction network.
[0124] The fourth input unit 105 is used to input the modal feature data extracted by the last feature extraction layer in each feature extraction network into the target task network;
[0125] The second acquisition unit 106 is used to acquire the task results output by the target task network.
[0126] It should be noted that the specific processing procedures and their technical effects of the first obtaining unit 101, the first input unit 102, the second input unit 103, the third input unit 104, the fourth input unit 105, and the second obtaining unit 106 can be referred to in this embodiment. Figure 1 The relevant explanations of steps S101 to S106 in the method shown will not be repeated here.
[0127] Optionally, each feature extraction network has the same number of feature extraction layers, and each has at least two layers.
[0128] Optionally, the number of feature fusion layers in the feature fusion network is N, and the number of feature extraction layers in each feature extraction network is M, where M minus N equals 1.
[0129] Optionally, the first feature extraction layer in each feature extraction network extracts modal feature data from the corresponding preprocessed modal data.
[0130] Optionally, each feature extraction layer after the first layer in each feature extraction network, after obtaining the modal feature data input from the previous feature extraction layer in the same feature extraction network and the feature fusion data input from the previous feature fusion layer, concatenates the input modal feature data and feature fusion data to obtain the concatenated data, and extracts the modal feature data from the concatenated data.
[0131] Optionally, the feature extraction network group includes a first feature extraction network and a second feature extraction network, and the feature fusion network is a network based on a multi-head attention mechanism.
[0132] Optionally, before the last feature extraction layer, each feature extraction layer performs a block-based operation on the modal feature data after extracting the modal feature data, obtaining block-based feature data, and inputting the block-based feature data into the feature fusion layer of the same layer; wherein, the block-based feature data includes multiple block-based feature data.
[0133] Optionally, after obtaining the segmented feature data input from the feature extraction layers of the two feature extraction networks, any feature fusion layer, based on the segmented feature data input from the feature extraction layer of the first feature extraction network, pays attention to the segmented feature data input from the feature extraction layer of the second feature extraction network to obtain first attention data, and sends the first attention data to the next feature extraction layer in the first feature extraction network; based on the segmented feature data input from the feature extraction layer of the second feature extraction network, pays attention to the segmented feature data input from the feature extraction layer of the first feature extraction network to obtain second attention data, and sends the second attention data to the next feature extraction layer in the second feature extraction network.
[0134] Optionally, the target task network includes one pooling layer and two fully connected layers.
[0135] Optionally, the first obtaining unit 101 includes: a third obtaining unit, a preprocessing unit, and a fourth obtaining unit;
[0136] The third acquisition unit is used to acquire multiple single-mode data output by the multi-mode sensor;
[0137] The preprocessing unit is used to preprocess each of the acquired single-modal data according to a predefined preprocessing method.
[0138] The fourth acquisition unit is used to acquire the matching preprocessed modal data.
[0139] The data processing system proposed in this embodiment includes: a feature fusion network and a feature extraction network group. The feature extraction network group includes at least two feature extraction networks, each corresponding to different single-modal data. The feature fusion network includes at least one feature fusion layer, and at least one feature extraction network includes at least two feature extraction layers. This invention can obtain multiple matching preprocessed modal data. Each preprocessed modal data is input to the first feature extraction layer of each feature extraction network. Before the last feature extraction layer in each feature extraction network, the modal feature data extracted by any feature extraction layer is input to the next feature extraction layer in the same feature extraction network, and also to the feature fusion layer in the same layer. The feature fusion data generated by the fusion processing of the received different modal feature data by any feature fusion layer is input to the next feature extraction layer in each feature extraction network. The modal feature data extracted by the last feature extraction layer in each feature extraction network is input to the target task network to obtain the task result output by the target task network. The modal data input into the target task network by this invention can largely retain the features of the original single modality, and can also add additional feature information of other extracted modal data, thus effectively improving the task processing accuracy of the target task network.
[0140] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0141] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A data processing method, characterized in that, An application is made in a data processing system, the data processing system comprising: a feature fusion network and a feature extraction network group, wherein the feature extraction network group includes at least two feature extraction networks, each of which corresponds to different single-modal data, the single-modal data being any one of fingerprint data, audio data, video data, image data, text data, LiDAR point cloud data, and face data; wherein the feature fusion network includes at least one feature fusion layer, and at least one of the feature extraction networks includes at least two feature extraction layers; the data processing method includes: Obtain multiple matching preprocessed modal data; Each of the preprocessed modal data is respectively input into the first feature extraction layer of each of the feature extraction networks; Before the last feature extraction layer in each of the aforementioned feature extraction networks, the modal feature data extracted by any feature extraction layer is input to the next feature extraction layer in the same feature extraction network, and also input to the feature fusion layer in the same layer. The feature fusion data generated by fusing different modal feature data received by any feature fusion layer is input into the next feature extraction layer in each feature extraction network. Each feature extraction layer after the first layer in each feature extraction network, after obtaining the modal feature data input from the previous feature extraction layer in the same feature extraction network and the feature fusion data input from the previous feature fusion layer, concatenates the input modal feature data and feature fusion data to obtain concatenated data, and extracts modal feature data from the concatenated data. The modal feature data extracted from the last feature extraction layer in each of the aforementioned feature extraction networks is input into the target task network to obtain the task result output by the target task network. The target task network is a neural network used to implement the task, and the task is any one of lip reading, lip movement detection, person recognition, target detection, and target recognition.
2. The data processing method according to claim 1, characterized in that, The feature extraction networks described herein all have the same number of feature extraction layers, and each has at least two layers.
3. The data processing method according to claim 2, characterized in that, The feature fusion network has N layers, and each feature extraction network has M layers, where M is the number of feature fusion layers and N is the number of feature extraction layers. The value obtained by subtracting N from M is 1.
4. The data processing method according to claim 1, characterized in that, The first feature extraction layer in each of the feature extraction networks extracts modal feature data from the corresponding preprocessed modal data.
5. The data processing method according to claim 1, characterized in that, The feature extraction network group includes a first feature extraction network and a second feature extraction network, and the feature fusion network is a network based on a multi-head attention mechanism.
6. The data processing method according to claim 5, characterized in that, Before the last feature extraction layer, each feature extraction layer performs a block-based operation on the modal feature data after extracting the modal feature data, obtaining block-based feature data, and inputting the block-based feature data into the feature fusion layer of the same layer; wherein, the block-based feature data includes multiple block-based feature data.
7. The data processing method according to claim 6, characterized in that, After obtaining the segmented feature data input from the feature extraction layers of the two feature extraction networks, any feature fusion layer, based on the segmented feature data input from the feature extraction layer of the first feature extraction network and paying attention to the segmented feature data input from the feature extraction layer of the second feature extraction network, obtains first attention data and sends the first attention data to the next feature extraction layer in the first feature extraction network; based on the segmented feature data input from the feature extraction layer of the second feature extraction network and paying attention to the segmented feature data input from the feature extraction layer of the first feature extraction network, it obtains second attention data and sends the second attention data to the next feature extraction layer in the second feature extraction network.
8. The data processing method according to claim 1, characterized in that, The target task network includes one pooling layer and two fully connected layers.
9. A data processing system, characterized in that, The data processing system includes: a feature fusion network and a feature extraction network group. The feature extraction network group includes at least two feature extraction networks, each corresponding to different single-modal data. The single-modal data is any one of fingerprint data, audio data, video data, image data, text data, LiDAR point cloud data, and face data. The feature fusion network includes at least one feature fusion layer, and at least one feature extraction network includes at least two feature extraction layers. The data processing system further includes: a first acquisition unit, a first input unit, a second input unit, a third input unit, a fourth input unit, and a second acquisition unit. The first obtaining unit is used to obtain multiple matching preprocessed modal data; The first input unit is used to input each of the preprocessed modal data into the first feature extraction layer of each feature extraction network respectively; The second input unit is used to input the modal feature data extracted by any feature extraction layer into the next feature extraction layer in the same feature extraction network before the last feature extraction layer in each of the feature extraction networks, and to input it into the feature fusion layer of the same layer. The third input unit is used to input the feature fusion data generated by fusing different modal feature data received by any feature fusion layer into the next feature extraction layer in each feature extraction network. Each feature extraction layer after the first layer in each feature extraction network, after obtaining the modal feature data input from the previous feature extraction layer in the same feature extraction network and the feature fusion data input from the previous feature fusion layer, concatenates the input modal feature data and feature fusion data to obtain concatenated data, and extracts modal feature data from the concatenated data. The fourth input unit is used to input the modal feature data extracted by the last feature extraction layer in each of the feature extraction networks into the target task network. The target task network is used as a neural network to implement the task, which is any one of lip reading, lip movement detection, person recognition, target detection, and target recognition. The second obtaining unit is used to obtain the task results output by the target task network.