Multimodal feature alignment method, model training method and related device

By encoding and training speech and text features using multimodal feature alignment parameters and generative adversarial techniques, the problem of misalignment between speech and text features is solved, thereby improving the performance and alignment effect of the task model.

CN117034188BActive Publication Date: 2026-04-21HKUST IFLYTEK (SHANGHAI) TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HKUST IFLYTEK (SHANGHAI) TECH CO LTD
Filing Date
2023-07-04
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

In existing technologies, speech and text features cannot be aligned in the representation space, affecting the performance of models for tasks such as speech translation, speech recognition, speech synthesis, and simultaneous speech interpretation.

Method used

Speech and text features are encoded using multimodal feature alignment parameters, and feature type discrimination training is performed using generative adversarial techniques to achieve alignment of different modal features in the representation space. A multimodal feature alignment model is then used for feature alignment.

Benefits of technology

It improved the performance of downstream task models, enhanced the effect of multimodal feature alignment, and strengthened the alignment ability of speech and text features in the representation space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117034188B_ABST
    Figure CN117034188B_ABST
Patent Text Reader

Abstract

The application provides a multimodal feature alignment method, a model training method and related devices. The method comprises: performing an encoding operation on modal data to obtain first modal encoding features; using a predetermined multimodal feature alignment parameter to perform multimodal feature alignment on the first modal encoding features in a representation space to obtain modal alignment features corresponding to the modal data; the multimodal feature alignment parameter is determined through feature type discrimination training, wherein the feature type discrimination training is performed to make it impossible to accurately determine the feature type of the modal alignment features after performing feature alignment on sample speech features and sample text features according to the multimodal feature alignment parameter. The present scheme can use the multimodal feature alignment parameter to align different modal features in the representation space, thereby improving the performance of a downstream task model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of natural language processing technology, and in particular to a multimodal feature alignment method, a model training method, and related apparatus. Background Technology

[0002] When training models for tasks such as speech translation, speech recognition, speech synthesis, and simultaneous interpretation, it is necessary to apply speech features corresponding to speech and text features corresponding to text. Current technologies typically use speech encoders to encode speech to obtain corresponding speech features and text encoders to encode text to obtain corresponding text features. However, speech and text are two related but significantly different modalities. Speech signals are continuous, while text signals are discrete. Therefore, the speech features and text features encoded by the encoder cannot be aligned in the representation space. Consequently, when training models for tasks such as speech translation, speech recognition, speech synthesis, and simultaneous interpretation using these speech features, the inability to align them in the representation space will negatively impact the performance of the task models.

[0003] Therefore, how to align features of different modalities in the representation space is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0004] Based on the defects and shortcomings of the existing technology, this application proposes a multimodal feature alignment method, a model training method, and related devices, which can realize the alignment of different modal features in the representation space, thereby improving the performance of downstream task models.

[0005] The technical solution proposed in this application is as follows:

[0006] According to a first aspect of the embodiments of this application, a multimodal feature alignment method is provided, comprising:

[0007] The modal data is encoded to obtain a first modal coding feature; wherein the modal data includes: speech data and / or text data;

[0008] Using predetermined multimodal feature alignment parameters, the first modality coding feature is aligned in the representation space to obtain the modality alignment feature corresponding to the modality data;

[0009] The multimodal feature alignment parameters are determined through feature type discrimination training. The goal of the feature type discrimination training is to ensure that after aligning the sample speech features and sample text features according to the multimodal feature alignment parameters, the feature type of the modal alignment feature cannot be accurately determined.

[0010] Optionally, the sample speech features are obtained by replacing at least one feature element in the sample speech coding features with a feature element in the sample text coding features, and the sample text features are obtained by replacing at least one feature element in the sample text coding features with a feature element in the sample speech coding features.

[0011] Optionally, the multimodal feature alignment parameters include: modality coding parameters corresponding to the modality data, and shared coding parameters;

[0012] Using predetermined multimodal feature alignment parameters, the first modality coding features are aligned in the representation space to obtain the modality alignment features corresponding to the modality data, including:

[0013] Using the modality coding parameters, the first modality coding feature is processed to obtain the initial modality alignment feature corresponding to the modality data;

[0014] Using the shared coding parameters, the initial modality alignment features are aligned in the representation space to obtain the modality alignment features corresponding to the modality data.

[0015] Optionally, using pre-determined multimodal feature alignment parameters, the first modality coding features are aligned in the representation space to obtain the modality alignment features corresponding to the modality data, including:

[0016] Using a pre-trained multimodal feature alignment model, the first modality coding feature is aligned in the representation space to obtain the modality alignment feature corresponding to the modality data;

[0017] The multimodal feature alignment model is used to align the input features in the representation space according to the multimodal feature alignment parameters.

[0018] Optionally, the training process of the multimodal feature alignment model includes:

[0019] Both the first sample speech coding features and the first sample text coding features are subjected to feature element masking operations to obtain the first sample speech features and the first sample text features.

[0020] The first sample speech features and the first sample text features are input into a pre-constructed multimodal feature alignment model to obtain the first sample alignment features corresponding to the first sample speech features and the second sample alignment features corresponding to the first sample text features.

[0021] The first sample alignment feature and the second sample alignment feature are input into a preset multimodal resolver to obtain a first recognition result and a second recognition result; the first recognition result represents the feature type of the first sample alignment feature, and the second recognition result represents the feature type of the second sample alignment feature;

[0022] Based on the first sample alignment feature, the second sample alignment feature, the first recognition result, and the second recognition result, the multimodal feature alignment model is first trained to obtain the trained multimodal feature alignment model; the first training aims to minimize the similarity between the first recognition result and the true feature type of the first sample alignment feature, as well as the similarity between the second recognition result and the true feature type of the second sample alignment feature, and to make it impossible to distinguish between the first sample alignment feature and the second sample alignment feature by the first recognition result and the second recognition result.

[0023] Optionally, while performing feature element masking on both the first sample speech encoding features and the first sample text encoding features, the method also includes:

[0024] Replace at least one feature element in the first sample speech coding feature with a feature element in the first sample text coding feature, and replace at least one feature element in the first sample text coding feature with a feature element in the first sample speech coding feature.

[0025] Optionally, the multimodal resolver is trained through a second training process, which includes:

[0026] Both the second sample speech coding features and the second sample text coding features are subjected to feature element masking to obtain the second sample speech features and the second sample text features.

[0027] The second sample speech features and the second sample text features are input into a pre-built multimodal feature alignment model to obtain the third sample alignment feature corresponding to the second sample speech features and the fourth sample alignment feature corresponding to the second sample text features;

[0028] The third sample alignment feature and the fourth sample alignment feature are input into a preset multimodal resolver to obtain a third recognition result and a fourth recognition result; the third recognition result represents the feature type of the third sample alignment feature, and the fourth recognition result represents the feature type of the fourth sample alignment feature;

[0029] The parameters of the multimodal discriminator are adjusted with the goal of maximizing the similarity between the third recognition result and the true feature type of the third sample alignment feature, as well as the similarity between the fourth recognition result and the true feature type of the fourth sample alignment feature, to obtain the trained multimodal discriminator.

[0030] Optionally, based on the first sample alignment features, the second sample alignment features, the first recognition result, and the second recognition result, the multimodal feature alignment model is first trained to obtain a trained multimodal feature alignment model, including:

[0031] The masking loss is calculated using the alignment features of the first sample and the speech coding features of the first sample.

[0032] The text loss is calculated using the sample text data corresponding to the second sample alignment feature and the first sample text encoding feature;

[0033] Using the first identification result and the second identification result, calculate and generate the adversarial loss;

[0034] Based on the masking loss, the text loss, the generative adversarial loss, and the predetermined loss weights, the multimodal feature alignment loss is calculated, and the parameters of the multimodal feature alignment model are adjusted based on the multimodal feature alignment loss to obtain the trained multimodal feature alignment model.

[0035] Optionally, the pre-built multimodal feature alignment model includes: an adaptive layer corresponding to the speech modality, an adaptive layer corresponding to the text modality, and a multimodal shared encoder;

[0036] The first sample speech features and the first sample text features are input into a pre-constructed multimodal feature alignment model to obtain a first sample alignment feature corresponding to the first sample speech features and a second sample alignment feature corresponding to the first sample text features, including:

[0037] The first sample speech features are input into the adaptive layer corresponding to the speech modality to obtain the first sample initial alignment features, and the first sample text features are input into the adaptive layer corresponding to the text modality to obtain the second sample initial alignment features.

[0038] Both the first sample initial alignment feature and the second sample initial alignment feature are input into the multimodal shared encoder to obtain the first sample alignment feature and the second sample alignment feature.

[0039] Optionally, based on the first sample alignment features, the second sample alignment features, the first recognition result, and the second recognition result, the multimodal feature alignment model is first trained to obtain a trained multimodal feature alignment model, including:

[0040] Using the initial alignment features of the first sample, the speech coding features of the first sample, the initial alignment features of the second sample, and the text coding features of the first sample, calculate the speech masking loss and the text masking loss;

[0041] Based on the speech masking loss, the parameters of the adaptive layer corresponding to the speech modality are adjusted; based on the text masking loss, the parameters of the adaptive layer corresponding to the text modality are adjusted.

[0042] Using the first sample alignment feature and the first sample speech coding feature, calculate the masking loss; using the sample text data corresponding to the second sample alignment feature and the first sample text coding feature, calculate the text loss; using the first recognition result and the second recognition result, calculate the generation adversarial loss.

[0043] Based on the masking loss, the text loss, the generative adversarial loss, and the predetermined loss weights, a multimodal feature alignment loss is calculated, and the parameters of the multimodal shared encoder are adjusted based on the multimodal feature alignment loss.

[0044] According to a second aspect of the embodiments of this application, a multimodal feature alignment apparatus is provided, comprising:

[0045] A modal data encoding module is used to encode modal data to obtain a first modal encoding feature; wherein the modal data includes: speech data and / or text data;

[0046] The modal feature alignment module is used to perform multimodal feature alignment on the first modal coding feature in the representation space using pre-determined multimodal feature alignment parameters to obtain the modal alignment feature corresponding to the modal data;

[0047] The multimodal feature alignment parameters are determined through feature type discrimination training. The goal of the feature type discrimination training is to ensure that after aligning the sample speech features and sample text features according to the multimodal feature alignment parameters, the feature type of the modal alignment feature cannot be accurately determined.

[0048] According to a third aspect of the embodiments of this application, a model training method is provided, comprising:

[0049] Obtain the modal alignment features corresponding to the model training data; wherein, the modal alignment features corresponding to the model training data are obtained using the above-mentioned multimodal feature alignment method; the model training data includes speech training data and text training data;

[0050] The modality alignment features are input into a pre-built task model, and the task model is trained to obtain a trained task model; the task training includes at least one of speech recognition training, speech translation training, speech synthesis training, and speech simultaneous interpretation training.

[0051] According to a fourth aspect of the embodiments of this application, a model training apparatus is provided, comprising:

[0052] The data processing module is used to obtain the modality alignment features corresponding to the model training data; wherein, the modality alignment features corresponding to the model training data are obtained using the above-mentioned multimodal feature alignment method; the model training data includes speech training data and text training data;

[0053] The model training module is used to input the modality alignment features into a pre-built task model, perform task training on the task model, and obtain a trained task model; the task training includes at least one of speech recognition training, speech translation training, speech synthesis training, and speech simultaneous interpretation training.

[0054] According to a fifth aspect of the embodiments of this application, an electronic device is provided, including: a memory and a processor;

[0055] The memory is connected to the processor and is used to store programs;

[0056] The processor is used to implement the above-described multimodal feature alignment method by running a program in the memory.

[0057] According to a sixth aspect of the embodiments of this application, a storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, it implements the above-described multimodal feature alignment method.

[0058] The multimodal feature alignment method proposed in this application includes: encoding modal data to obtain a first modal encoded feature; wherein the modal data includes: speech data and / or text data; using pre-determined multimodal feature alignment parameters, performing multimodal feature alignment on the first modal encoded feature in the representation space to obtain the modal aligned feature corresponding to the modal data; the multimodal feature alignment parameters are determined through feature type discrimination training, wherein the feature type discrimination training aims to ensure that after feature alignment of sample speech features and sample text features according to the multimodal feature alignment parameters, it is impossible to accurately determine the feature type of the modal aligned feature. Using the technical solution of this application, the alignment of different modal features in the representation space can be achieved using multimodal feature alignment parameters, thereby improving the performance of downstream task models. Furthermore, the generative adversarial technique used in the determination process of the multimodal feature alignment parameters can improve the effect of multimodal feature alignment. Attached Figure Description

[0059] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0060] Figure 1 This is a flowchart illustrating a multimodal feature alignment method provided in an embodiment of this application;

[0061] Figure 2 This is a schematic diagram of the processing flow for training a multimodal feature alignment model provided in an embodiment of this application;

[0062] Figure 3 This is a schematic diagram of the structure of a multimodal feature alignment model provided in an embodiment of this application;

[0063] Figure 4 This is a schematic diagram of the processing flow for training a multimodal resolver provided in an embodiment of this application;

[0064] Figure 5 This is a schematic flowchart of a model training method provided in an embodiment of this application;

[0065] Figure 6 This is a schematic diagram of the structure of a multimodal feature alignment device provided in an embodiment of this application;

[0066] Figure 7 This is a schematic diagram of the structure of a model training device provided in an embodiment of this application;

[0067] Figure 8This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0068] The technical solution of this application embodiment is applicable to multimodal data preprocessing application scenarios. By adopting the technical solution of this application embodiment, the alignment of different modal features in the representation space can be achieved, thereby improving the performance of downstream task models and enhancing the effect of multimodal feature alignment.

[0069] With the increasing demands for human-computer interaction and globalization, tasks such as speech recognition, speech translation, speech synthesis, and simultaneous speech interpretation have become increasingly important. For these tasks, it is typically necessary to pre-train corresponding task models to perform the tasks. Training these models requires sample data containing both speech and text. During training, a speech encoder encodes the speech in the sample data to obtain corresponding speech features, and a text encoder encodes the text in the sample data to obtain corresponding text features. These encoded speech and text features are then input into the task model for training. For example, a speech recognition model is trained for speech recognition, a speech translation model for speech translation, and a speech synthesis model for speech synthesis.

[0070] However, speech and text are two related but significantly different modalities. Speech signals are continuous, while text signals are discrete. Therefore, speech features and text features encoded by the encoder cannot be aligned in the representation space. Consequently, when training task models for tasks such as speech translation, speech recognition, speech synthesis, and simultaneous speech interpretation using speech and text features, the inability to align speech and text features in the representation space will affect the performance of the task models.

[0071] Based on this, this application proposes a multimodal feature alignment method. This technical solution can use multimodal feature alignment parameters to align different modal features in the representation space, thereby solving the problem that speech features and text features cannot be aligned in the representation space in the prior art, which affects the performance of the task model.

[0072] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0073] Exemplary methods

[0074] This application proposes a multimodal feature alignment method, which can be executed by an electronic device. This electronic device can be any device with data and instruction processing capabilities, such as a computer, smart terminal, or server. See also... Figure 1 As shown, the method includes:

[0075] S101. Encode the modal data to obtain the first modal coding feature.

[0076] In the training process of various task models in the field of natural language processing for human-computer interaction, it is necessary to align the modal features of various modal data in the training samples to improve the training effect of the task models. To achieve modal feature alignment, the modal data first needs to be encoded to obtain the first modal encoded features corresponding to the modal data. The modal data includes speech data and / or text data. For speech data, a speech encoder is used to obtain speech encoded features; for text data, a text encoder is used to obtain text encoded features.

[0077] Specifically, in this embodiment, both the speech encoder and the text encoder are pre-trained. To ensure the encoding performance of the speech and text encoders, the pre-trained speech encoder is used for encoding speech data, and the pre-trained text encoder is used for encoding text data. For example, the pre-trained speech encoder can use the w2v-BERT model, and the pre-trained text encoder can use the spanBERT model. Speech data is input into the pre-trained speech encoder for encoding to obtain speech encoding features. Text data is input into the pre-trained text encoder for encoding to obtain text encoding features.

[0078] S102. Using pre-determined multimodal feature alignment parameters, perform multimodal feature alignment on the first modality coding feature in the representation space to obtain the modality alignment feature corresponding to the modality data.

[0079] Specifically, in this embodiment, multimodal feature alignment parameters are predetermined. Using these parameters, the encoded first modality features are aligned in the representation space to obtain the modality-aligned features corresponding to the modality data. In this embodiment, the multimodal feature alignment parameters are predetermined by training the sample features corresponding to the sample data using generative adversarial techniques to determine the feature types.

[0080] In this embodiment, the sample data includes sample speech data and sample text data. Before determining the multimodal feature alignment parameters, it is necessary to determine the sample features of the sample data, namely, the sample speech features corresponding to the sample speech data and the sample text features corresponding to the sample text data. Then, the current multimodal feature alignment parameters are used to perform modal feature alignment on the sample speech features and sample text features respectively, thereby obtaining the modal alignment features after the sample speech data is aligned and the modal alignment features after the sample text data is aligned. The discriminator in the generative adversarial network corresponding to the generative adversarial technique is used to determine the feature type of the modal alignment features after the sample speech data is aligned and the modal alignment features after the sample text data is aligned. Only when the discriminator in the generative adversarial network cannot accurately determine the feature type of the modal alignment features can it be said that the features of the speech modality and the features of the text modality have achieved alignment in the representation space, that is, the features of the speech modality and the features of the text modality share the representation method, so that the discriminator cannot determine the feature type based on the feature representation method. Therefore, when training the multimodal feature alignment parameters for feature type discrimination using sample data, the goal is to continuously adjust the multimodal feature alignment parameters so that after aligning the sample speech features and sample text features according to the multimodal feature alignment parameters, the feature type of the modal alignment features cannot be accurately determined.

[0081] Furthermore, since the alignment of multimodal feature alignment parameters with modal features is merely an alignment in the representation space of the features—that is, using the same representation method for all modal features (i.e., this representation method can have both text-related contextual information representation capabilities and speech-related sequence information representation capabilities)—rather than semantic alignment, the sample data used to train the multimodal feature alignment parameters in this embodiment can be unpaired sample speech data and sample text data. This means the semantics expressed by the sample speech data and the sample text data do not match. The amount of unpaired sample data is larger than that of paired sample data. Therefore, this embodiment can use unpaired sample speech data and sample text data to train the multimodal feature alignment parameters, which can improve the training effect of the multimodal feature alignment parameters and thus improve the alignment effect of multimodal features.

[0082] As described above, the multimodal feature alignment method proposed in this application encodes modal data to obtain a first modal encoded feature. The modal data includes speech data and / or text data. Using pre-determined multimodal feature alignment parameters, the first modal encoded feature is aligned in the representation space to obtain the modal aligned feature corresponding to the modal data. The multimodal feature alignment parameters are determined through feature type discrimination training. The goal of this training is to ensure that after aligning sample speech features and sample text features according to the multimodal feature alignment parameters, the feature type of the aligned feature cannot be accurately determined. By employing the technical solution of this embodiment, the alignment of different modal features in the representation space can be achieved using multimodal feature alignment parameters, thereby improving the performance of downstream task models. Furthermore, the generative adversarial technique used in the determination of the multimodal feature alignment parameters can enhance the effectiveness of multimodal feature alignment.

[0083] As an optional implementation, another embodiment of this application discloses that, in the sample speech data and sample text data in the sample data trained for multimodal feature alignment parameters, the determination of the sample speech features corresponding to the sample speech data and the sample text features corresponding to the sample text data can be achieved by directly using a pre-trained speech encoder to encode the sample speech data to obtain sample speech encoding features, and using a pre-trained text encoder to encode the sample text data to obtain sample text encoding features. The sample speech encoding features are then used as the sample speech features corresponding to the sample speech data, and the sample text encoding features are used as the sample text features corresponding to the sample text data.

[0084] Alternatively, the determination of sample speech features corresponding to sample speech data and sample text features corresponding to sample text data can also be achieved using the following steps: First, the sample speech data needs to be encoded using a pre-trained speech encoder to obtain sample speech encoded features. Then, the sample text data needs to be encoded using a pre-trained text encoder to obtain sample text encoded features. Next, at least one feature element in the sample speech encoded features is replaced with a feature element in the sample text encoded features (e.g., replacing the second feature element in the sample speech encoded features with the second feature element in the sample text encoded features, and replacing the fifth feature element in the sample speech encoded features with the seventh feature element in the sample text encoded features). The replaced feature is then used as the sample speech feature corresponding to the sample speech data. Similarly, at least one feature element in the sample text encoded features is replaced with a feature element in the sample speech encoded features (e.g., replacing the third feature element in the sample text encoded features with the fifth feature element in the sample speech encoded features). The replaced feature is then used as the sample text feature corresponding to the sample text data.

[0085] In this embodiment, by exchanging feature elements between sample speech coding features and sample text coding features, the sample speech features can contain not only features of sample speech data but also features of sample text data, and the sample text features can contain not only features of sample text data but also features of sample speech data. When training multimodal feature alignment parameters using sample speech features and sample text features, the alignment effect of multimodal feature alignment parameters can be improved.

[0086] As an optional implementation, another embodiment of this application discloses that the pre-determined multimodal feature alignment parameters include: modal coding parameters corresponding to the modal data, and shared coding parameters. Correspondingly, step S102 specifically includes the following steps:

[0087] First, the modality coding parameters are used to process the first modality coding features to obtain the initial modality alignment features corresponding to the modality data.

[0088] Specifically, in this embodiment, based on the data type of the pre-encoded modal data, corresponding modal coding parameters are selected. For example, if the modal data is speech data, then its data type is speech; if the modal data is text data, then its data type is text. When performing multimodal feature alignment on the first modal coding feature in this embodiment, firstly, the modal coding parameters corresponding to the modal data are used to perform preliminary processing on the first modal coding feature corresponding to the modal data, thereby obtaining the initial modal alignment feature corresponding to the modal data. The modal coding parameters corresponding to the modal data have a preliminary modal feature alignment capability in the representation space, but this capability is shallower compared to the shared coding parameters of multimodal data. Therefore, this embodiment uses the modal coding parameters corresponding to the modal data to perform preliminary modal feature alignment on the first modal coding feature corresponding to the modal data.

[0089] Second, by using shared coding parameters, the initial modality alignment features are aligned in the representation space to obtain the modality alignment features corresponding to the modality data.

[0090] Specifically, after the first modality coding features undergo preliminary processing, shared coding parameters are used to perform multimodal feature alignment on the initial modality alignment features in the representation space. That is, after using the modality coding parameters corresponding to the modality data to perform shallow multimodal feature alignment on the representation space for the first modality coding features corresponding to the modality data, the shared coding parameters of the multimodality are then used to perform deep multimodal feature alignment on the representation space, thereby obtaining the modality alignment features corresponding to the modality data.

[0091] In this embodiment, by setting the modality coding parameters corresponding to the modality data, the first modality coding feature is initially aligned with the multimodal features in the representation space, thereby achieving fine-tuning of the first modality coding feature. When using shared coding parameters for multimodal feature alignment, the accuracy and efficiency of multimodal feature alignment can be improved.

[0092] As an optional implementation, another embodiment of this application discloses step S102 in the above embodiment, which uses pre-determined multimodal feature alignment parameters to perform multimodal feature alignment on the first modality coding feature in the representation space to obtain the modality alignment feature corresponding to the modality data, specifically including:

[0093] By using a pre-trained multimodal feature alignment model, the first modality encoded features are aligned with multimodal features in the representation space to obtain the modality aligned features corresponding to the modality data.

[0094] This embodiment pre-trains a multimodal feature alignment model, which is used to perform multimodal feature alignment on the input features in the representation space according to pre-determined multimodal feature alignment parameters. Specifically, the first modality encoded features, which are pre-encoded with the modal data, are input into the pre-trained multimodal feature alignment model. The multimodal feature alignment model performs multimodal feature alignment on the first modality encoded features in the representation space to obtain the modality aligned features corresponding to the modality data, so that the modality aligned features have both text-related contextual information representation capabilities and speech-related sequence information representation capabilities.

[0095] As an optional implementation, see [link to relevant documentation]. Figure 2 and Figure 3 As shown in another embodiment of this application, the training process of a pre-trained multimodal feature alignment model is disclosed, including the following steps:

[0096] S201. Perform feature element masking operation on both the first sample speech coding features and the first sample text coding features to obtain the first sample speech features and the first sample text features.

[0097] Specifically, in this embodiment, a first sample data is pre-collected to train the multimodal feature alignment model. The first sample data includes first sample speech data X = {x1, x2, ..., x...}. U} and the first sample text data T = {t1, t2, ..., t N The first sample speech data X is input into a pre-trained speech encoder. The pre-trained speech encoder is used to encode the first sample speech data X to obtain the first sample speech coding feature H. sp That is, H sp =Enc sp(X). The first sample text data T is input into a pre-trained text encoder, and the pre-trained text encoder is used to encode the first sample text data T to obtain the first sample text encoding features H. te That is, H te =Enc te (T).

[0098] This embodiment uses the speech coding feature H of the first sample. sp and the first sample text encoding feature H te Perform feature element masking operations separately, that is, encode feature H for the first sample speech. sp At least one feature element in the first sample is masked to obtain the first sample speech feature, and the first sample text is encoded with the feature H. te At least one feature element in the sample is masked to obtain the first sample text features. For example... Figure 3 As shown, the speech coding features H of the first sample are... sp The third feature element x3 and the fourth feature element x4 in the first sample text encoding feature H were masked. te The fifth feature element t5 in the data was masked.

[0099] S202. Input the first sample speech features and the first sample text features into the pre-constructed multimodal feature alignment model to obtain the first sample alignment features corresponding to the first sample speech features and the second sample alignment features corresponding to the first sample text features.

[0100] In this embodiment, after obtaining the first sample speech features and the first sample text features through masking operations, both the first sample speech features and the first sample text features are input into a pre-constructed multimodal feature alignment model. The multimodal feature alignment model performs multimodal feature alignment on the first sample alignment features and the first sample text features in the representation space, so that the multimodal feature alignment model can learn the sequential information representation ability of speech features in the representation space, as well as the contextual information representation ability of text features in the representation space.

[0101] The multimodal feature alignment model learns the representational capabilities of the first sample speech features and the first sample text features in the representation space by predicting the masked feature elements in the first sample speech features and the masked feature elements in the first sample text features. The multimodal feature alignment model obtains the first sample aligned feature Z corresponding to the first sample speech features through multimodal feature alignment operations on the first sample speech features and the first sample text features in the representation space. sp The second sample alignment feature Z corresponding to the first sample text features te .

[0102] S203. Input the first sample alignment feature and the second sample alignment feature into the preset multimodal resolver to obtain the first recognition result and the second recognition result.

[0103] This embodiment pre-configures a multimodal discriminator, which is a pre-trained multimodal discriminator employing a discriminator structure from generative adversarial networks (GANs) to identify the feature types corresponding to the input features. Specifically, the first sample is aligned with feature Z. sp Second sample alignment feature Z te All inputs are fed into a multimodal resolver, which aligns the first sample's features Z. sp Perform feature type identification to obtain the first identification result, and align feature Z with the second sample. te Feature type identification is performed to obtain a second identification result. The first identification result represents the alignment feature Z of the first sample. sp The feature type, the second recognition result characterizes the second sample alignment feature Z te The feature type. In this embodiment, the recognition result can be represented by a probability method. For example, the recognition result can be the probability that the feature type is speech, or it can be the probability that the feature type is text.

[0104] S204. Based on the first sample alignment features, the second sample alignment features, the first recognition result, and the second recognition result, perform the first training on the multimodal feature alignment model to obtain the trained multimodal feature alignment model.

[0105] This embodiment is based on the first sample alignment feature Z. sp Second sample alignment feature Z te The first sample speech coding feature H sp The first sample text data T, the first recognition result, and the second recognition result are used to train the multimodal feature alignment model, thereby obtaining the trained multimodal feature alignment model. During the first training of the multimodal feature alignment model, the training objectives are: to minimize the similarity between the true feature types of the first recognition result and the first sample alignment feature, and the similarity between the true feature types of the second recognition result and the second sample alignment feature; and to ensure that the first and second sample alignment features cannot be distinguished by the first and second recognition results, i.e., the multimodal discriminator cannot accurately determine the feature types of the first and second sample alignment features.

[0106] Specifically, if the recognition result is the probability that the feature type is speech, then the first recognition result is the first sample aligned feature Z. sp The probability that the feature type is speech type, and the second recognition result is the second sample alignment feature Z. teIf the probability of the feature type being speech is considered, then the goal should be to make the probability of the first recognition result as close to 0 as possible, and the probability of the second recognition result as close to 1 as possible. If the recognition result is the probability of the feature type being text, then the first recognition result is the first sample alignment feature Z. sp The probability that the feature type is text type, and the second recognition result is the second sample alignment feature Z. te If the probability of the feature type being text is given, then the goal should be to make the probability of the first recognition result as close to 1 as possible and the probability of the second recognition result as close to 0 as possible.

[0107] Specifically, the steps for the first training of the multimodal feature alignment model are as follows:

[0108] First, the masking loss is calculated using the alignment features of the first sample and the speech coding features of the first sample.

[0109] In this embodiment, a speech masking loss function is pre-set, and the first sample alignment feature Z is calculated using this speech masking loss function. sp With the first sample speech coding features H sp The masking loss between them. Specifically, the speech masking loss function is:

[0110]

[0111] Among them, L smlm-sh Z represents the alignment feature of the first sample. sp With the first sample speech coding features H sp The masking loss between the two samples, U represents the alignment feature Z of the first sample. sp The total number of feature elements in the middle, H represents the speech coding feature of the first sample. sp The u-th feature element, Z represents the alignment feature of the first sample. sp The u-th feature element.

[0112] Second, the text loss is calculated using the sample text data corresponding to the second sample alignment features and the first sample text encoding features.

[0113] In this embodiment, a text loss function is pre-set, and this text loss function is used to calculate the second sample alignment feature Z. te and the first sample text encoding feature H te The text loss is calculated between the corresponding sample text data (i.e., the first sample text data T). Specifically, the text loss function is:

[0114] L ctc =-logp ctc (T|Z te )

[0115] Among them, L ctc Z represents the alignment feature of the second sample. te The text loss between the first sample text data T and the first sample text data T.

[0116] Third, using the first and second identification results, the adversarial loss is calculated and generated.

[0117] This embodiment pre-sets a generative adversarial loss function. Using this function, along with the first and second recognition results, the generative adversarial loss of the multimodal feature alignment model is calculated. Specifically, the generative adversarial loss function is:

[0118]

[0119] Among them, L align Let K represent the generative adversarial loss of the multimodal feature alignment model, and K represent the total number of training iterations for the multimodal feature alignment model. This represents the alignment feature of the second sample during the k-th training iteration of the multimodal feature alignment model. This represents the alignment features of the first sample during the k-th training iteration of the multimodal feature alignment model. This represents the second recognition result during the k-th training iteration of the multimodal feature alignment model. This represents the first recognition result during the k-th training iteration of the multimodal feature alignment model.

[0120] Fourth, based on the masking loss, text loss, generative adversarial loss, and predetermined loss weights, the multimodal feature alignment loss is calculated, and the parameters of the multimodal feature alignment model are adjusted based on the multimodal feature alignment loss to obtain the trained multimodal feature alignment model.

[0121] In this embodiment, based on the importance of each loss, the corresponding loss weights for masking loss, text loss, and generative adversarial loss are pre-determined. Each loss is multiplied by its corresponding weight and then summed to obtain the final multimodal feature alignment loss of the multimodal feature alignment model. Specifically, the formula for calculating the multimodal feature alignment loss is as follows:

[0122] L sh =α·L smlm-sh +β·L ctc +γ·L align

[0123] Among them, L sh Let L represent the multimodal feature alignment loss, and α represent the masking loss. smlm-sh The loss weights, β represents the text loss L ctc The loss weight, γ, represents the generative adversarial loss L. alignThe loss weight.

[0124] This embodiment aims to minimize the calculated multimodal feature alignment loss by adjusting the parameters of the multimodal feature alignment model until the loss reaches a pre-set range. When the multimodal feature alignment loss reaches the pre-set range, it indicates that the accuracy of the multimodal feature alignment model in aligning features has reached the standard, and the model has completed training.

[0125] As an optional implementation, another embodiment of this application discloses that during the training process of the multimodal feature alignment model, while performing step S201, that is, performing feature element masking operations on both the first sample speech coding features and the first sample text coding features, the following steps are also included:

[0126] Replace at least one feature element in the first sample speech coding feature with a feature element in the first sample text coding feature, and replace at least one feature element in the first sample text coding feature with a feature element in the first sample speech coding feature.

[0127] In this embodiment, the speech coding features H of the first sample are used... sp At least one feature element is selected from the sample for masking, and then at least one feature element is selected to replace the first sample text encoding feature H. te The feature elements in the image are used to obtain the first sample speech feature after masking and replacing the feature elements, denoted as Swap(H). sp ).For example, Figure 3 In the middle, the speech coding features H of the first sample are... sp The third and fourth feature elements in the image are masked, and the speech encoding feature H of the first sample is used. sp The second feature element in the model is replaced with the first sample text encoding feature H. te The second feature element in the code encodes the first sample speech feature H. sp The sixth feature element in the text is replaced with the first sample text encoding feature H. te The fifth feature element in.

[0128] From the text encoding features H of the first sample te At least one feature element is selected from the sample for masking, and then at least one feature element is selected to replace the first sample speech coding feature H. sp The feature elements in the image are used to obtain the first sample text feature after masking and replacing the feature elements, denoted as Swap(H). te ).For example, Figure 3 In the first sample text encoding feature H teThe fifth feature element in the image is masked, and the first sample text encoding feature H is used. te The third feature element in the data is replaced with the first sample speech coding feature H. sp The third feature element in the encoding will encode the first sample text feature H. te The sixth feature element in the code is replaced with the first sample speech coding feature H. sp The sixth feature element in.

[0129] This embodiment uses the first sample speech coding feature H sp With the first sample text encoding feature H te The feature element swapping in the model ensures that the first sample speech feature input into the multimodal feature alignment model includes the first sample text encoding feature H. te The first sample text feature, H, is included in the first sample speech coding feature input into the multimodal feature alignment model. sp The feature elements in the model can improve the alignment effect of modal features between speech modalities and text modalities in the representation space.

[0130] As an optional implementation, see [link to relevant documentation]. Figure 4 As shown in another embodiment of this application, in the above embodiments, the pre-set multimodal resolver is trained through a second training process, wherein the second training process may specifically include the following steps:

[0131] S401. Perform feature element masking operation on both the second sample speech coding features and the second sample text coding features to obtain the second sample speech features and the second sample text features.

[0132] Specifically, in this embodiment, second sample data is pre-collected to train a preset multimodal discriminator. The second sample data includes second sample speech data and second sample text data. The second sample speech data is encoded using a pre-trained speech encoder to obtain second sample speech encoding features, and the second sample text data is encoded using a pre-trained text encoder to obtain second sample text encoding features.

[0133] In this embodiment, at least one feature element is selected from the encoded second sample speech features and the encoded second sample text features respectively to perform a masking operation, thereby obtaining the second sample speech features and the second sample text features.

[0134] Furthermore, in this embodiment, while performing a masking operation on the second sample speech coding features and the second sample text coding features, at least one feature element in the second sample speech coding features can be replaced with a feature element in the second sample text coding features, and at least one feature element in the second sample text coding features can be replaced with a feature element in the second sample speech coding features. The mask of the second sample speech coding features and the features after the feature element replacement are used as the second sample speech features, and the mask of the second sample text coding features and the features after the feature element replacement are used as the second sample text features.

[0135] S402. Input the second sample speech features and the second sample text features into the pre-constructed multimodal feature alignment model to obtain the third sample alignment feature corresponding to the second sample speech features and the fourth sample alignment feature corresponding to the second sample text features.

[0136] In this embodiment, both the second sample speech features and the second sample text features are input into a pre-constructed multimodal feature alignment model. The multimodal feature alignment model performs multimodal feature alignment on the second sample speech features and the second sample text features in the representation space, and outputs the third sample alignment feature corresponding to the second sample speech features and the fourth sample alignment feature corresponding to the second sample text features.

[0137] S403. Input the third sample alignment feature and the fourth sample alignment feature into the preset multimodal resolver to obtain the third recognition result and the fourth recognition result.

[0138] In this embodiment, the third sample alignment feature and the fourth sample alignment feature are input into a preset multimodal resolver to obtain a third recognition result and a fourth recognition result. The third recognition result represents the feature type of the third sample alignment feature, and the fourth recognition result represents the feature type of the fourth sample alignment feature.

[0139] S404. With the goal of maximizing the similarity between the true feature types of the third recognition result and the aligned features of the third sample, and the true feature types of the fourth recognition result and the aligned features of the fourth sample, the parameters of the multimodal discriminator are adjusted to obtain the trained multimodal discriminator.

[0140] This embodiment aims to maximize the similarity between the true feature type of the third identification result and the third sample alignment feature, as well as the similarity between the true feature type of the fourth identification result and the fourth sample alignment feature. In other words, it aims to enable the multimodal discriminator to accurately distinguish the feature types of the third sample alignment feature and the fourth sample alignment feature. The parameters of the multimodal discriminator are adjusted to improve the recognition accuracy of the multimodal discriminator.

[0141] Specifically, if the recognition result is the probability that the feature type is speech, then the third recognition result is the probability that the feature type of the third sample alignment feature is speech, and the fourth recognition result is the probability that the feature type of the fourth sample alignment feature is speech. In this case, the goal should be to make the probability of the third recognition result as close to 1 as possible and the probability of the fourth recognition result as close to 0 as possible. If the recognition result is the probability that the feature type is text, then the third recognition result is the probability that the feature type of the third sample alignment feature is text, and the fourth recognition result is the probability that the feature type of the fourth sample alignment feature is text. In this case, the goal should be to make the probability of the third recognition result as close to 0 as possible and the probability of the fourth recognition result as close to 1 as possible.

[0142] Furthermore, this embodiment pre-sets a loss function for the multimodal resolver. Based on this loss function and the third and fourth recognition results, the recognition loss of the multimodal resolver is calculated. The parameters of the multimodal resolver are adjusted with the goal of minimizing this recognition loss. The loss function of the multimodal resolver is as follows:

[0143]

[0144] Among them, L D Let D(X) represent the recognition loss of the multimodal resolver, K represent the total number of training iterations for the multimodal resolver, and D(X) represent the recognition loss of the multimodal resolver. k D(T) represents the third recognition result during the k-th training of the multimodal discriminator. k ) represents the fourth recognition result during the k-th training of the multimodal discriminator.

[0145] Furthermore, in this embodiment, the training of the multimodal feature alignment model and the multimodal discriminator can be carried out in an alternating training manner. That is, a batch of sample data is collected in advance as the second sample data, and the multimodal discriminator is trained for the first time using the second sample data. Then, a batch of sample data is collected as the first sample data, and the multimodal feature alignment model is trained for the first time using the first sample data. Then, a batch of sample data is collected as the second sample data, and the multimodal discriminator is trained for the second time. Then, a batch of sample data is collected as the first sample data, and the multimodal feature alignment model is trained for the second time, and so on, thereby realizing the alternating training of the multimodal feature alignment model and the multimodal discriminator.

[0146] As an optional implementation, see [link to relevant documentation]. Figure 3 As shown, another embodiment of this application discloses a pre-constructed multimodal feature alignment model, including: an adaptive layer corresponding to the speech modality (i.e., Figure 3 The speech adaptation layer and the text modality-corresponding adaptation layer (i.e., the text mode adaptation layer) are all part of this process. Figure 3The text adaptive layer in the model and the multimodal shared encoder are used. Correspondingly, in the above embodiment, the first sample speech features and the first sample text features are input into a pre-built multimodal feature alignment model to obtain the first sample alignment features corresponding to the first sample speech features and the second sample alignment features corresponding to the first sample text features. Specifically, this includes the following steps:

[0147] First, the speech features of the first sample are input into the adaptive layer corresponding to the speech modality to obtain the initial alignment features of the first sample, and the text features of the first sample are input into the adaptive layer corresponding to the text modality to obtain the initial alignment features of the second sample.

[0148] In this embodiment, the first sample speech feature Swap(H) is used. sp The input is fed into the adaptive layer corresponding to the speech modality. The adaptive layer corresponding to this speech modality applies the first sample speech features Swap(H) to the input. sp Preliminary processing, namely preliminary multimodal feature alignment, is performed to obtain the initial alignment features A of the first sample. sp That is, A sp =Adp sp (Swap(H sp This embodiment uses the first sample text feature Swap(H) to... te The input is fed into the adaptive layer corresponding to the text modality. The adaptive layer corresponding to the text modality applies the Swap(H) feature of the first sample text. te Preliminary processing, namely preliminary multimodal feature alignment, is performed to obtain the initial alignment features A of the second sample. te That is, A te =Adp te (Swap(H te )).

[0149] Second, the initial alignment features of the first sample and the initial alignment features of the second sample are both input into the multimodal shared encoder to obtain the first sample alignment features and the second sample alignment features.

[0150] In this embodiment, the initial alignment feature A of the first sample is... sp Second sample initial alignment feature A te All inputs are fed into a multimodal shared encoder, which aligns the initial features A of the first sample. sp Multimodal feature alignment is performed in the representation space to obtain the first sample alignment feature Z. sp That is, Z sp =Enc sh (A sp The multimodal shared encoder aligns the initial features A of the second sample. te Multimodal feature alignment is performed in the representation space to obtain the second sample alignment feature Z.te ,Right now,

[0151] Z te =Enc sh (A te ).

[0152] Furthermore, since the pre-constructed multimodal feature alignment model includes: an adaptive layer corresponding to the speech modality, an adaptive layer corresponding to the text modality, and a multimodal shared encoder, then, correspondingly, step S204 of the above embodiment, which involves performing a first training on the multimodal feature alignment model based on the first sample alignment features, the second sample alignment features, the first recognition result, and the second recognition result, to obtain the trained multimodal feature alignment model, may specifically include the following steps:

[0153] First, the speech masking loss and text masking loss are calculated using the initial alignment features of the first sample, the speech coding features of the first sample, the initial alignment features of the second sample, and the text coding features of the first sample.

[0154] Specifically, this embodiment pre-sets a speech masking loss function and a text masking loss function. Using the speech masking loss function, the initial alignment feature A of the first sample can be calculated. sp With the first sample speech coding features H sp The speech masking loss is calculated between the given values. The speech masking loss function is:

[0155]

[0156] Among them, L smlm-a Indicates the initial alignment feature A of the first sample. sp With the first sample speech coding features H sp The speech masking loss between the two samples, where U represents the initial alignment feature A of the first sample. sp The total number of feature elements in the middle, H represents the speech coding feature of the first sample. sp The u-th feature element, Indicates the initial alignment feature A of the first sample. sp The u-th feature element.

[0157] Using the text masking loss function, the initial alignment feature A of the second sample can be calculated. te With the first sample text encoding feature H te The text masking loss is calculated between the two. The text masking loss function is:

[0158]

[0159] Among them, L tmlm Indicates the initial alignment feature A of the second sample.te With the first sample text encoding feature H te The text masking loss between the two samples, where N represents the initial alignment feature A of the second sample. te The total number of feature elements in the middle, H represents the text encoding feature of the first sample. te The u-th feature element, Indicates the initial alignment feature A of the second sample. te The u-th feature element.

[0160] Second, based on speech masking loss, the parameters of the adaptive layer corresponding to the speech modality are adjusted, and based on text masking loss, the parameters of the adaptive layer corresponding to the text modality are adjusted.

[0161] This embodiment aims to minimize the calculated speech masking loss by adjusting the parameters of the adaptive layer corresponding to the speech modality, and aims to minimize the calculated text masking loss by adjusting the parameters of the adaptive layer corresponding to the text modality.

[0162] Third, using the first sample alignment feature and the first sample speech coding feature, the masking loss is calculated; using the sample text data corresponding to the second sample alignment feature and the first sample text coding feature, the text loss is calculated; and using the first recognition result and the second recognition result, the adversarial loss is calculated.

[0163] Fourth, based on the masking loss, text loss, generative adversarial loss, and predetermined loss weights, calculate the multimodal feature alignment loss, and adjust the parameters of the multimodal shared encoder based on the multimodal feature alignment loss.

[0164] Specifically, the third and fourth steps have been described in detail in the above embodiments, and will not be repeated in this embodiment.

[0165] This application proposes a model training method that can be executed by an electronic device. This electronic device can be any device with data and instruction processing capabilities, such as a computer, a smart terminal, or a server. See also... Figure 5 As shown, the method includes:

[0166] S501. Obtain the modality alignment features corresponding to the model training data.

[0167] Specifically, in this embodiment, model training data is pre-acquired, and the multimodal feature alignment method described in the above embodiment is used to determine the modal alignment features corresponding to the model training data. The pre-acquired model training data includes speech training data and text training data; therefore, the corresponding modal alignment features include speech modal alignment features corresponding to the speech training data and text modal alignment features corresponding to the text training data.

[0168] S502. Input the modality alignment features into the pre-built task model, train the task model, and obtain the trained task model.

[0169] Modality alignment features are input into a pre-built task model, and the task model is trained to obtain a trained task model. The task model can include at least one of the following: a speech recognition model, a speech translation model, a speech synthesis model, and a simultaneous speech interpretation model. The corresponding task training includes at least one of these three methods. For example, if the task model is a speech recognition model, then the modality alignment features corresponding to the model training data are used to train the speech recognition model for speech recognition, resulting in a trained speech recognition model. If the task model is a speech translation model, then the modality alignment features corresponding to the model training data are used to train the speech translation model for speech translation, resulting in a trained speech translation model. If the task model is a speech synthesis model, then the modality alignment features corresponding to the model training data are used to train the speech synthesis model for speech synthesis, resulting in a trained speech synthesis model. If the task model is a simultaneous speech interpretation model, then the modality alignment features corresponding to the model training data are used to train the simultaneous speech interpretation model for simultaneous speech interpretation, resulting in a trained simultaneous speech interpretation model.

[0170] As can be seen from the above description, the model training method proposed in this application can use a multimodal feature alignment method to align the modal features of model training data of different modalities in the representation space, thereby improving the performance of downstream task models. Furthermore, the generative adversarial technique used in the determination process of the multimodal feature alignment parameters in the multimodal feature alignment method can improve the effect of multimodal feature alignment.

[0171] Exemplary device

[0172] Corresponding to the above-mentioned multimodal feature alignment method, this application also discloses a multimodal feature alignment device, see [link to relevant documentation]. Figure 6 As shown, the device includes:

[0173] The modal data encoding module 100 is used to encode modal data to obtain first modal encoding features; wherein, the modal data includes: speech data and / or text data;

[0174] The modal feature alignment module 110 is used to perform multimodal feature alignment on the first modal coding feature in the representation space using pre-determined multimodal feature alignment parameters to obtain the modal alignment feature corresponding to the modal data;

[0175] The multimodal feature alignment parameters are determined through feature type discrimination training. The goal of this training is to ensure that after aligning the sample speech features and sample text features according to the multimodal feature alignment parameters, the feature type of the modal alignment feature cannot be accurately determined.

[0176] As can be seen from the above description, the multimodal feature alignment device proposed in this application can achieve alignment of different modal features in the representation space by using multimodal feature alignment parameters, thereby improving the performance of downstream task models. Furthermore, the generative adversarial technique used in the determination process of multimodal feature alignment parameters can improve the effect of multimodal feature alignment.

[0177] As an optional implementation, another embodiment of this application discloses that the sample speech features are obtained by replacing at least one feature element in the sample speech coding features with a feature element in the sample text coding features, and the sample text features are obtained by replacing at least one feature element in the sample text coding features with a feature element in the sample speech coding features.

[0178] As an optional implementation, another embodiment of this application discloses that the multimodal feature alignment parameters include: modal coding parameters corresponding to the modal data, and shared coding parameters;

[0179] Correspondingly, the modal feature alignment module 110 is specifically used for:

[0180] By using modality coding parameters, the first modality coding features are processed to obtain the initial modality alignment features corresponding to the modality data;

[0181] By utilizing shared coding parameters, the initial modality alignment features are aligned in the representation space to obtain the modality alignment features corresponding to the modality data.

[0182] As an optional implementation, another embodiment of this application discloses a modal feature alignment module 110, which is specifically used to perform multimodal feature alignment on the first modal coding features in the representation space using a pre-trained multimodal feature alignment model to obtain the modal alignment features corresponding to the modal data.

[0183] The multimodal feature alignment model is used to align the input features in the representation space according to the multimodal feature alignment parameters.

[0184] As an optional implementation, another embodiment of this application discloses that the multimodal feature alignment device of this embodiment further includes: a mask module, a first input module, a second input module, and a first training module.

[0185] The masking module is used to perform feature element masking operations on both the first sample speech coding features and the first sample text coding features to obtain the first sample speech features and the first sample text features.

[0186] The first input module is used to input the first sample speech features and the first sample text features into a pre-built multimodal feature alignment model to obtain the first sample alignment features corresponding to the first sample speech features and the second sample alignment features corresponding to the first sample text features.

[0187] The second input module is used to input the first sample alignment feature and the second sample alignment feature into a preset multimodal resolver to obtain a first recognition result and a second recognition result; the first recognition result represents the feature type of the first sample alignment feature, and the second recognition result represents the feature type of the second sample alignment feature;

[0188] The first training module is used to perform a first training on the multimodal feature alignment model based on the first sample alignment feature, the second sample alignment feature, the first recognition result, and the second recognition result, so as to obtain the trained multimodal feature alignment model. The first training aims to minimize the similarity between the first recognition result and the true feature type of the first sample alignment feature, as well as the similarity between the second recognition result and the true feature type of the second sample alignment feature, and to make it impossible to distinguish between the first sample alignment feature and the second sample alignment feature through the first recognition result and the second recognition result.

[0189] As an optional implementation, another embodiment of this application discloses that the multimodal feature alignment device of this embodiment further includes: a feature replacement module, used to replace at least one feature element in the first sample speech coding feature with a feature element in the first sample text coding feature, and to replace at least one feature element in the first sample text coding feature with a feature element in the first sample speech coding feature.

[0190] As an optional implementation, another embodiment of this application discloses that the multimodal feature alignment device of this embodiment further includes: a second training module.

[0191] The masking module is also used to perform feature element masking operations on both the second sample speech coding features and the second sample text coding features to obtain the second sample speech features and the second sample text features.

[0192] The first input module is also used to input the second sample speech features and the second sample text features into a pre-built multimodal feature alignment model to obtain the third sample alignment feature corresponding to the second sample speech features and the fourth sample alignment feature corresponding to the second sample text features.

[0193] The second input module is also used to input the third sample alignment feature and the fourth sample alignment feature into a preset multimodal resolver to obtain the third recognition result and the fourth recognition result; the third recognition result represents the feature type of the third sample alignment feature, and the fourth recognition result represents the feature type of the fourth sample alignment feature;

[0194] The second training module is used to adjust the parameters of the multimodal discriminator with the goal of maximizing the similarity between the true feature types of the third recognition result and the aligned features of the third sample, as well as the similarity between the fourth recognition result and the aligned features of the fourth sample, so as to obtain the trained multimodal discriminator.

[0195] As an optional implementation, another embodiment of this application discloses that, in the multimodal feature alignment device of this embodiment, the first training module is specifically used for:

[0196] The masking loss is calculated using the alignment features and speech coding features of the first sample.

[0197] The text loss is calculated using the sample text data corresponding to the second sample alignment features and the first sample text encoding features.

[0198] Using the first and second identification results, the adversarial loss is calculated and generated;

[0199] The multimodal feature alignment loss is calculated based on the masking loss, text loss, generative adversarial loss, and pre-determined loss weights. The parameters of the multimodal feature alignment model are then adjusted based on the multimodal feature alignment loss to obtain the trained multimodal feature alignment model.

[0200] As an optional implementation, another embodiment of this application discloses a pre-constructed multimodal feature alignment model, including: an adaptive layer corresponding to the speech modality, an adaptive layer corresponding to the text modality, and a multimodal shared encoder; correspondingly, the first input module is specifically used for:

[0201] The first sample speech features are input into the adaptive layer corresponding to the speech modality to obtain the first sample initial alignment features, and the first sample text features are input into the adaptive layer corresponding to the text modality to obtain the second sample initial alignment features.

[0202] The first sample initial alignment feature and the second sample initial alignment feature are both input into the multimodal shared encoder to obtain the first sample alignment feature and the second sample alignment feature.

[0203] As an optional implementation, another embodiment of this application discloses a first training module, which is further used for:

[0204] Using the initial alignment features of the first sample, the speech coding features of the first sample, the initial alignment features of the second sample, and the text coding features of the first sample, calculate the speech masking loss and the text masking loss;

[0205] Based on speech masking loss, the parameters of the adaptive layer corresponding to the speech modality are adjusted; based on text masking loss, the parameters of the adaptive layer corresponding to the text modality are adjusted.

[0206] Using the first sample alignment feature and the first sample speech coding feature, the masking loss is calculated; using the sample text data corresponding to the second sample alignment feature and the first sample text coding feature, the text loss is calculated; using the first recognition result and the second recognition result, the generative adversarial loss is calculated.

[0207] Based on the masking loss, text loss, generative adversarial loss, and predetermined loss weights, the multimodal feature alignment loss is calculated, and the parameters of the multimodal shared encoder are adjusted based on the multimodal feature alignment loss.

[0208] The multimodal feature alignment device provided in this embodiment belongs to the same concept as the multimodal feature alignment method provided in the above embodiments of this application. It can execute the multimodal feature alignment method provided in any of the above embodiments of this application and has the corresponding functional modules and beneficial effects for executing the multimodal feature alignment method. Technical details not described in detail in this embodiment can be found in the specific processing content of the multimodal feature alignment method provided in the above embodiments of this application, and will not be repeated here.

[0209] Corresponding to the above-described model training method, this application also discloses a model training apparatus, see [link to relevant documentation]. Figure 7 As shown, the device includes:

[0210] The data processing module 200 is used to obtain the modality alignment features corresponding to the model training data; wherein, the modality alignment features corresponding to the model training data are obtained using the multimodal feature alignment method of the above embodiment; the model training data includes speech training data and text training data;

[0211] The model training module 210 is used to input modality alignment features into a pre-built task model, train the task model, and obtain a trained task model; the task training includes at least one of speech recognition training, speech translation training, speech synthesis training, and speech simultaneous interpretation training.

[0212] As can be seen from the above description, the model training device proposed in this application can use the multimodal feature alignment method to align the modal features of model training data of different modalities in the representation space, thereby improving the performance of downstream task models. Furthermore, the generative adversarial technique used in the determination process of the multimodal feature alignment parameters in the multimodal feature alignment method can improve the effect of multimodal feature alignment.

[0213] The model training apparatus provided in this embodiment belongs to the same concept as the model training method provided in the above embodiments of this application. It can execute the model training method provided in any of the above embodiments of this application and has the corresponding functional modules and beneficial effects for executing the model training method. Technical details not described in detail in this embodiment can be found in the specific processing content of the model training method provided in the above embodiments of this application, and will not be repeated here.

[0214] Exemplary electronic devices, storage media, and computer program products

[0215] Corresponding to the above-mentioned multimodal feature alignment method, this application also discloses an electronic device, see [link to relevant documentation]. Figure 8 As shown, the electronic device includes:

[0216] Memory 300 and processor 310;

[0217] The memory 300 is connected to the processor 310 and is used to store programs;

[0218] The processor 310 is configured to implement the multimodal feature alignment method disclosed in any of the above embodiments by running a program stored in the memory 300.

[0219] Specifically, the aforementioned electronic device may further include: a bus, a communication interface 320, an input device 330, and an output device 340.

[0220] The processor 310, memory 300, communication interface 320, input device 330, and output device 340 are interconnected via a bus. Among them:

[0221] A bus can include a pathway for transmitting information between various components of a computer system.

[0222] The processor 310 can be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits used to control the execution of the program of the present application. It can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0223] Processor 310 may include a main processor, as well as a baseband chip, modem, etc.

[0224] The memory 300 stores a program for executing the technical solution of this application, and may also store an operating system and other critical business functions. Specifically, the program may include program code, which includes computer operation instructions. More specifically, the memory 200 may include read-only memory (ROM), other types of static storage devices capable of storing static information and instructions, random access memory (RAM), other types of dynamic storage devices capable of storing information and instructions, disk storage, flash memory, etc.

[0225] Input device 330 may include a device for receiving user input data and information, such as a keyboard, mouse, camera, scanner, light pen, voice input device, touch screen, pedometer, or gravity sensor.

[0226] Output device 340 may include devices that allow information to be output to a user, such as a display screen, printer, speaker, etc.

[0227] The communication interface 320 may include a device that uses any transceiver to communicate with other devices or communication networks, such as Ethernet, Radio Access Network (RAN), Wireless Local Area Network (WLAN), etc.

[0228] The processor 310 executes the program stored in the memory 300 and calls other devices, which can be used to implement the various steps of the multimodal feature alignment method provided in the above embodiments of this application.

[0229] Another embodiment of this application provides a storage medium storing a computer program that, when executed by a processor, implements the various steps of the multimodal feature alignment method provided in any of the above embodiments.

[0230] In addition to the methods and devices described above, embodiments of this application may also be computer program products, which include computer program instructions that, when executed by processor 310, cause processor 310 to perform the various steps of the multimodal feature alignment method provided in the above embodiments.

[0231] Computer program products can be written in any combination of one or more programming languages ​​to perform the operations of the embodiments of this application. The programming languages ​​include object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0232] Furthermore, embodiments of this application may also be computer-readable storage media storing computer program instructions, which, when executed by a processor, cause the processor 310 to perform the various steps of the multimodal feature alignment method provided in the above embodiments.

[0233] The computer-readable storage medium may be any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may, for example, include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0234] Specifically, the specific working content of each part of the aforementioned electronic device, computer program product, and storage medium, as well as the specific processing content of the computer program product or the computer program on the aforementioned storage medium when run by the processor, can all be found in the various embodiments of the aforementioned multimodal feature alignment method, and will not be repeated here.

[0235] For the foregoing method embodiments, in order to simplify the description, they are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, because according to this application, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0236] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For apparatus embodiments, since they are basically similar to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.

[0237] The steps in the methods of the various embodiments of this application can be adjusted, merged, or deleted in order according to actual needs, and the technical features described in each embodiment can be replaced or combined.

[0238] The modules and sub-modules in the apparatus and terminal in the various embodiments of this application can be merged, divided, and deleted according to actual needs.

[0239] It should be understood, in the several embodiments provided in this application, that the disclosed terminals, devices, and methods can be implemented in other ways. For example, the terminal embodiments described above are merely illustrative; for instance, the division of modules or sub-modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple sub-modules or modules may be combined or integrated into another module, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interfaces, devices, or modules, and may be electrical, mechanical, or other forms.

[0240] The modules or submodules described as separate components may or may not be physically separate. The components that constitute a module or submodule may or may not be physical modules or submodules; that is, they may be located in one place or distributed across multiple network modules or submodules. Some or all of the modules or submodules can be selected to achieve the purpose of this embodiment's solution, depending on actual needs.

[0241] Furthermore, the functional modules or sub-modules in the various embodiments of this application can be integrated into one processing module, or each module or sub-module can exist physically separately, or two or more modules or sub-modules can be integrated into one module. The integrated modules or sub-modules described above can be implemented in hardware or in the form of software functional modules or sub-modules.

[0242] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0243] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software unit executed by a processor, or a combination of both. The software unit can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0244] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0245] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A multi-modal feature alignment method, characterized in that, The method comprises: performing an encoding operation on modal data to obtain first modal encoding features; wherein the modal data comprises voice data and / or text data; performing multi-modal feature alignment on the first modal encoding features in a representation space by using pre-determined multi-modal feature alignment parameters to obtain modal alignment features corresponding to the modal data; wherein the modal alignment features have text-related context information representation capability and voice-related sequence information representation capability; The multi-modal feature alignment parameters are determined through feature type discrimination training, wherein the feature type discrimination training is performed to make it impossible to accurately determine the feature type of the modal alignment features after performing feature alignment on sample voice features and sample text features according to the multi-modal feature alignment parameters, which is the target; the sample voice features are obtained by replacing at least one feature element in sample voice encoding features with a feature element in sample text encoding features, and the sample text features are obtained by replacing at least one feature element in sample text encoding features with a feature element in sample voice encoding features.

2. The method of claim 1, wherein, The multi-modal feature alignment parameters comprise modal encoding parameters corresponding to the modal data and shared encoding parameters; performing multi-modal feature alignment on the first modal encoding features in a representation space by using pre-determined multi-modal feature alignment parameters to obtain modal alignment features corresponding to the modal data, comprising: processing the first modal encoding features by using the modal encoding parameters to obtain initial modal alignment features corresponding to the modal data; performing multi-modal feature alignment on the initial modal alignment features in a representation space by using the shared encoding parameters to obtain modal alignment features corresponding to the modal data.

3. The method of claim 1, wherein, performing multi-modal feature alignment on the first modal encoding features in a representation space by using pre-determined multi-modal feature alignment parameters to obtain modal alignment features corresponding to the modal data, comprising: performing multi-modal feature alignment on the first modal encoding features in a representation space by using a pre-trained multi-modal feature alignment model to obtain modal alignment features corresponding to the modal data; The multi-modal feature alignment model is used to perform multi-modal feature alignment on input features in a representation space according to the multi-modal feature alignment parameters.

4. The method of claim 3, wherein, The training process of the multi-modal feature alignment model comprises: performing feature element masking operations on first sample voice encoding features and first sample text encoding features to obtain first sample voice features and first sample text features; inputting the first sample voice features and the first sample text features into a pre-constructed multi-modal feature alignment model to obtain first sample alignment features corresponding to the first sample voice features and second sample alignment features corresponding to the first sample text features; inputting the first sample alignment feature and the second sample alignment feature into a preset multi-modal discriminator to obtain a first recognition result and a second recognition result; the first recognition result represents a feature type of the first sample alignment feature, and the second recognition result represents a feature type of the second sample alignment feature; performing first training on the multi-modal feature alignment model according to the first sample alignment feature, the second sample alignment feature, the first recognition result and the second recognition result to obtain a trained multi-modal feature alignment model; the first training aims to minimize similarities between the first recognition result and a true feature type of the first sample alignment feature and between the second recognition result and a true feature type of the second sample alignment feature, and to make the first recognition result and the second recognition result unable to distinguish the first sample alignment feature and the second sample alignment feature.

5. The method of claim 4, wherein, The feature element masking operation on the first sample speech coding feature and the first sample text coding feature also includes: replacing at least one feature element in the first sample speech coding feature with a feature element in the first sample text coding feature, and replacing at least one feature element in the first sample text coding feature with a feature element in the first sample speech coding feature.

6. The method of claim 4, wherein, The multi-modal discriminator is trained through a second training process, and the second training process includes: performing feature element masking operation on the second sample speech coding feature and the second sample text coding feature to obtain a second sample speech feature and a second sample text feature; inputting the second sample speech feature and the second sample text feature into a pre-constructed multi-modal feature alignment model to obtain a third sample alignment feature corresponding to the second sample speech feature and a fourth sample alignment feature corresponding to the second sample text feature; inputting the third sample alignment feature and the fourth sample alignment feature into a preset multi-modal discriminator to obtain a third recognition result and a fourth recognition result; the third recognition result represents a feature type of the third sample alignment feature, and the fourth recognition result represents a feature type of the fourth sample alignment feature; adjusting parameters of the multi-modal discriminator to obtain a trained multi-modal discriminator, aiming to maximize similarities between the third recognition result and a true feature type of the third sample alignment feature and between the fourth recognition result and a true feature type of the fourth sample alignment feature.

7. The method of claim 4, wherein, performing first training on the multi-modal feature alignment model according to the first sample alignment feature, the second sample alignment feature, the first recognition result and the second recognition result to obtain a trained multi-modal feature alignment model, including: calculating a masking loss by using the first sample alignment feature and the first sample speech coding feature; calculating a text loss by using the second sample alignment feature and sample text data corresponding to the first sample text coding feature; and Using the first identification result and the second identification result, calculate and generate the adversarial loss; Based on the masking loss, the text loss, the generative adversarial loss, and the predetermined loss weights, the multimodal feature alignment loss is calculated, and the parameters of the multimodal feature alignment model are adjusted based on the multimodal feature alignment loss to obtain the trained multimodal feature alignment model.

8. The method of claim 4, wherein, The pre-built multimodal feature alignment model includes: an adaptive layer corresponding to the speech modality, an adaptive layer corresponding to the text modality, and a multimodal shared encoder; The first sample speech features and the first sample text features are input into a pre-constructed multimodal feature alignment model to obtain a first sample alignment feature corresponding to the first sample speech features and a second sample alignment feature corresponding to the first sample text features, including: The first sample speech features are input into the adaptive layer corresponding to the speech modality to obtain the first sample initial alignment features, and the first sample text features are input into the adaptive layer corresponding to the text modality to obtain the second sample initial alignment features. Both the first sample initial alignment feature and the second sample initial alignment feature are input into the multimodal shared encoder to obtain the first sample alignment feature and the second sample alignment feature.

9. The method of claim 8, wherein, Based on the first sample alignment features, the second sample alignment features, the first recognition result, and the second recognition result, the multimodal feature alignment model is first trained to obtain the trained multimodal feature alignment model, including: Using the initial alignment features of the first sample, the speech coding features of the first sample, the initial alignment features of the second sample, and the text coding features of the first sample, calculate the speech masking loss and the text masking loss; Based on the speech masking loss, the parameters of the adaptive layer corresponding to the speech modality are adjusted; based on the text masking loss, the parameters of the adaptive layer corresponding to the text modality are adjusted. Using the first sample alignment feature and the first sample speech coding feature, calculate the masking loss; using the sample text data corresponding to the second sample alignment feature and the first sample text coding feature, calculate the text loss; using the first recognition result and the second recognition result, calculate the generation adversarial loss. Based on the masking loss, the text loss, the generative adversarial loss, and the predetermined loss weights, a multimodal feature alignment loss is calculated, and the parameters of the multimodal shared encoder are adjusted based on the multimodal feature alignment loss.

10. A multi-modal feature alignment apparatus, characterized by, include: A modal data encoding module is used to encode modal data to obtain a first modal encoding feature; wherein the modal data includes: speech data and / or text data; The modal feature alignment module is used to perform multimodal feature alignment on the first modal coding feature in the representation space using pre-determined multimodal feature alignment parameters to obtain the modal alignment feature corresponding to the modal data; wherein, the modal alignment feature has the ability to represent text-related contextual information and speech-related sequence information. The multi-modal feature alignment parameter is determined through feature type discrimination training, where the feature type discrimination training is to make it impossible to accurately determine the feature type of the modality alignment feature after performing feature alignment on the sample speech feature and the sample text feature according to the multi-modal feature alignment parameter, as the target; the sample speech feature is obtained by replacing at least one feature element in a sample speech coding feature with a feature element in a sample text coding feature, and the sample text feature is obtained by replacing at least one feature element in a sample text coding feature with a feature element in a sample speech coding feature.

11. A model training method, comprising: The method comprises: obtaining a modality alignment feature corresponding to model training data; where the modality alignment feature corresponding to the model training data is obtained by using the multi-modal feature alignment method according to any one of claims 1-9; the model training data comprises speech training data and text training data; inputting the modality alignment feature into a pre-constructed task model to perform task training on the task model, and obtaining a trained task model; the task training comprises at least one of speech recognition training, speech translation training, speech synthesis training, and speech simultaneous interpretation training.

12. A model training apparatus, comprising: The method comprises: a data processing module configured to obtain a modality alignment feature corresponding to model training data; where the modality alignment feature corresponding to the model training data is obtained by using the multi-modal feature alignment method according to any one of claims 1-9; the model training data comprises speech training data and text training data; a model training module configured to input the modality alignment feature into a pre-constructed task model to perform task training on the task model, and obtain a trained task model; the task training comprises at least one of speech recognition training, speech translation training, speech synthesis training, and speech simultaneous interpretation training.

13. An electronic device, comprising: The method comprises: a memory and a processor; wherein the memory is connected with the processor and is configured to store programs; the processor is configured to realize the multi-modal feature alignment method according to any one of claims 1-9 by running the programs in the memory.

14. A storage medium, characterized by The storage medium has a computer program stored thereon, and the computer program is executed by the processor to realize the multi-modal feature alignment method according to any one of claims 1-9.