Method and apparatus for augmenting audio-text multimodal data in training and inference situations
By generating augmented audio-text learning data and using it to enhance the performance of audio-text multimodal models, the method addresses the scarcity of datasets, leading to improved learning and reasoning capabilities in multimodal scenarios.
Patent Information
- Application Number
- PCT/KR2023/017075
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-30
- Filing Date
- 2023-10-31
- Publication Date
- 2025-05-08
AI Technical Summary
The performance of audio-text multimodal models is hindered by the scarcity of public datasets, making it difficult to learn and reason effectively in multimodal scenarios.
The method involves generating augmented audio-text learning data by combining pairs of audio and text data, and then using this enhanced data to improve the performance of audio-text multimodal models through supervised learning and inference processes.
This approach effectively increases the quantity and quality of audio-text multimodal data, leading to improved performance of audio-text multimodal models in learning and reasoning tasks, and enables the construction of more accurate multimodal models.
Smart Images

Figure KR2023017075_08052025_PF_FP_ABST
Abstract
Description
Method for augmenting audio-text multimodal data in learning and inference situations and device therefor
[0001] The present invention relates to a method and a device for augmenting audio-text multimodal data in learning and inference situations. More specifically, the present invention relates to a data augmentation method and a device for enhancing the performance of an artificial intelligence model for audio-text multimodal data, for which the number of publicly available datasets is extremely limited.
[0002] As the performance of artificial intelligence models improves, interest in learning and utilizing multimodal artificial intelligence models that utilize various modalities similar to humans is growing day by day. Representative examples include ChatGPT, which receives not only conventional text input but also images and outputs results through large-scale updates.
[0003] Meanwhile, among multimodal artificial intelligence models, the audio-text multimodal model is a comprehensive term for a model that receives audio or text as input and extracts information between modalities. Its main tasks include audio-based text search, text-based audio search, and audio caption generation, and it is showing limitless potential for use in the multimedia market that provides OTT services, and numerous companies and institutions are actively conducting research and development.
[0004] Since this audio-text multimodal model is a type of artificial intelligence model, training data is essential for learning. AudioCaps, a public audio-text open dataset, contains about 40,000 audio data and about 46,000 text data. Compared to MSCOCO, an image-text open dataset (about 330,000 image data and about 1.5 million text data), it contains only an insufficient amount of data, making it difficult to learn the audio-text multimodal model. Various data augmentation techniques are being discussed to solve this problem, but there is a problem that most of them target a single modality (audio or text).
[0005] Despite these difficulties, the audio-text multimodal model must be improved in performance due to its usability, and the present invention relates to a new and progressive technique for this purpose.
[0006] The technical problem to be solved by the present invention is to provide a method for augmenting audio-text multimodal data in a learning and inference situation that can improve the performance of an audio-text multimodal model by augmenting audio-text multimodal data used for learning an audio-text multimodal model with multi-modality as the target, and a device therefor.
[0007] Another technical problem to be solved by the present invention is to provide a method for augmenting audio-text multimodal data in a learning and inference situation, which can drastically improve the performance of an audio-text multimodal model by augmenting and processing inference data not only during the learning process of the audio-text multimodal model but also during the inference process, and a device therefor.
[0008] The technical problems of the present invention are not limited to the technical problems mentioned above, and other technical problems not mentioned will be clearly understood by those skilled in the art from the description below.
[0009] In order to achieve the above technical task, according to an embodiment of the present invention, a method for augmenting audio-text multimodal data in a learning and inference situation includes (a) a first step of generating audio-text augmented learning data by processing any two audio-text learning data among N (N is a positive integer greater than or equal to 2) pieces of audio-text learning data each consisting of a pair of any audio and a text matched to the any audio, (b) a second step of inputting the N pieces of audio-text learning data and the generated audio-text augmented learning data into an audio-text multimodal model and training the model, and (c) a third step of inputting any inference data into the trained audio-text multimodal model, augmenting and processing the inputted any inference data, and outputting a final prediction value for the any inference data.
[0010] According to one embodiment, the first step may include: (a-1) a step 1-1 of generating augmented audio learning data by adding values of two audio learning data for any two audio-text learning data, (a-2) a step 1-2 of generating augmented text learning data by randomly concatenating two text learning data for any two audio-text learning data, and (a-3) a step 1-3 of generating audio-text augmented learning data by matching the generated augmented audio learning data and augmented text learning data.
[0011] According to one embodiment, the addition of the values of the two audio learning data in the step 1-1 may be performed according to one of the first addition method in the original audio learning data stage and the second addition method in the Mel Spectrogram stage, which is the transformation of the original audio learning data into a time-frequency axis by Short Time Fourier Transform (STFT).
[0012] In one embodiment, the first addition and the second addition can be determined probabilistically.
[0013] According to one embodiment, the random inference data in the third step may be any one of random audio inference data and random text inference data.
[0014] According to one embodiment, the audio-text multimodal model is a model including M layers (M is a positive integer greater than or equal to 2), and in this case, the third step may include at least one of: (c-1) a 3-1 step of augmenting the input arbitrary inference data with a plurality of inference data before passing through the first layer among the M layers, and (c-2) a 3-2 step of calculating one or more first medians by calculating a weighted average of a predetermined number of output values of the plurality of augmented inference data passed through an arbitrary layer after the first layer among the M layers.
[0015] According to one embodiment, if there is one first intermediate value calculated in step 3-2, after step 3-2, (c-3) a step 3-3 may be further included for outputting one output value of the calculated one first intermediate value passed through the last layer among the M layers as the final predicted value.
[0016] According to one embodiment, if there are multiple first intermediate values calculated in step 3-2, after step 3-2, (c-4) a step 3-4 of calculating a second intermediate value by weighting all of the output values that have passed the multiple first intermediate values calculated through an arbitrary layer after the arbitrary layer that has passed through step 3-2 among the M layers, and (c-5) a step 3-5 of outputting a single output value that has passed the multiple first intermediate values calculated through the M layers up to the last layer as the final predicted value may be included.
[0017] In order to achieve the above technical task, according to another embodiment of the present invention, a device for augmenting audio-text multimodal data in a learning and inference situation comprises one or more processors, a network interface, a memory for loading a computer program to be executed by the processors, and a storage for storing large-capacity network data and the computer program, wherein the computer program executes, by the one or more processors, (A) a first step of processing any two audio-text learning data among N (N is a positive integer equal to or greater than 2) pieces of audio-text learning data each consisting of a pair of any audio and a text matched to the any audio to generate audio-text augmented learning data, (B) a second step of inputting the N pieces of audio-text learning data and the generated audio-text augmented learning data into an audio-text multimodal model to train it, and (C) a third step of inputting any inference data into the trained audio-text multimodal model, augmenting and processing the inputted any inference data, and outputting a final prediction value for the any inference data.
[0018] According to another embodiment of the present invention for achieving the above technical task, a computer program stored in a medium is combined with a computing device, and executes the following steps: (AA) a first step of processing any two audio-text learning data among N (N is a positive integer greater than or equal to 2) audio-text learning data each consisting of a pair of any audio and a text matched to the any audio to generate audio-text augmented learning data; (BB) a second step of inputting the N audio-text learning data and the generated audio-text augmented learning data into an audio-text multimodal model to train the model; and (CC) a third step of inputting any inference data into the trained audio-text multimodal model, augmenting and processing the inputted arbitrary inference data, and outputting a final prediction value for the arbitrary inference data.
[0019] According to the present invention as described above, for audio-text multimodal data with an absolutely insufficient number of learning data, an amount of audio-text augmented learning data can be generated through a unique method called PairMix, thereby having the effect of dramatically improving the performance of an audio-text multimodal model that uses the data for learning.
[0020] In addition, the inference data is augmented and used for inference through a unique method called Mid TTA / Multi TTA not only during the learning process but also during the inference process that outputs the actual predicted value, which can further improve the performance of the audio-text multimodal model and at the same time, it has the effect of building an audio-text multimodal model that is robust to errors.
[0021] The effects of the present invention are not limited to the effects mentioned above, and other effects not mentioned will be clearly understood by those skilled in the art from the description below.
[0022] FIG. 1 is a diagram exemplarily illustrating the overall configuration of a device for augmenting audio-text multimodal data in a learning and inference situation according to a first embodiment of the present invention.
[0023] Figure 2 is an example diagram of audio-text multimodal data.
[0024] FIG. 3 is a flowchart illustrating representative steps of a method for augmenting audio-text multimodal data in a learning and inference situation according to a second embodiment of the present invention.
[0025] FIG. 4 is a flowchart detailing a first step of generating audio-text augmented learning data in a method for augmenting audio-text multimodal data in a learning and inference situation according to a second embodiment of the present invention.
[0026] Figure 5 is a diagram illustrating two audio-text multimodal data and data enhanced by processing them.
[0027] FIG. 6 is a drawing showing each of the Mel Spectrogram images converted into a time-frequency axis by performing a short-time Fourier transform on the two audio learning data shown in FIG. 5, and also showing a Mel Spectrogram image obtained by adding these Mel Spectrogram images.
[0028] FIG. 7 is a flowchart illustrating a third step of inputting inference data and outputting a final prediction value in a method for augmenting audio-text multimodal data in a learning and inference situation according to a second embodiment of the present invention.
[0029] Figure 8 is a schematic diagram illustrating step 3-3 corresponding to Mid TTA.
[0030] Figure 9 is a schematic diagram illustrating steps 3-4 and 3-5 corresponding to Multi TTA.
[0031] FIG. 10 and FIG. 11 are performance evaluation results for a method for augmenting audio-text multimodal data in a learning and inference situation according to a second embodiment of the present invention.
[0032] The purpose, technical configuration, and resulting operational effects of the present invention will be more clearly understood through the following detailed description based on the drawings attached to the specification of the present invention. The following describes embodiments of the present invention in detail with reference to the attached drawings.
[0033] The embodiments disclosed herein should not be construed or used to limit the scope of the present invention. Those skilled in the art will readily appreciate that the descriptions of the embodiments herein, including the embodiments, have a wide range of applications. Therefore, any embodiments described in the detailed description of the present invention are intended to serve as illustrative examples to better illustrate the present invention and are not intended to limit the scope of the present invention to the embodiments.
[0034] The functional blocks depicted in the drawings and described below are merely examples of possible implementations. Other implementations may utilize other functional blocks without departing from the spirit and scope of the detailed description. Furthermore, while one or more functional blocks of the present invention are depicted as individual blocks, one or more of the functional blocks of the present invention may be a combination of various hardware and software configurations that perform the same function.
[0035] Additionally, the expression "including certain components" is an "open" expression, simply indicating the presence of those components, and should not be construed as excluding additional components.
[0036] Furthermore, when it is said that a component is "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but there may also be other components in between.
[0037] Hereinafter, detailed embodiments of the present invention will be described with reference to the drawings.
[0038] FIG. 1 is a drawing exemplarily illustrating the overall configuration of a device (100) for augmenting audio-text multimodal data in a learning and inference situation according to a first embodiment of the present invention.
[0039] However, this is only a preferred embodiment for achieving the purpose of the present invention, and some components may be added or deleted as needed, and the role performed by one component may be performed by another component as well.
[0040] A device (100) for augmenting audio-text multimodal data in a learning and inference situation according to a first embodiment of the present invention may include a processor (10), a network interface (20), a memory (30), a storage (40), and a data bus (50) connecting them, and it will be understood that it may further include additional components required to achieve the purpose of the present invention.
[0041] The processor (10) controls the overall operation of each component. The processor (10) may be any one of a CPU (Central Processing Unit), an MPU (Micro Processor Unit), an MCU (Micro Controller Unit), or an artificial intelligence processor of a type widely known in the technical field to which the present invention pertains. In addition, the processor (10) may perform operations for at least one application or program for performing a method for augmenting audio-text multimodal data in a learning and inference situation according to the second embodiment of the present invention. For this purpose, the processor (10) may include an artificial intelligence model exhibiting a predetermined structure, which will be described later.
[0042] The network interface (20) supports wired and wireless Internet communication of the device (100) for augmenting audio-text multimodal data in a learning and inference situation according to the first embodiment of the present invention, and may also support other known communication methods. Accordingly, the network interface (20) may be configured to include a corresponding communication module.
[0043] The memory (30) stores various information, commands, and / or information, and can load one or more computer programs (41) from the storage (40) to perform a method of augmenting audio-text multimodal data in a learning and inference situation according to the second embodiment of the present invention. In Fig. 1, RAM is illustrated as one of the memories (30), but it goes without saying that various storage media can be used as the memory (30).
[0044] Storage (40) can non-temporarily store one or more computer programs (41) and large-capacity network information (42). This storage (40) can be any one of non-volatile memory such as Read Only Memory (ROM), Erasable Programmable ROM (EPROM), Electrically Erasable Programmable ROM (EEPROM), flash memory, hard disk (HDD), solid state drive (SSD), removable disk, or any form of computer-readable recording medium widely known in the art to which the present invention pertains.
[0045] A computer program (41) can be loaded into a memory (30) and executed by one or more processors (10) to: (A) process any two audio-text learning data among N (N is a positive integer greater than or equal to 2) audio-text learning data, each of which is a pair of any audio and a text matched to the random audio, to generate audio-text augmented learning data; (B) input the N audio-text learning data and the generated audio-text augmented learning data into an audio-text multimodal model and train the model; and (C) input any inference data into the trained audio-text multimodal model, augment and process the input random inference data, and output a final prediction value for the random inference data.
[0046] The operations performed by the computer program (41) briefly mentioned above can be viewed as a function of the computer program (41), and a more detailed description will be provided later in the description of a method for augmenting audio-text multimodal data in a learning and inference situation according to the second embodiment of the present invention.
[0047] The data bus (50) serves as a path for transferring commands and / or information between the processor (10), network interface (20), memory (30), and storage (40) described above.
[0048] The device (100) for augmenting audio-text multimodal data in a learning and inference situation according to the first embodiment of the present invention, which has been briefly described above, may be in the form of an independent device, for example, an electronic device or a server (including a cloud), and the electronic device may be not only a desktop PC or a server device that is fixedly installed and used in one place, but also a portable device that is easy to carry, such as a smart phone, a tablet PC, a notebook PC, a PDA, a PMP, etc. Any electronic device that has a CPU corresponding to the processor (10) installed and has only a network function and an audio (microphone, etc.) and text (input means, etc.) input function may be used.
[0049] Hereinafter, assuming that the device (100) for augmenting audio-text multimodal data in a learning and inference situation according to the first embodiment of the present invention is in the form of a "server" among the electronic devices that are independent devices, a process for providing a method for augmenting audio-text multimodal data in a learning and inference situation according to the second embodiment of the present invention through a dedicated application installed on a user terminal (not shown) of a user who wishes to augment audio-text multimodal data in the learning and inference process will be described with reference to FIGS. 2 to 11.
[0050] Figure 2 is an example of audio-text multimodal data.
[0051] Referring to Figure 2, audio-text multimodal data can be confirmed. The sentence "A bell chimes melodically" is text data, and the data that looks like a waveform right below the text data is audio data.
[0052] Here, audio data may be data recorded through a microphone or the like when the above sentence is read aloud, and audio-text multimodal data is data composed of a pair of a specific text and audio data for that text.
[0053] FIG. 3 is a flowchart illustrating representative steps of a method for augmenting audio-text multimodal data in a learning and inference situation according to a second embodiment of the present invention.
[0054] However, this is only a preferred embodiment for achieving the purpose of the present invention, and some steps may be added or deleted as needed, and one step may be included in another step and performed.
[0055] Meanwhile, it is assumed that each step is performed through a device (100) for augmenting audio-text multimodal data in a learning and inference situation according to the first embodiment of the present invention, and since it is assumed that the device (100) for augmenting audio-text multimodal data in a learning and inference situation according to the first embodiment of the present invention is in the form of a “server,” a dedicated application installed in a user terminal (not shown) will be regarded as having the same meaning as the device (100) for augmenting audio-text multimodal data in a learning and inference situation according to the first embodiment of the present invention, and for the convenience of explanation, all of them will be referred to as “devices (100).”
[0056] First, the device (100) processes any two audio-text learning data among N (N is a positive integer greater than or equal to 2) audio-text learning data pairs consisting of any audio and any text matched to the audio to generate audio-text augmented learning data (S310), which is referred to as the first step.
[0057] Here, the N audio-text learning data may be audio-text learning data included in AudioCAPS, a public audio-text open dataset, or may be audio-text learning data created by other users themselves. In any case, it is acceptable to have multiple audio-text learning data to be used for learning.
[0058] The device (100) can process any two audio-text learning data among N audio-text learning data to generate one audio-text augmented learning data, which will be described below with reference to FIG. 4.
[0059] FIG. 4 is a flowchart detailing a first step of generating audio-text augmented learning data in a method for augmenting audio-text multimodal data in a learning and inference situation according to a second embodiment of the present invention.
[0060] However, this is only a preferred embodiment for achieving the purpose of the present invention, and some steps may be added or deleted as needed, and one step may be included in another step and performed.
[0061] First, the device (100) adds the values of two audio learning data for any two audio-text learning data (Element-wise Sum) to generate augmented audio learning data (S310-1), which is referred to as step 1-1, and then concatenates two text learning data in a random order for any two audio-text learning data (Concatenation) to generate augmented text learning data (S310-2), which is referred to as step 1-2.
[0062] FIG. 5 shows not only the audio-text multimodal data exemplarily illustrated in FIG. 2, but also audio-text multimodal data consisting of a pair of text data and voice data for the sentence "A machine beeps continuously", on the left side of the dotted line. If you look at the right side of the dotted line, you can see new augmented audio data in which two audio data are added, and new augmented text data in which two text data are randomly sequenced.
[0063] Meanwhile, since these data are data used for learning, their names include the word learning, and accordingly, they are text learning data, audio learning data, augmented text learning data, and augmented audio learning data. Since there is no priority between text learning data and audio learning data, it is natural that either step 1-1 or step 1-2 can be performed first, or both steps can be performed simultaneously.
[0064] To explain step 1-1 a little more specifically, step 1-1 is to generate augmented audio learning data by adding the values of two audio learning data. Two methods can be followed in adding the values of the two audio learning data. One is the first addition method in which the values of the two audio learning data are added at the original audio learning data stage, as illustrated in Fig. 5, and the other is the second addition method in which the values are added at the Mel Spectrogram stage in which the original audio learning data is transformed into the time-frequency axis by Short Time Fourier Transform (STFT).
[0065] Here, the first addition method can be performed according to the following mathematical expression 1.
[0066] Mathematical formula 1:
[0067] In the above mathematical expression 1, N is the number of audio learning data, a i is the i-th audio learning data signal (Waveform), λ i is a parameter that determines which audio training data signals to add, M is a function that converts the audio training data signals into a Mel Spectrogram, and S w refers to the final augmented audio learning data augmented from the original audio learning data.
[0068] The first addition method performed according to mathematical formula 1 adds at the original audio learning data stage, whereas mathematical formula 1 includes M, a function that converts to a mel spectrogram. This is because, in order for an artificial intelligence model, more specifically an audio-text multimodal model, to input and process audio learning data, a step of applying a mel spectrogram that visualizes the audio learning data and converts it into image data is required, and this is not added at the mel spectrogram stage.
[0069] Meanwhile, the second addition method is not a method of adding at the original audio learning data level like the first addition method, but a method of adding at the Mel Spectrogram level by applying Mel Spectrogram, as shown in Fig. 6 as an example.
[0070] Figure 6 shows the Mel Spectrogram images converted into the time-frequency axis by performing a short-time Fourier transform on the two audio learning data shown in Figure 5, and also shows the Mel Spectrogram image obtained by adding these Mel Spectrogram images.
[0071] Here, the second addition method can be performed according to the following mathematical formula 2.
[0072] Mathematical formula 2:
[0073] In the above mathematical expression 2, S mrefers to the final augmented audio learning data augmented in the Mel Spectrogram unit.
[0074] Meanwhile, when adding the values of two audio learning data, the device (100) can probabilistically determine whether to proceed by selecting either the first addition method or the second addition method according to the following mathematical expression 3.
[0075] Mathematical formula 3:
[0076] In the above mathematical expression 3, S is the final augmented audio learning data, γ is a value (0 or 1) randomly selected from a Bernoulli distribution (p=0.5), and is a parameter that determines which of the two addition methods is probabilistically selected.
[0077] In the first and second steps of augmenting text learning data, there is no room for multiple methods, such as audio learning data, and augmented text learning data can be generated by randomly concatenating two text learning data according to the following mathematical formula 4.
[0078] Mathematical formula 4:
[0079] In the above mathematical expression 4, N is the number of text learning data, Concat is a function that connects N text learning data, and t i is the original i-th text training data, and t is the final augmented text training data.
[0080] The device (100) matches the augmented audio learning data and augmented text learning data generated as described above to generate audio-text augmented learning data (S310-3), which is referred to as step 1-3.
[0081] As mentioned earlier, audio-text multimodal data is composed of a pair of text data and audio data for the same sentence, and the same goes for augmented data. The data composed of a pair of augmented text learning data and augmented audio learning data with two audio learning data added, such as "A bell chimes melodically amachine beeps continuously" on the right side of the dotted line in Fig. 6, becomes audio-text augmented learning data.
[0082] According to the description of these steps 1-1 to 1-3, it is possible to generate one audio-text augmented learning data using two audio-text learning data, and it goes without saying that the generated audio-text augmented learning data can be used as two audio-text learning data that are integrated into N, which is the number of audio-text learning data, to generate another audio-text augmented learning data.
[0083] Meanwhile, when generating audio-text augmented learning data by processing any two audio-text learning data among N audio-text learning data, it is also possible that audio-text learning data that has already been used to generate one audio-text augmented learning data can be reused to generate another audio-text augmented learning data.
[0084] According to the above explanation, even if the number of audio-text learning data is absolutely insufficient, the number of learning data can be increased by continuously generating new audio-text augmented learning data as much as the user wants or as much as the number set in the device (100), which can ultimately contribute to improving the performance of the audio-text multimodal model. Therefore, in the method for augmenting audio-text multimodal data in a learning and inference situation according to the second embodiment of the present invention, the first step of generating audio-text augmented learning data is named PairMix.
[0085] Let us return to the description of Figure 3.
[0086] If audio-text augmented learning data has been generated, the device (100) inputs N audio-text learning data and the audio-text augmented learning data generated in the first step into an audio-text multimodal model to train it (S320), which is referred to as the second step.
[0087] Here, the audio-text multimodal model corresponds to an artificial intelligence model described in the description of the processor (10) above, and any model implemented in any way is acceptable as long as it is an artificial intelligence model including M (M is a positive integer greater than or equal to 2) layers, such as CNN (Convolutional Neural Networks), RNN (Recurrent Neural Networks), LSTM (Long Short-Term Memory), Transformer, AutoEncoder, Reinforcement Learning Model, etc.
[0088] Meanwhile, both the N audio-text learning data and the audio-text augmented learning data generated in the first stage serve as learning data for the audio-text multimodal model, and the learning method can be any method, such as supervised learning, unsupervised learning, reinforcement learning, semi-supervised learning, or self-supervised learning.
[0089] When the device (100) performs up to the second step, the learning of the audio-text multimodal model is completed, and the third step described below is a description of the inference process for outputting an actual prediction value through the learned audio-text multimodal model.
[0090] The device (100) inputs arbitrary inference data into the trained audio-text multimodal model, augments and processes the input arbitrary inference data, and outputs a final prediction value for the arbitrary inference data (S330), which is referred to as the third step.
[0091] Here, the arbitrary inference data may be either arbitrary audio inference data or arbitrary text inference data using the device (100), and when audio inference data is input, the device (100) will output corresponding text data or subtitle data as the final prediction value, and when text inference data is input, the device (100) will output corresponding audio data as the final prediction value.
[0092] Nevertheless, regardless of the type of arbitrary inference data, a common description can be applied to the third step, which will be explained below with reference to Fig. 7.
[0093] FIG. 7 is a flowchart illustrating a third step of inputting inference data and outputting a final prediction value in a method for augmenting audio-text multimodal data in a learning and inference situation according to a second embodiment of the present invention.
[0094] However, this is only a preferred embodiment for achieving the purpose of the present invention, and some steps may be added or deleted as needed, and one step may be included in another step and performed.
[0095] First, the device (100) augments the input random inference data into multiple inference data before passing through the first layer among the M layers (S330-1), which is called step 3-1.
[0096] As mentioned above, the audio-text multimodal model is a model including M layers. The device (100) augments the first layer among the M layers, that is, the layer through which the inference data is input and passes for the first time, with multiple inference data before passing through it. Two or more of the multiple layers is sufficient, and the number may be determined by the selection or setting of the user or the device (100).
[0097] If the input inference data here is audio inference data, known augmentation methods such as SpecAugment and Mixup can be used. If the input inference data is text inference data, known augmentation methods such as EDA (Easy Data Augmentation) and Back Translation can be used. Any other known method can be used.
[0098] Afterwards, the device (100) calculates one or more first median values by calculating a weighted average (Weighted Sum) of a predetermined number of output values of the multiple inference data augmented by the device (100) that pass through any layer after the first layer among the M layers, and this is called step 3-2.
[0099] In step 3-1, if one inference data is augmented into multiple inference data, the augmented multiple inference data will pass through the first layer among the M layers, and then pass through one or more arbitrary layers thereafter to extract and specify features.
[0100] Meanwhile, the characteristics extracted and specified here will be the same as the number of augmented inference data unless there are special circumstances. For example, if the initially input inference data increases to six inference data and passes through the layer, six output values will be output from the layer, and step 3-2 is to calculate one or more first median values by taking a weighted average of a predetermined number of these six output values.
[0101] For example, if a weighted average is calculated between two output values for six output values, three first medians can be produced, if a weighted average is calculated between three output values, two first medians can be produced, and if a weighted average is calculated for all six output values, one first median can be produced. Therefore, it is said that the number of first medians produced is one or more, and the predetermined number here can also be selected or set by the user or the device (100).
[0102] The steps after step 3-2 may differ depending on the number of first intermediate values, and accordingly, step 3-2.5 (S330-2.5) may be further performed to determine whether there are multiple first intermediate values calculated after step 3-2.
[0103] First, in the case where the number of first intermediate values produced is not multiple, that is, when there is only one first intermediate value, the device (100) outputs one output value that passes one first intermediate value produced to the last layer among the M layers as the final predicted value (S330-3), and this is called step 3-3.
[0104] A schematic diagram of the 3-3 step is illustrated as an example in Fig. 8. Referring to Fig. 8, before an input inference data (x) passes through the first layer (f1), it is augmented into three inference data (x1 to x3) according to the 3-1 step, and the three augmented inference data (x1 to x3) pass sequentially from the first layer (f1) and then pass through any layer (f h ) after passing through the weighted average, you can see that one median is produced, and the produced one median is placed in the remaining layers (f h+1 ~ f M ) is sequentially passed through and output as the final predicted value (Pred).
[0105] This method is not like the augmentation method used in some conventional inference processes in which augmentation is performed only in the last layer (e.g., Conventional TTA), but augmentation is also performed for the layers before that. Therefore, in the method for augmenting audio-text multimodal data in a learning and inference situation according to the second embodiment of the present invention, it is named Mid TTA to distinguish it from these.
[0106] This time, when the number of first intermediate values produced is multiple, that is, when there are multiple first intermediate values, the device (100) calculates one second intermediate value by weighting all of the output values that pass through an arbitrary layer after an arbitrary layer that passes through the M layers in step 3-2, and this is called step 3-4. One output value that passes through the M layers to the last layer of the produced second intermediate values is output as the final predicted value (S330-5), and this is called step 3-5.
[0107] A schematic diagram of steps 3-4 and 3-5 is illustrated as an example in Fig. 9. Referring to Fig. 9, before an input inference data (x) passes through the first layer (f1), it is augmented into six inference data (x1 to x6) according to step 3-1, and the six augmented inference data (x1 to x6) pass sequentially from the first layer (f1) and then pass through any layer (f h ), you can see that three first medians are produced by taking a weighted average of the two output values, and the three first medians are then passed through the remaining layers (f h+1 ~ f h´ ) is sequentially passed through the second intermediate value, and the second intermediate value is passed through the last layer (f h´+1 ~ f M ) and output as the final predicted value (Pred).
[0108] In this method, unlike the augmentation method used in some of the conventional inference processes, augmentation is performed not only on the last layer, but also on the layers before that, and since the intermediate value used to output the predicted value includes not only the first intermediate value but also the second intermediate value calculated therefrom, the method for augmenting audio-text multimodal data in a learning and inference situation according to the second embodiment of the present invention is named Multi TTA to distinguish it from these.
[0109] The method of augmenting audio-text multimodal data in a learning and inference situation according to the second embodiment of the present invention through the third step described above enables data to be augmented and used for inference not only during the learning process but also during the inference process, and thus has various effects, which will be described below.
[0110] So far, a method for augmenting audio-text multimodal data in a learning and inference situation according to the second embodiment of the present invention has been described. According to the present invention, for audio-text multimodal data with an absolutely insufficient number of learning data, an amount of audio-text augmented learning data can be generated through a unique method called PairMix, thereby dramatically improving the performance of an audio-text multimodal model that uses this for learning. In addition, not only during the learning process but also during the inference process that outputs the actual predicted value, the inference data is augmented through a unique method called Mid TTA / Multi TTA and used for inference, thereby further improving the performance of the audio-text multimodal model and, at the same time, constructing an audio-text multimodal model that is robust to errors.
[0111] FIG. 10 and FIG. 11 are performance evaluation results for a method for augmenting audio-text multimodal data in a learning and inference situation according to a second embodiment of the present invention.
[0112] Referring to Figure 10, it can be confirmed that the prediction accuracy was higher in the case of using only PairMix, which corresponds to the first stage, compared to the prior art for various evaluation items, and the prediction accuracy was highest in the case of using Multi TTA.
[0113] In addition, referring to FIG. 11, it can be confirmed that the prediction accuracy is significantly higher when only PairMix and up to Multi TTA are used compared to the prior art, both when predicting audio data when the inference data is text data (left) and when predicting text data when the inference data is audio data (right).
[0114] Finally, the device (100) for augmenting audio-text multimodal data in a learning and inference situation according to the first embodiment of the present invention and the method for augmenting audio-text multimodal data in a learning and inference situation according to the second embodiment of the present invention may be implemented as a computer program stored on a computer-readable medium according to the third embodiment of the present invention, in which case, in combination with a computing device, (AA) a first step of processing any two audio-text learning data among N (N is a positive integer greater than or equal to 2) audio-text learning data consisting of a pair of any audio and a text matched to the random audio to generate audio-text augmented learning data, (BB) a second step of inputting the N audio-text learning data and the generated audio-text augmented learning data into an audio-text multimodal model to train it, and (CC) a third step of inputting any inference data into the trained audio-text multimodal model, augmenting and processing the inputted random inference data, and outputting a final prediction value for the random inference data, and although not described in detail for the sake of redundancy, the present invention It should be understood that all technical features applied to the device (100) for augmenting audio-text multimodal data in a learning and inference situation according to the first embodiment and the method for augmenting audio-text multimodal data in a learning and inference situation according to the second embodiment of the present invention can be equally applied to a computer program stored on a computer-readable medium according to the third embodiment of the present invention.
[0115] Although embodiments of the present invention have been described with reference to the attached drawings, those skilled in the art will appreciate that the present invention can be implemented in other specific forms without altering the technical concept or essential features thereof. Therefore, the embodiments described above should be understood to be illustrative in all respects and not restrictive.
Claims
1. A method for augmenting audio-text multimodal data by a device including a processor and a memory, (a) A first step of generating audio-text augmented learning data by processing any two audio-text learning data among N (N is a positive integer greater than or equal to 2) audio-text learning data pairs each consisting of a random audio and a text matched to the random audio; (b) a second step of inputting the N audio-text learning data and the generated audio-text augmented learning data into an audio-text multimodal model to train the model; and (c) a third step of inputting arbitrary inference data into the learned audio-text multimodal model, augmenting and processing the input arbitrary inference data, and outputting a final prediction value for the arbitrary inference data; A method for augmenting audio-text multimodal data including:
2. In paragraph 1, The above first step is, (a-1) Step 1-1 of generating augmented audio learning data by adding the values of two audio learning data for any two audio-text learning data; (a-2) Step 1-2 of generating augmented text learning data by randomly concatenating two text learning data for any two audio-text learning data; and Steps 1-3 of generating audio-text augmented learning data by matching the augmented audio learning data and augmented text learning data generated above; A method for augmenting audio-text multimodal data comprising one or more of:
3. In paragraph 2, The addition of the values of the two audio learning data in the above step 1-1 is The values of the above two audio learning data are either in the first addition method at the original audio learning data stage or in the second addition method at the Mel Spectrogram stage, which is converted to the time-frequency axis by performing a Short Time Fourier Transform (STFT) on the original audio learning data. A method for augmenting audio-text multimodal data.
4. In paragraph 3, The above first addition method and the second addition method are determined probabilistically, A method for augmenting audio-text multimodal data.
5. In paragraph 1, Any inference data in the above third step is, Either arbitrary audio inference data or arbitrary text inference data, A method for augmenting audio-text multimodal data.
6. In paragraph 1, The above audio-text multimodal model is, A model containing M layers (where M is a positive integer greater than or equal to 2), In this case, the third step is, (c-1) Step 3-1 of augmenting the input random inference data into multiple inference data before passing through the first layer among the M layers; and (c-2) Step 3-2 of calculating one or more first medians by calculating a weighted average (Weighted Sum) of a predetermined number of output values of the plurality of augmented inference data passed through any layer after the first layer among the M layers; A method for augmenting audio-text multimodal data comprising one or more of:
7. In paragraph 6, If the first median value calculated in step 3-2 above is one, After the above step 3-2, (c-3) Step 3-3 of outputting one output value of the first intermediate value calculated above, which is passed through the last layer among the M layers, as the final predicted value; A method for augmenting audio-text multimodal data including more.
8. In paragraph 6, If there are multiple first medians calculated in step 3-2 above, After the above step 3-2, (c-4) Step 3-4 of calculating a second intermediate value by weighting all of the output values of the plurality of first intermediate values calculated above and passing through any layer after the arbitrary layer passed through step 3-2 among the M layers; and (c-5) Step 3-5 of outputting one output value of the second intermediate value calculated above, which is passed through the last layer among the M layers, as the final predicted value; A method for augmenting audio-text multimodal data comprising one or more of:
9. One or more processors; network interface; A memory that loads a computer program to be executed by the processor; and Including storage for storing large-capacity network data and the above computer program, The above computer program is executed by the one or more processors, (A) A first step of generating audio-text augmented learning data by processing any two audio-text learning data among N (N is a positive integer greater than or equal to 2) audio-text learning data pairs each consisting of a random audio and a text matched to the random audio; (B) a second step of inputting the N audio-text learning data and the generated audio-text augmented learning data into an audio-text multimodal model to train the model; and (C) A third step of inputting arbitrary inference data into the learned audio-text multimodal model, augmenting and processing the input arbitrary inference data, and outputting a final prediction value for the arbitrary inference data; A device for augmenting audio-text multimodal data that runs.
10. In combination with a computing device, (AA) A first step of generating audio-text augmented learning data by processing any two audio-text learning data among N (N is a positive integer greater than or equal to 2) audio-text learning data pairs each consisting of a random audio and a text matched to the random audio; (BB) A second step of inputting the N audio-text learning data and the generated audio-text augmented learning data into an audio-text multimodal model to train the model; and (CC) A third step of inputting arbitrary inference data into the learned audio-text multimodal model, augmenting and processing the input arbitrary inference data, and outputting a final prediction value for the arbitrary inference data; A computer program stored on a computer-readable medium that executes the program.
Citation Information
Patent Citations
Drone system
KR1020240130983A
Regularization Techniques for End-To-End Speech Recognition
US20190130896A1
Intelligent Training Set Augmentation for Natural Language Processing Tasks
US20220067277A1
Augmented training data for end-to-end models
US20220189461A1