Learning device, learning method, and learning program

The learning device enhances cross-modal task estimation by extracting and concatenating features from unimodal and multimodal data with embedded modality information, ensuring accurate task estimation despite data limitations.

JP7747216B2Active Publication Date: 2025-10-01NIPPON TELEGRAPH & TELEPHONE CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024534811
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-07-19
Publication Date
2025-10-01
Estimated Expiration
2042-07-19

AI Technical Summary

Technical Problem

Conventional cross-modal task estimation methods suffer from decreased accuracy when large amounts of multimodal pair data are unavailable, and the inability to utilize unimodal data, especially when modal data is missing, leading to limited training and inference challenges.

Method used

A learning device that extracts encoding features from both unimodal and multimodal data, embeds segment information to identify modality, concatenates these features, and calculates model parameters to estimate cross-modal tasks, enabling accurate estimation even with limited data.

Benefits of technology

The device ensures high accuracy in cross-modal task estimation using both multimodal and unimodal data, preventing accuracy drops and allowing estimation even when modal data is missing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007747216000014
    Figure 0007747216000014
  • Figure 0007747216000015
    Figure 0007747216000015
  • Figure 0007747216000016
    Figure 0007747216000016
Patent Text Reader

Abstract

A training device (100) is configured to: extract temporal-directional encoding features on the basis of input data that is one or both of single-modal data, which is data of a single modality, and multi-modal pair data containing a plurality of different modalities; embed segment information, which is information for identifying the type of modality of the input data, into the encoding features on the basis of a predetermined condition; use input conditions of the segment-embedded features, which are obtained by embedding the segment information, to couple the plurality of segment-embedded features in the temporal direction to create modal-coupled features; and calculate model parameters by using one or both of the segment-embedded features and the modal-coupled features, and an estimation vector estimated for a cross-modal task on the basis of ground truth data.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a learning device, a learning method, and a learning program. [Background technology]

[0002] In estimating cross-modal tasks, a known method for optimizing the parameters of each neural network using a neural network and a large amount of data is to simultaneously input paired data of different modalities (hereinafter simply referred to as "multi-modal paired data") into a neural network and have the neural network learn the features of each modality (see, for example, Non-Patent Documents 1 and 2). Note that the above-mentioned cross-modal task refers to a task that has a common output for input data of different modalities.

[0003] Furthermore, the aforementioned multimodal paired data is, for example, "sequential image data that is image data of a person's face" or "audio data of a person's voice" included in one video data having a correct answer label of a certain emotion category in an emotion recognition task (for example, a task of classifying human emotions into categories such as "sadness" and "happiness"). When an emotion recognition task is performed in the estimation of a cross-modal task using conventional techniques, video data including multimodal paired data may be used as training data or input data for inference. [Prior art documents] [Non-patent literature]

[0004] [Non-Patent Document 1] Valentin Vielzeuf and Stephane Pateux and Frederic Jurie: Temporal multimodal fusion for video emotion classification in the wild, In Proc. ACM International Conference on Multimodal Interaction (ICMI), P.569--57, 2017 [Non-patent document 2] Panagiotis Tzirakis and George Trigeorgis and Mihalis A. Nicolaou and Bjorn W. Schuller and Stefanos Zafeiriou: End-to-End Multimodal Emotion Recognition Using Deep Neural Networks, IEEE JOURNAL OF SELECTED TOPICS IN SIGNAL PROCESSING Vol.11, Number.8, P.1301-1309, 2017 Summary of the Invention [Problem to be solved by the invention]

[0005] However, in conventional techniques, the accuracy of cross-modal task estimation decreases when a large amount of multimodal pair data cannot be prepared, and in cases where a modal is missing, it is sometimes impossible to estimate the task from data of one modality.

[0006] For example, in conventional technology, cross-modal task estimation involves inputting multimodal pair data simultaneously into a neural network, which poses a problem in that the estimation device only operates when multimodal pair data is input simultaneously, making it impossible to utilize single-modal data (hereinafter simply referred to as "unimodal data").

[0007] On the other hand, while parameter estimation and updating using large amounts of data is effective in neural network training, it is difficult to collect large amounts of multimodal pair data. This results in limited training data, which can lead to a decrease in task accuracy. Furthermore, there are limited cases where multimodal pair data can be prepared at the time of inference, and there is a problem that inference is impossible when there are missing modals. [Means for solving the problem]

[0008] In order to solve the above problems and achieve the object, the learning device of the present invention is characterized by comprising: an extraction unit that extracts encoding features having a time series direction based on input data of both or either unimodal data, which is data of a single modality, and multimodal pair data including a plurality of different modalities; an embedding unit that embeds segment information, which is information that identifies the modal type of the input data, into the encoding features based on predetermined conditions; a concatenation unit that concatenates a plurality of the segment embedding features in the time series direction as modal concatenation features based on input conditions of the segment embedding features into which the segment information is embedded; and a calculation unit that calculates model parameters using the segment embedding features and / or the modal concatenation features, and an estimated vector of a cross-modal task estimated based on ground truth data. [Effects of the Invention]

[0009] The present invention has the advantage that it is possible to suppress a decrease in accuracy when estimating cross-modal tasks when a large amount of multimodal pair data cannot be prepared, and that it is possible to estimate tasks from data of one modal when there is a modal missing. [Brief explanation of the drawings]

[0010] [Figure 1] FIG. 1 is a diagram illustrating an example of an outline of a learning method of a learning device according to an embodiment. [Figure 2]FIG. 2 is a diagram illustrating an example of the device configuration of the learning device according to the embodiment. [Figure 3] FIG. 3 is a diagram illustrating an example of a segment embedding feature linking process according to the embodiment. [Figure 4] FIG. 4 is a diagram illustrating an example of a modified example of the learning device according to the embodiment. [Figure 5] FIG. 5 is a diagram illustrating an example of a flowchart of a learning method according to the embodiment. [Figure 6] FIG. 6 is a diagram illustrating an example of a computer in which the learning device according to the embodiment is implemented. DETAILED DESCRIPTION OF THE INVENTION

[0011] Hereinafter, a mode for carrying out the present embodiment (hereinafter referred to as "embodiment") will be described with reference to the drawings. Note that the present embodiment is not limited to the content described below.

[0012] [1. Overview of learning methods] The learning device 100 of this embodiment has a mechanism that can input both multimodal pair data and unimodal data as training data, trains a neural network based on the training data, and estimates an estimation vector for a cross-modal task. The learning device 100 then calculates and applies neural network parameters to the training data. As a result, the learning device 100 estimates a task from both multimodal pair data and unimodal data during inference.

[0013] First, an example of an outline of a learning method by the learning device 100 of this embodiment will be described using FIG. 1. First, the learning device 100 acquires, as learning data 10, unimodal data 1 (Equation (1)), unimodal data 2 (Equation (2)), and multimodal paired data 3 (Equation (3)), which are expressed by the following equations. Note that hereafter, unimodal data 1 will be consistently referred to as "unimodal data 1 (Equation (1))," unimodal data 2 as "unimodal data 2 (Equation (2))," and multimodal paired data 3 as "multimodal paired data 3 (Equation (3))." Note that the parameters included in Equations (1) to (3) will be described in detail in later sections.

[0014]

number

number

number

[0015] Next, the extraction unit 131 of the model parameter learning unit 130 extracts and outputs an encoding feature having a time series direction based on the acquired training data 10. Then, the embedding unit 132 of the model parameter learning unit 130 embeds segment information, which is information identifying the modal type of the training data, into the encoding feature based on predetermined conditions, and outputs a segment embedding feature. Next, the connection unit 133 of the model parameter learning unit 130 connects the segment embedding features in the time series direction and adds zero-filled portions (zero-filled data that has a mechanism for not performing learning in the calculation unit 135 described below; hereinafter, simply referred to as "zero-filled portions") based on the input conditions of the segment embedding features, and outputs the modal connection feature or the segment embedding feature after the zero-filled portions have been added.

[0016] Furthermore, the estimation unit 134 of the model parameter learning unit 130 converts the modal link feature or the segment embedding feature after adding zero-filled portions using an arbitrary neural network function, and estimates a vector corresponding to the ground truth data Y as an estimated vector for the cross-modal task. Subsequently, the calculation unit 135 of the model parameter learning unit 130 calculates the model parameter θ using the estimated vector for the cross-modal task.

[0017] Then, the cross-modal task estimation unit 140 receives the model parameter θ as an input and estimates an estimation vector Z corresponding to the unimodal data s and the multimodal pair data m included in the inference input data 20.

[0018] 2. Configuration of the learning device Next, the configuration of the learning device 100 according to this embodiment will be described with reference to Fig. 2. As shown in Fig. 2, the learning device 100 includes a communication unit 110, a storage unit 120, a model parameter learning unit 130, and a cross-modal task estimation unit 140. Although not shown, the learning device 100 may also include an input unit (e.g., a keyboard, a mouse, etc.) that accepts various operations, and a display unit (e.g., a display, etc.) that displays various information. Next, the detailed functions of each unit will be described below.

[0019] (Communication unit 110) The communication unit 110 is realized by a NIC (Network Interface Card) or the like, and controls communication via an electric communication line such as a LAN (Local Area Network), the Internet, etc. The communication unit 110 is connected to a network by wire or wirelessly as necessary, and can transmit and receive information bidirectionally.

[0020] (Storage unit 120) The storage unit 120 stores data and programs necessary for various processes by the model parameter learning unit 130 and the cross-modal task estimation unit 140. The storage unit 120 also includes a modal data storage unit 121, an encoding feature storage unit 122, a segment embedding feature storage unit 123, a modal concatenation feature storage unit 124, a model parameter storage unit 125, and an estimated vector storage unit 126. The storage unit 120 is realized by a semiconductor memory element such as a RAM (Random Access Memory) or a flash memory, or a storage device such as a hard disk or an optical disk.

[0021] (Modal data storage unit 121) The modal data storage unit 121 stores multimodal pair data and unimodal data. For example, in an emotion recognition task, the modal data storage unit 121 stores, as unimodal data, voice modal data in which human voices are labeled with emotion categories, image modal data in which sequential image data of human facial movements are labeled with emotion categories, etc. Note that the information stored in the modal data storage unit 121 is not limited to the above-mentioned modal data, etc., and other modal data may be stored.

[0022] (Encoding feature storage unit 122) The encoding feature storage unit 122 stores encoding features having a time series direction that are extracted from any modal data by an extraction unit 131 described later.

[0023] (Segment embedding feature storage unit 123) The segment embedding feature storage unit 123 stores segment embedding features in which the embedding unit 132, which will be described later, embeds segment information, which is information for identifying the modal type of the learning data, into the encoding features.

[0024] (Modal link feature storage unit 124) When a plurality of segment embedding features are input, the modal concatenation feature storage unit 124 stores the modal concatenation feature that is output by the concatenation unit 133 after concatenating the plurality of segment embedding features in the time series direction and adding zero-filled portions. Furthermore, the modal concatenation feature storage unit 124 also stores segment embedding features after adding zero-filled portions that are not subjected to concatenation processing.

[0025] (Model parameter storage unit 125) The model parameter storage unit 125 stores the model parameter θ calculated by the calculation unit 135, which will be described later.

[0026] (Estimated vector storage unit 126) The estimated vector storage unit 126 stores an estimated vector of a cross-modal task estimated by an estimation unit 134 (described later) and an estimated vector Z estimated by a cross-modal task estimation unit 140.

[0027] (Model parameter learning unit 130) The model parameter learning unit 130 includes an extraction unit 131, an embedding unit 132, a connection unit 133, an estimation unit 134, and a calculation unit 135. The model parameter learning unit 130 has an internal memory for temporarily storing programs that define various processing procedures and the like and processing data, and is realized by electronic circuits such as a CPU (Central Processing Unit) or an MPU (Micro Processing Unit), or integrated circuits such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array).

[0028] (Extraction part 131) The extraction unit 131 extracts an encoding feature having a time series direction based on input data, which may be unimodal data, which is data of a single modality, or multimodal pair data, which includes multiple different modalities, or both. Furthermore, the model parameter learning unit 130 has multiple extraction units 131 as an extraction unit, and extracts a single encoding feature when the input data is single unimodal data, and extracts encoding features according to the number of modal types included in the input data when the input data is two or more unimodal data or multimodal pair data, or both. The extraction unit 131 can set operating conditions based on the number and type of input learning data.

[0029] Furthermore, the extraction unit 131 extracts encoding features based on any neural network corresponding to any modal type. For example, if the purpose is to extract image modal features, the neural network of the extraction unit 131 is configured with four convolutional layers deployed in the time series direction, and if the purpose is to extract audio modal features, it is configured with four RNN layers. Note that the neural network of the extraction unit 131 is not limited to the aforementioned four convolutional layers and four RNN layers, and other neural networks may be configured.

[0030] (embedded portion 132) The embedding unit 132 embeds segment information, which is information identifying the modality type of the input data, into the encoding feature based on a predetermined condition. Specifically, the embedding unit 132 embeds a vector, which has the same sequence length as the encoding feature to be input and includes a fixed value that differs for each modality, into the encoding feature.

[0031] The segment information is a vector having an arbitrary fixed value for each modality with the same sequence length as the input encoding feature (for example, if the encoding feature is an image modal, it is a vector of [1,1,1,1,1,1 → time sequence direction], and if it is an audio modal, it is a vector of [2,2,2,2,2,2 → time sequence direction]). The embedding unit 132 then assigns the segment information to the encoding feature by, for example, adding the segment information to the encoding feature.

[0032] (Connection part 133) The concatenation unit 133 concatenates a plurality of segment embedding features in the time series direction as modal concatenation features based on the input conditions of the segment embedding features in which segment information is embedded in the encoding features. Specifically, when a plurality of segment embedding features are input, the concatenation unit 133 concatenates the plurality of segment embedding features in the time series direction.

[0033] For example, the concatenation unit 133 receives both the segment embedding feature 1 and the segment embedding feature 2 as input, concatenates the segment embedding feature 1 and the segment embedding feature 2 in the time series direction, and outputs the result as a modal concatenation feature. On the other hand, when only one of the segment embedding feature 1 or the segment embedding feature 2 is received as input, the concatenation unit 133 outputs the segment embedding feature 1 or the segment embedding feature 2 after adding zero-filled portions without performing concatenation processing.

[0034] Furthermore, the concatenation of segment embedding features and the addition of zero-filling portions performed by the concatenation unit 133 will be described with reference to Fig. 3. Fig. 3 illustrates the processing of the concatenation unit 133 in each of the situations of "input of only segment embedding feature 1", "input of only segment embedding feature 2", and "input of both segment embedding features 1 and 2".

[0035] First, in the situations of "only segment embedding feature 1 input" and "only segment embedding feature 2 input," since there is only one segment embedding feature, the concatenation unit 133 does not perform concatenation processing and only adds zero-filling parts. Note that the concatenation unit 133 adds zero-filling parts in the time series direction so that all segment embedding features have the same sequence length.

[0036] On the other hand, in the case of "simultaneous input of segment embedding features 1 and 2," since there are multiple segment embedding features, the concatenation unit 133 concatenates segment embedding feature 1 and segment embedding feature 2 in the time series direction and adds zero-padding parts so that all segment embedding features have the same series length.

[0037] (Estimation part 134) Now, we return to the explanation of Fig. 2. The estimation unit 134 performs transformation using an arbitrary neural network function based on the segment embedding features and / or the modal coupling features, and estimates a vector corresponding to the ground truth data Y as an estimated vector for the cross-modal task.

[0038] The estimation unit 134 selects an arbitrary neural network suited to the type of modality. For example, for a classification task, the estimation unit 134 has four LSTM layers and two fully connected layers, and outputs a probability for each classification category. Note that the estimation unit 134 is not limited to vectors related to the classification categories described above, and can also estimate vectors related to tasks of other categories.

[0039] (Calculation unit 135) The calculation unit 135 performs iterative learning using the estimated vector of the cross-modal task estimated by the estimation unit 134 based on the segment embedding features and / or modal concatenation features and the ground truth data Y, and calculates the model parameter θ. Specifically, the calculation unit 135 measures the error between the estimated vector of the cross-modal task and the ground truth data Y for the model parameter θ, performs iterative learning, and updates the model parameter θ multiple times to minimize the error. For example, when the task is a classification task, the calculation unit 135 performs the update process of the model parameter θ using cross entropy or the like. Note that the calculation unit 135 is not limited to the model parameter update method described above, and can use other model parameter update methods for neural networks.

[0040] (Cross-modal task estimation unit 140) The cross-modal task estimation unit 140 receives as input the model parameter θ calculated by the calculation unit 135 and estimates an estimation vector Z for the inference input data 20. Note that the cross-modal task estimated by the cross-modal task estimation unit 140 is a "classification task that outputs a probability for each category" or a "regression task that outputs a vector," and the cross-modal task estimation unit 140 can arbitrarily set the task.

[0041] [3. Modifications] Next, a modified example of the learning method of the learning device 100 in this embodiment will be described with reference to Fig. 4. First, we will explain X, which is included in the unimodal data 1 (Equation (1)), unimodal data 2 (Equation (2)), and multimodal pair data 3 (Equation (3)), which are parameters used in this modified example.

[0042] In this modification, it is assumed that there are two types of unimodal data and one type of multimodal paired data, that is, the inputs are unimodal data 1, unimodal data 2, and multimodal paired data 3. The types of unimodal data and multimodal paired data are not limited to the above-mentioned numbers, and may be other numbers.

[0043] The X in the aforementioned unimodal data is data of a single modality, such as image data, audio data, or text data. On the other hand, the X in multimodal paired data is data that pairs two or more different modalities. For the sake of explanation, unimodal data 1 will be defined as "image data," unimodal data 2 as "audio data," and multimodal paired data 3 as "paired data of image data and audio data." Furthermore, the X in unimodal data 1 (Equation (1)), unimodal data 2 (Equation (2)), and multimodal paired data 3 (Equation (3)) will be expressed as "X of unimodal data 1 (Equation (4))," "X of unimodal data 2 (Equation (5))," and "X of multimodal paired data 3 (Equation (6))," respectively, using the following equations (4) to (6).

[0044]

number

number

number

[0045] The unimodal data 1 and unimodal data 2 may be unimodal data extracted from the multimodal paired data 3, or may be data not included in the multimodal paired data 3.

[0046] Next, we will explain Y, which is included in unimodal data 1 (Equation (1)), unimodal data 2 (Equation (2)), and multimodal pair data 3 (Equation (3)). Y is the correct answer data corresponding to any task (hereinafter referred to as "correct answer data Y"), and for example, in the case of a classification task, it is a category label, and in the case of a regression task, it is a correct answer value or vector. In an emotion classification task, classification categories such as "happiness," "anger," "sadness," and "neutral" are common to each data.

[0047] In this modification, it is assumed that unimodal data 1 (equation (1)), unimodal data 2 (equation (2)), and multimodal pair data 3 (equation (3)) are all data for the same task, and that the format of the correct answer data Y is unified. The correct answer data Y is then expressed by the following equations (7) to (9), respectively.

[0048]

number

number

number

[0049] Next, each functional unit of the model parameter learning unit 130 in this modified example will be described with reference to Fig. 4. The model parameter learning unit 130 in this modified example includes a plurality of extraction units 131: extraction unit 131A that receives X equation (4) of unimodal data 1, which is the image modal, and X equation (6) of multimodal paired data 3, and extraction unit 131B that receives X equation (5) of unimodal data 2, which is the audio modal, and X equation (6) of multimodal paired data 3. The number of extraction units 131 included in the model parameter learning unit 130 is not limited to the two mentioned above, and can be set arbitrarily depending on the number of modal types included in the unimodal data or multimodal paired data.

[0050] Then, using the above-mentioned learning data 10 as input, the extraction unit 131A extracts and outputs the encoding feature 1, and the extraction unit 131B extracts and outputs the encoding feature 2. Specifically, when the following equation (10) included in equation (4) for X of the unimodal data 1 is the sequential image data equation (10) of the unimodal data 1, and the following equation (11) included in equation (6) for X of the multimodal paired data 3 is the sequential image data equation (11) of the multimodal paired data 3, the extraction unit 131A extracts and outputs the encoding feature 1 based on the sequential image data equation (10) of the unimodal data 1 and the sequential image data equation (11) of the multimodal paired data 3. Note that in this modification, the extraction unit 131A can use the sequential image data equation (10) of the unimodal data 1 and the sequential image data equation (11) of the multimodal paired data 3 as input data in any order.

[0051]

number

number

[0052] On the other hand, when the following formula (12) included in formula (5) for X of the unimodal data 2 is formula (12) for the speech data of the unimodal data 2, and the following formula (13) included in formula (6) for X of the multimodal paired data 3 is formula (13) for the speech data of the multimodal paired data 3, the extraction unit 131B extracts and outputs encoding feature 2 based on formula (12) for the speech data of the unimodal data 2 and formula (13) for the speech data of the multimodal paired data 3. Note that in this modification, the extraction unit 131B can use formula (12) for the speech data of the unimodal data 2 and formula (13) for the speech data of the multimodal paired data 3 as input data in any order.

[0053]

number

number

[0054] Furthermore, extraction units 131A and 131B change their operations based on learning data 10 that is input data. Specifically, when only X equation (4) of unimodal data 1 is input, extraction unit 131B does not operate and only extraction unit 131A operates, and when X equation (4) of unimodal data 1 and X equation (5) of unimodal data 2 are input, respectively, extraction units 131A and 131B both operate. On the other hand, when X equation (6) of multimodal paired data 3 is input, extraction units 131A and 131B operate simultaneously.

[0055] Next, the embedding unit 132 takes as input both the encoding feature 1 output by the extraction unit 131A and the encoding feature 2 output by the extraction unit 131B, assigns segment information indicating which modal feature it is to each encoding feature, and outputs them as segment embedding feature 1 and segment embedding feature 2.

[0056] Furthermore, the embedding unit 132 changes its operation based on the encoding feature that is used as input data. Specifically, when the extraction unit 131A receives only Equation (4) of unimodal data 1 as input, the embedding unit 132 outputs segment embedding feature 1 based on encoding feature 1 output by the extraction unit 131A. On the other hand, when the extraction unit 131B receives only Equation (6) of unimodal data 2 as input, the embedding unit 132 outputs segment embedding feature 2 based on encoding feature 2 output by the extraction unit 131B.

[0057] Furthermore, when the extraction units 131A and 131B respectively input sequential image data of the multimodal pair data (Equation (11)) and audio data of the multimodal pair data 3 (Equation (13)), the embedding unit 132 inputs both the encoding feature 1 output by the extraction unit 131A and the encoding feature 2 output by the extraction unit 131B, and outputs both the segment embedding feature 1 and the segment embedding feature 2.

[0058] Next, the concatenation unit 133 concatenates and adds zero-filling parts in the time series direction for both the segment embedding feature 1 and the segment embedding feature 2 output by the embedding unit 132, and outputs them as a modal concatenation feature.

[0059] Next, the estimation unit 134 receives the above-mentioned modal link feature as input, transforms it using an arbitrary neural network function, estimates an estimated vector for the cross-modal task, which is a vector for an arbitrary task corresponding to the correct answer data Y, and outputs it.

[0060] Next, the calculation unit 135 receives the estimated vector of the cross-modal task as an input and calculates the model parameter θ. The correct data Y used by the calculation unit 135 is the above-mentioned formulas (7) to (9). The model parameter θ calculated by the calculation unit 135 may be a parameter corresponding to the three units, namely, the extraction unit 131A, the extraction unit 131B, and the estimation unit 134, or may be a parameter corresponding only to the estimation unit 134.

[0061] The cross-modal task estimation unit 140 then estimates the estimation vector Z based on the above-mentioned model parameter θ and the inference input data 20. Note that the cross-modal task estimation unit 140 can use either unimodal data s or multimodal pair data m as the inference input data 20. Specifically, the above-mentioned unimodal data s is data of a single modality, such as image data, audio data, or text data.

[0062] On the other hand, multimodal pair data m is data that pairs two or more modals different from multimodal pair data m, and is data that expresses one piece of data in multiple different modals, such as sequential image data and audio data extracted from one piece of video data.

[0063] [4. Processing Procedure] Next, the procedure of the learning method by the learning device 100 will be described with reference to Fig. 5. First, the model parameter learning unit 130 acquires multimodal pair data or unimodal data as learning data (step S11). Next, the extraction unit 131 extracts encoding features based on the multimodal pair data or unimodal data (step S12). Subsequently, the embedding unit 132 embeds segment information into the encoding features (step S13).

[0064] Next, the concatenation unit 133 determines that multiple segment embedding features have been input (Yes in step S14). In this case, the concatenation unit 133 concatenates the multiple segment embedding features in the time series direction (step S15). On the other hand, if the concatenation unit 133 determines that multiple segment embedding features have not been input, it proceeds to the next step without performing the concatenation process (No in step S14). Next, the concatenation unit 133 adds zero-filling portions to a single segment embedding feature or a modal concatenation feature obtained by concatenating multiple segment embedding features (step S16).

[0065] Next, the estimation unit 134 estimates a cross-modal task estimation vector for an arbitrary task corresponding to the supervised data Y based on the segment embedding feature or the modal concatenation feature (step S17). Subsequently, the calculation unit 135 repeatedly performs learning based on the cross-modal task estimation vector and the supervised data Y, updates the model parameter θ multiple times to minimize the error, and calculates the model parameter θ (step S18). Then, the cross-modal task estimation unit 140 estimates the cross-modal task estimation vector Z from the inference input data using the model parameter θ (step S19).

[0066] [5. Effects] As described above, the learning device 100 extracts encoding features with a time series direction based on input data of either or both unimodal data, which is data of a single modality, and multimodal pair data, which includes multiple different modalities; embeds segment information, which is information identifying the modality type of the input data, into the encoding features based on predetermined conditions; connects multiple segment embedding features in the time series direction as modal concatenation features based on the input conditions of the segment embedding features with the embedded segment information; and calculates model parameters using either or both of the segment embedding features and the modal concatenation features, and an estimated vector of the cross-modal task estimated based on the ground truth data. Therefore, this embodiment has the following advantages.

[0067] The learning device 100 provides the effect of enabling highly accurate estimation of cross-modal tasks optimized for both multimodal paired data and unimodal data.

[0068] Furthermore, the learning device 100 operates regardless of whether multimodal pair data or unimodal data is input, and provides the advantage of being able to estimate cross-modal tasks.

[0069] Furthermore, the learning device 100 provides the effect of being able to prevent a decrease in task accuracy when only limited data is available.

[0070] [6. Hardware Configuration] The components of each device shown in the figure are conceptual functional units and do not necessarily have to be physically configured as shown. In other words, the specific form of distribution and integration of each device is not limited to that shown, and all or part of each device can be functionally or physically distributed and integrated in any unit depending on various loads, usage conditions, etc. Furthermore, all or any part of the processing functions performed by each device can be realized by a CPU and a program analyzed and executed by the CPU, or can be realized as hardware using wired logic.

[0071] Furthermore, among the processes described in this embodiment, all or part of the processes described as being performed automatically can also be performed manually using known methods. In addition, the information including the processing procedures, control procedures, specific names, various data, and parameters shown in the drawings can be changed as desired unless otherwise specified.

[0072] [program] In one embodiment, study device 100 can be implemented by installing a study program that executes the aforementioned study as package software or online software on a desired computer. For example, by executing the study program on an information processing device, it can function as study device 100. The information processing device referred to here includes desktop and notebook personal computers. Other information processing devices also include mobile communication terminals such as smartphones, mobile phones, and PHS (Personal Handyphone Systems), as well as slate terminals such as PDAs (Personal Digital Assistants).

[0073] 6 is a diagram showing an example of a computer in which learning device 100 is realized. Computer 1000 includes, for example, memory 1010 and CPU 1020. Computer 1000 also includes hard disk drive interface 1030, disk drive interface 1040, serial port interface 1050, video adapter 1060, and network interface 1070. These components are connected by bus 1080.

[0074] The memory 1010 includes a ROM (Read Only Memory) 1011 and a RAM 1012. The ROM 1011 stores a boot program such as a BIOS (Basic Input Output System). The hard disk drive interface 1030 is connected to a hard disk drive 1090. The disk drive interface 1040 is connected to a disk drive 1100. A removable storage medium such as a magnetic disk or optical disk is inserted into the disk drive 1100. The serial port interface 1050 is connected to a mouse 1110 and a keyboard 1120, for example. The video adapter 1060 is connected to a display 1130, for example.

[0075] The hard disk drive 1090 stores, for example, an OS 1091, an application program 1092, a program module 1093, and program data 1094. That is, the program that defines each process of the learning device 100 is implemented as a program module 1093 in which computer-executable code is written. The program module 1093 is stored, for example, in the hard disk drive 1090. For example, a program module 1093 for executing processes similar to those of the functional configuration of the learning device 100 is stored in the hard disk drive 1090. Note that the hard disk drive 1090 may be replaced by an SSD (Solid State Drive).

[0076] Furthermore, setting data used in the processing of the above-described embodiment is stored as program data 1094, for example, in the memory 1010 or the hard disk drive 1090. Then, the CPU 1020 reads the program module 1093 or the program data 1094 stored in the memory 1010 or the hard disk drive 1090 into the RAM 1012 as necessary, and executes the processing of the above-described embodiment.

[0077] The program module 1093 and program data 1094 are not limited to being stored in the hard disk drive 1090, but may also be stored in, for example, a removable storage medium and read by the CPU 1020 via the disk drive 1100 or the like. Alternatively, the program module 1093 and program data 1094 may be stored in another computer connected via a network (such as a LAN or a WAN (Wide Area Network)). The program module 1093 and program data 1094 may then be read by the CPU 1020 from the other computer via the network interface 1070.

[0078] [7. Other] Although the present embodiment has been described above, the present embodiment is not limited by the description and drawings that form part of the disclosure. In other words, other embodiments, examples, operational techniques, etc. made by those skilled in the art based on the present embodiment are all included in the scope of the present embodiment. [Explanation of symbols]

[0079] 10 Training data 20 Input data for inference 100 Learning Device 110 Communications Department 120 Storage section 121 Modal Data Storage Unit 122 Encoding feature memory 123 Segment Embedding Feature Memory 124 Modal connection feature memory 125 Model parameter storage unit 126 Estimated Vector Memory Unit 130 Model parameter learning unit 131 Extraction part 131A Extraction part 131B Extraction part 132 Embedded part 133 Connecting part 134 Estimation Department 135 Calculation Unit 140 Cross-modal task estimation unit 1000 computers 1010 memory 1011 ROM 1012 RAM 1020 CPU 1030 hard disk drive interface 1040 disk drive interface 1050 serial port interface 1060 video adapter 1070 Network Interface 1080 Bus 1090 hard disk drive 1091 OS 1092 Application Program 1093 Program Module 1094 Program Data 1100 disk drive 1110 Mouse 1120 keyboard

Claims

1. an extraction unit that extracts encoded features having a time series direction based on input data of either or both of unimodal data, which is data of a single modality, and multimodal pair data including a plurality of different modalities; an embedding unit that embeds segment information, which is information that identifies the modal type of the input data, into the encoded features based on a predetermined condition; a concatenation unit that concatenates a plurality of the segment embedding features in a time series direction as modal concatenation features based on an input condition of the segment embedding features in which the segment information is embedded; a calculation unit that calculates model parameters using the segment embedding features or the modal coupling features, or both, and an estimated vector of a cross-modal task estimated based on ground truth data; A learning device comprising:

2. The extraction unit includes a plurality of extraction units, extracting a single encoding feature when the input data is a single unimodal data; extracting the encoding features according to the number of modal types included in the input data when the input data includes two or more of the unimodal data and / or the multimodal pair data; 2. The learning device according to claim 1 .

3. the extraction unit extracts the encoding features based on a neural network corresponding to the modal type.

3. The learning device according to claim 2.

4. the embedding unit embeds, into the encoded features, a vector having the same sequence length as the encoded features to be input, the vector including fixed values ​​that differ for each modal.

2. The learning device according to claim 1 .

5. the concatenation unit concatenates the plurality of segment embedding features in a time series direction when the plurality of segment embedding features are input; 2. The learning device according to claim 1 .

6. an estimation unit that performs transformation using a function of an arbitrary neural network based on both or either one of the segment embedding features and the modal coupling features, and estimates a vector corresponding to the ground truth data as an estimated vector for the cross-modal task; 2. The learning device according to claim 1 .

7. extracting encoded features having a time series direction based on input data of either or both of unimodal data, which is data of a single modality, and multimodal pair data including a plurality of different modalities; embedding segment information, which is information for identifying the modal type of the input data, into the encoded features based on a predetermined condition; a step of linking the plurality of segment embedding features in a time series direction as modal linking features based on an input condition of the segment embedding features into which the segment information is embedded; calculating model parameters using the segment embedding features and / or the modal coupling features and an estimated vector of a cross-modal task estimated based on ground truth data; A learning method comprising:

8. extracting encoded features having a time series direction based on input data of either or both of unimodal data, which is data of a single modality, and multimodal pair data including a plurality of different modalities; embedding segment information, which is information for identifying the modal type of the input data, into the encoded features based on a predetermined condition; a step of linking the plurality of segment embedding features in a time series direction as modal linking features based on an input condition of the segment embedding features into which the segment information is embedded; Calculating model parameters using the segment embedding features or the modal coupling features, or both, and an estimated vector of a cross-modal task estimated based on ground truth data; A learning program characterized by comprising:

Citation Information

Patent Citations

  • Multi-modal fusion emotion recognition system and method based on multi-task learning and attention mechanism and experimental evaluation method

    CN113420807A