Training method, training program, and information processing device

WO2026203070A1PCT designated stage Publication Date: 2026-10-01NT T INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/011870
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2026-10-01

Smart Images

  • Figure JP2025011870_01102026_PF_FP_ABST
    Figure JP2025011870_01102026_PF_FP_ABST
Patent Text Reader

Abstract

A training method executed by an information processing device (10) and comprising the steps of: calculating the contrast loss of individual sample sets a first type of data and a second type of data on the basis of classification results obtained by classifying each sample set into positive or negative examples according to the degree of similarity of the sample set; and updating, on the basis of the contrast loss, a parameter of a multimodal AI for processing the first type of input data and the second type of input data.
Need to check novelty before this filing date? Find Prior Art

Description

Training method, training program, and information processing device

[0001] The present invention relates to a training method, a training program, and an information processing device.

[0002] One example of multimodal AI (Artificial Intelligence) is CLIP (Contrastive Language-Image Pretraining), which integrates and processes data from two modal elements: language and vision.

[0003] Alec Radford et al., “Learning Transferable Visual Models From Natural Language Supervision”, Alec Radford+, arxiv, 2021

[0004] However, multimodal AI, such as CLIP mentioned above, has room for improvement in that paired data with pair labels that map between modalities are essential when fine-tuning the model. This fine-tuning includes additional training and tuning for specific domains. In other words, it may also include modifying the model for data specific to a particular domain or task.

[0005] Therefore, the present invention aims to provide a training method, a training program, and an information processing device that can realize the training of a multimodal AI even when there is insufficient paired data.

[0006] To solve the above-mentioned problems and achieve the objective, the training method of the present invention is a training method performed on an information processing device, which includes the following steps: for each sample set of first type data and second type data, calculate a comparison loss for each sample set based on a classification result in which the sample set is classified as a positive example according to the similarity of the sample sets; and update the parameters of a multimodal AI that processes the first type and the second type input data based on the comparison loss.

[0007] According to the present invention, training of multimodal AI can be achieved even when there is insufficient paired data.

[0008] Figure 1 is a block diagram showing an example of the functional configuration of an information processing device. Figure 2 is a diagram showing an example of paired data and unimodal data. Figure 3 is a diagram (1) showing an example of comparative learning. Figure 4 is a diagram (2) showing an example of comparative learning. Figure 5 is a diagram showing one aspect of the problem. Figure 6 is a diagram (1) showing one aspect of the problem-solving approach. Figure 7 is a diagram (2) showing one aspect of the problem-solving approach. Figure 8 is a schematic diagram explaining the overall picture of the processing of each functional unit. Figure 9 is a flowchart showing the procedure for the setup process. Figure 10 is a flowchart (1) showing the procedure for the training process. Figure 11 is a flowchart (2) showing the procedure for the training process. Figure 12 is a schematic diagram explaining application example 1. Figure 13 is a schematic diagram explaining application example 2. Figure 14 is a schematic diagram explaining application example 3. Figure 15 is a diagram showing an example of the hardware configuration.

[0009] The following description will explain the forms for implementing the training method, training program, and information processing device related to this disclosure (hereinafter referred to as "Embodiments") with reference to the attached drawings. It should be noted that these embodiments represent only one example or one aspect, and the following description does not limit the structure, operation, function, properties, characteristics, methods, and applications related to this disclosure.

[0010] <Overall Configuration> Figure 1 is a block diagram showing an example of the functional configuration of the information processing device 10. Figure 1 shows an information processing device 10 that provides a training function to update the parameters of a multimodal AI by calculating the comparison loss for each pair based on the classification results of positive examples corresponding to the similarity of pairs of unimodal data.

[0011] The information processing device 10 can provide the above-mentioned training functions as a cloud service by running a PaaS (Platform as a Service) type middleware or a SaaS (Software as a Service) type application.

[0012] As shown in Figure 1, the information processing device 10 can be connected to a client terminal 30 via a network NW in a communicative manner. For example, the network NW may be any type of communication network, such as the Internet or a LAN (Local Area Network), whether wired or wireless. Although Figure 1 shows an example where one client terminal 30 is connected to one information processing device 10, this does not prevent any number of client terminals 30 from being connected.

[0013] The client terminal 30 is a terminal device that receives the training functions described above. For example, the client terminal 30 can be used by anyone involved in multimodal AI, such as those involved in the design, development, operation, or maintenance of the system. The client terminal 30 may be implemented using any computer, including, for example, a personal computer, a smartphone, a tablet, or a wearable device.

[0014] While the above training function is described as being provided as a cloud service, it is not limited to this. For example, the above training function may be provided on-premises. Also, while the above training function is described as being provided as a client-server system, it is not limited to this. For example, the above training function may be provided as a standalone service by having an application running on the client terminal 30 execute processing corresponding to the above training function on the client terminal 30.

[0015] <Multimodal AI> Among multimodal AIs, there is a type called a foundational model. Hereafter, the multimodal foundational model may be referred to as the "MM (Multi Modal) foundational model".

[0016] Examples of such memory-memory (MM) based models include CLIP, AudioCLIP, and ImageBind. For example, CLIP is an MM based model that integrates and processes data from two types of modals: language and visual. AudioCLIP is an MM based model that integrates and processes data from three types of modals: language, speech, and visual. ImageBin is an MM based model that integrates and processes data from three or more types of modals, such as language, speech, and visual.

[0017] These memory modeling (MM)-based models are trained using paired data with pair labels that map between modalities. Figure 2 shows an example of paired data and unimodal data. As shown in Figure 2, paired data is MM data that has temporal or conceptual relationships with each other across multiple modalities, and may include data such as images, audio, and text contained in a video. On the other hand, unimodal data is individual modal data such as text data, audio data, and image data. For example, personal data could include posts and diaries on social networking services (SNS), radio broadcasts, and image folders on mobile devices. For example, organizational data could include images from factory cameras, recordings, meeting minutes, security camera images, and internal company documents. In such unimodal data, if pair labels are not set, the relationships between multiple modalities are unknown.

[0018] One aspect is that the absolute number of paired data is significantly smaller than the absolute number of unimodal data where the relationships between multiple modalities are unknown. Another aspect is that preparing paired data is also very costly. For example, the absolute number of video data is small compared to the absolute number of individual unimodal data such as text data, audio data, and image data, and recording them is also time-consuming. Furthermore, when generating paired data from text data, audio data, or image data where the relationships between multiple modalities are unknown, it is time-consuming to check whether multiple modalities have temporal or conceptually corresponding relationships with each other and to set pair labels, and data that cannot be paired is wasted. In addition, from the perspective of privacy and security, there may be gaps in the paired data for specific modalities.

[0019] <Contrastive Learning> Contrastive learning using paired data is performed on multimodal AI. Figures 3 and 4 are (1) and (2) which show examples of contrastive learning. Figure 3 schematically illustrates how contrastive learning is performed on CLIP, while Figure 4 schematically illustrates how contrastive learning is performed on AudioCLIP.

[0020] As shown in Figure 3, in the case of CLIP, image and text pair data is used as training data. For such training data, a dataset known as WIT (WebImageText) is used, which consists of pairs of images and text written as captions for those images extracted from web pages on the internet.

[0021] In such paired data, the image data is input to an image encoder, while the text data is input to a text encoder. The image encoder, upon receiving the image data, outputs a vector that embeds the image into the feature space. On the other hand, the text encoder, upon receiving the text data, outputs a vector that embeds the text into the feature space.

[0022] For example, Figure 3 shows Image I 1 and text T 1 The pair, Image I 2and text T 2 pair, ..., image I N and text T N A batch with batch size N including N pieces of pair data is illustrated as an example. In this case, by inputting each of the N pieces of pair data into the image encoder and the text encoder, a similarity matrix M1 of N×N embedding vectors can be obtained. It should be noted that the "similarity" mentioned herein is merely an example of "relationship". For example, the similarity may be indexed by the inner product between embedding vectors, cosine similarity, or the like, or may be indexed by distance or the like. In other words, the similarity matrix M1 may be generated as an example of the relevance matrix of N×N sample sets.

[0023] Here, in the training of CLIP, an objective function including contrastive loss is used. For example, in the example of the batch shown in FIG. 3, for the i-th image in the batch, the i-th text corresponds to the correct pair, so while the i-th text is taken as a positive example, all texts including the non-paired texts are taken as negative examples. That is, since one positive example and N negative examples are set for each piece of training data, there are N positive examples and N 2 negative examples generated in the entire batch. For example, in the example of the similarity matrix M1, the elements of the N diagonal components displayed in black and white inversion are taken as positive examples, and N 2 elements displayed as white background and black background are taken as negative examples.

[0024] Under such a similarity matrix M1, while maximizing the similarity of N pairs corresponding to positive examples, N 2 parameters of the image encoder and the text encoder that minimize the similarity of sample sets are trained.

[0025] For example, a first image I 1 in the example, according to the following formula (1), a first text T 1 is taken as a positive example, T 1 to T N all texts are taken as negative examples, and contrastive loss, such as cross-entropy error, is calculated in the row direction of the similarity matrix M1. By executing such calculation of row-direction contrastive loss for each of the N images, the contrastive loss L related to the imageI→T is obtained. On the other hand, the first text T 1 As an example, according to the following formula (2), the first image I 1 is set as a positive example, I 1 to I N all images of are set as negative examples, and contrastive loss is calculated in the column direction of the similarity matrix M1. By performing such calculation of column-direction contrastive loss for each of N texts, the contrastive loss related to texts T→I is obtained. In this way, for each row direction and column direction of the similarity matrix M1, a contrastive loss that brings features of positive examples closer while moving features of negative examples further apart is calculated. Then, according to the following formula (3), the contrastive loss T of CLIP CLIP , for example, parameter updating for minimizing the average value of the loss related to images and the loss related to texts is executed on the image encoder and the text encoder.

[0026]

[0027] As shown in FIG. 4, in the case of AudioCLIP, paired data of text, audio and image is used as training data. Here, even in the case of AudioCLIP, for the similarity matrix M1 of text data and image data, the same calculation method as the contrastive loss using the similarity matrix M1 shown in FIG. 3 is executed. Thereby, the contrastive loss T of AudioCLIP shown in the following formula (4) AudioCLIP of which the loss term L I→T and the loss term L T→I is obtained. Furthermore, for the similarity matrix M2 of text data and audio data and the similarity matrix M3 of image data and audio data, even if the modalities are swapped, the same calculation method as the contrastive loss using the similarity matrix M1 shown in FIG. 3 can be executed. Thereby, for the contrastive loss T of AudioCLIP shown in the following formula (4) AudioCLIP of which the loss term L T→A , the loss term L A→T , the loss term L I→A and the loss term L A→IThis is then calculated. The parameters that minimize the average of these six loss terms are updated in the text encoder, image encoder, and audio encoder.

[0028]

[0029] <One aspect of the challenge> As explained in the background technology section above, multimodal AI such as CLIP requires paired data with pair labels that map between modalities when fine-tuning the model, and there is room for improvement in this regard.

[0030] The fine-tuning referred to here includes additional training and tuning for specific domains. In other words, it may also include modifying the model for data specific to a particular domain or task.

[0031] Figure 5 shows one aspect of the problem. In Figure 5, the amount of unimodal data and multimodal data available for each phase of the learning process, in which three phases are performed on the base model—pre-training, additional training to train on specific domains, and task training to train on specific tasks—is schematicly represented in proportion to the amount of data in each phase.

[0032] As shown in Figure 5, in multimodal learning, it is clear that paired data with pair labels that establish a one-to-one correspondence between modalities is essential in all phases: pre-training, additional training, and task training. Due to this constraint on paired data, multimodal learning inevitably has the drawback of having a limited amount of usable data.

[0033] <One aspect of the problem-solving approach> Therefore, in this embodiment, a training function is provided that calculates the control loss for each sample set based on the classification result of positive examples corresponding to the relationship between the sample sets of the first type of data and the second type of data, and updates the parameters of the multimodal AI.

[0034] The following example uses similarity as an indicator to represent "relationship," but other indicators such as distance may also be used. Similarly, modality is given as an example of "type," but "type" may include varieties, categories, classes, etc.

[0035] Figures 6 and 7 are (1) and (2) illustrating one aspect of the problem-solving approach. Figures 6 and 7 show similarity matrices M for two types of modalities: the first modality and the second modality. Furthermore, in Figures 6 and 7, the similarity of the elements of the sample sets included in the similarity matrix M is indicated by the intensity of the colors. Darker colors indicate higher similarity, while lighter colors indicate lower similarity.

[0036] As merely an example, in the training function according to this embodiment, elements of sample sets included in the similarity matrix M whose similarity is equal to or greater than a threshold can be classified as positive examples. On the other hand, as merely an example, all elements of sample sets may be used as negative examples. This does not prevent the classification of elements of sample sets whose similarity is not equal to or greater than a threshold as negative examples. For example, Figure 6 shows an example where the similarity corresponding to the second highest hatching among the seven types of hatching shown in the legend is set as the threshold. In this case, as shown in Figure 6, the elements of the sample sets indicated by the thick-lined boxes in the similarity matrix M, namely the elements in row 1, column 1, row 1, column 3, row 2, column 2, row 2, column 5, row 4, column 4, row 5, column 2, row 5, column 8, row 7, column 6, and row 7, column 8, are classified as positive examples. On the other hand, all elements of sample sets are used as negative examples. Therefore, even if the sample sets in the similarity matrix M do not necessarily have pair labels set, the row and column-wise comparison losses of the similarity matrix M can be calculated.

[0037] As another example, in the training function according to this embodiment, elements of the sample sets included in the similarity matrix M that correspond to the top predetermined number of similarities are classified as positive examples. On the other hand, for negative examples, all elements of the sample sets may be used as examples. For example, in Figure 7, the elements of the sample sets in the similarity matrix M that correspond to the top predetermined number of similarities, for example, the top five, are numbered according to the rank of the similarity. In this case, as shown in Figure 7, the elements of the sample sets indicated by the thick-lined boxes in the similarity matrix M, namely the elements in row 1, column 1, row 2, column 2, row 4, column 4, row 5, column 8 and row 7, column 6, are classified as positive examples. On the other hand, all elements of the sample sets are used as negative examples. Therefore, even if the sample sets in the similarity matrix M do not necessarily have pair labels set, the row and column-wise similarity loss of the similarity matrix M can be calculated.

[0038] Therefore, according to the training function of this embodiment, training of a multimodal AI can be achieved even when there is insufficient paired data.

[0039] <Configuration of Information Processing Device 10> Next, the functional configuration of the information processing device 10 that provides the above-mentioned training function will be described. Figure 1 schematically shows the blocks related to the training function of the information processing device 10. As shown in Figure 1, the information processing device 10 has a communication control unit 11, a storage unit 13, and a control unit 15. Note that Figure 1 only shows a selection of the functional units related to the above-mentioned training function, and the information processing device 10 may also be equipped with functional units other than those shown.

[0040] The communication control unit 11 is a functional unit that controls communication with other devices such as the client terminal 30. In one embodiment, the communication control unit 11 can be implemented by a network interface card such as a LAN card. In one aspect, the communication control unit 11 receives training requests from the client terminal 30 that request the execution of additional training of the MM base model, or outputs a response to the training request, such as the trained MM base model, to the client terminal 30.

[0041] The storage unit 13 is a functional unit that stores various types of data. In one embodiment, the storage unit 13 may be implemented by internal, external, or auxiliary storage of the information processing device 10. For example, the storage unit 13 stores a first data set 13A, a second data set 13B, and base model data 13C. However, this does not prevent other data other than the first data set 13A, the second data set 13B, and the base model data 13C from being stored in the storage unit 13.

[0042] The first dataset 13A is a set of data from the first modality. The second dataset 13B is a set of data from the second modality. These first and second modalities may be different modalities and may be any modality, including text, audio, images, and sensor information. Furthermore, some of the samples in the first dataset 13A and the second dataset 13B may include paired data with paired labels. Moreover, this data may be transformed. For example, it may be a feature vector (feature quantity) output by inputting data into a base model.

[0043] The base model data 13C is data relating to the MM base model. For example, the base model data 13C may include hyperparameters related to the layer structure, such as neurons and synapses in the input, hidden, and output layers that form the pre-trained MM base model, as well as parameters related to the objective function, such as weights and biases of each layer. For example, a pre-trained MM base model can be obtained by controlled learning as explained using Figures 3 and 4. Furthermore, the base model data 13C may also store additional learning parameters added to the task network or base model, such as LoRA (Low-Rank Adaptation). Note that the base model data 13C does not necessarily have to be stored in the memory unit 13; a pre-trained MM base model can also be obtained from a library published on the network NW.

[0044] The control unit 15 is a functional unit that performs overall control of the information processing device 10. For example, the control unit 15 can be implemented by a hardware processor. As shown in Figure 1, the control unit 15 has a setting unit 15A, a generation unit 15B, a classification unit 15C, a calculation unit 15D, and an update unit 15E. The control unit 15 may also be implemented by hardwired logic or the like.

[0045] Before providing a detailed explanation of each of these functional units—the setting unit 15A, the generation unit 15B, the classification unit 15C, the calculation unit 15D, and the update unit 15E—we will explain the overall process performed by each functional unit.

[0046] Figure 8 is a schematic diagram illustrating the overall processing of each functional unit. Figure 8 shows an example where additional training of a specific domain is performed on a pre-trained MM base model. Furthermore, Figure 8 shows text data as an example of a sample from the first dataset, and image data as an example of a sample from the second dataset.

[0047] As shown in Figure 8, during additional training of the MM-based model, a batch is generated in which sample sets of text and image data with paired labels are mixed with sample sets of text data without paired labels and image data without paired labels. In this case, the number of samples between the two modalities, text and image, does not necessarily have to be the same.

[0048] In this way, samples of text data from the batch generated are input to the text encoder E11 of the MM-based model undergoing additional training, while samples of image data are input to the image encoder E12 of the MM-based model undergoing training. The parameters of these text encoder E11 and image encoder E12 may be updated during the training process when additional training is performed. For example, they may be updated epoch by epoch, or they may be updated at times other than epochs. The text encoder E11, which receives the text data samples, outputs a feature vector in which the text data is embedded in the feature space. On the other hand, the image encoder E12, which receives the image data, outputs a feature vector in which the image data is embedded in the feature space. Then, for each sample set in which text data samples and image data samples are combined within a batch, the similarity of the feature vectors of that sample set is calculated as the first similarity.

[0049] On the other hand, among the samples included in the batch, the text data samples are input to the pre-trained MM-based model's text encoder E01, while the image data samples are input to the pre-trained MM-based model's image encoder E02. The parameters of these text encoder E01 and image encoder E02 may not be updated with each additional training epoch, but may be frozen after pre-training. Thus, the text encoder E01, which receives the text data samples, outputs a feature vector in which the text data is embedded in the feature space. On the other hand, the image encoder E02, which receives the image data, outputs a feature vector in which the image data is embedded in the feature space. Then, for each sample pair in which text data samples and image data samples are combined within the batch, the similarity of the feature vectors of that sample pair is calculated as a second similarity.

[0050] Here, the similarity used to calculate the control loss within a batch can be calculated by combining the first similarity and the second similarity. As just one example of such similarity, the sum of the first similarity, weighted (1-μ), and the second similarity, weighted μ, can be calculated as shown in equation (5) below.

[0051]

[0052] Here, the threshold α used for comparison with similarity can, as an example, be set to the similarity of the top k% of all pairs obtained by combining all samples from the first and second datasets that do not have pair labels.

[0053] Then, among the elements of the sample sets included in the similarity matrix M generated based on the above similarity, the elements of the sample sets for which pair labels have been set are classified as positive examples. Furthermore, the elements of the sample sets whose similarity is greater than or equal to the threshold α are classified as pseudo-positive examples. On the other hand, all sample sets are used as negative examples. Then, according to equation (6) below, the parameters of the text encoder and image encoder are updated to maximize the similarity of the sample sets corresponding to positive examples and the similarity of the sample sets corresponding to pseudo-positive examples, while minimizing the similarity of the sample sets corresponding to negative examples. At this time, as shown in equation (6) below, the similarity of the sample sets corresponding to positive examples is assigned a weight λ, while the similarity of the sample sets corresponding to pseudo-positive examples may be assigned a weight (1-λ). Note that if the proportion of positive examples and pseudo-positive examples in a batch is less than the threshold β, the parameter update for that batch can be skipped.

[0054]

[0055] Returning to the explanation of Figure 1, the setting unit 15A has the function of setting a threshold α to compare with the similarity of the sample set. In one embodiment, if the first dataset 13A and the second dataset 13B have already been registered in the storage unit 13, the setting unit 15A can pre-set the threshold α before the above training request is received.

[0056] More specifically, the configuration unit 15A generates B batches ' from the datasets corresponding to each of the first and second modalities, specifically from samples that do not have pair labels assigned. The "batches '" mentioned here can correspond to the batch size of the batches generated during additional training. Subsequently, the configuration unit 15A performs the following processing for each of the B batches '. Specifically, for each M samples of the first modality contained in the b-th batch ', the configuration unit 15A inputs the sample into the encoder corresponding to the first modality in the pre-trained MM base model. This yields a second feature vector for each of the M samples of the first modality. In parallel, the configuration unit 15A inputs each of the N samples of the second modality contained in the b-th batch ' into the encoder corresponding to the second modality in the pre-trained MM base model. This yields a second feature vector for each of the N samples of the second modality. Then, the setting unit 15A calculates the similarity of the sample set corresponding to each element of the E sample sets included in the M x N similarity matrix, for example, the dot product of the second feature vector pair or the cosine similarity. This gives the similarity for each element of the E sample sets. After that, the setting unit 15A sorts the similarity of the elements of the E sample sets obtained for each B batch' in descending order. Then, the setting unit 15A sets the similarity of the top k% point among the sorted sample sets as the threshold α. Here, as an example, we have given an example of generating B batch' from samples without pair labels, but it goes without saying that it is also possible to generate batches from all samples without pair labels and determine the threshold α.

[0057] The generation unit 15B has the function of generating batches. In one embodiment, the generation unit 15B can start processing when it receives a training request from the client terminal 30 requesting the execution of additional training of the MM-based model. For example, the generation unit 15B repeats the following process until it reaches a predetermined number of epochs P. That is, the generation unit 15B obtains a subset of paired data samples with pair labels set from the datasets corresponding to each modality of the first dataset 13A and the second dataset 13B. In parallel with this, the generation unit 15B randomly generates a subset of samples without pair labels set from the datasets corresponding to each modality of the first dataset 13A and the second dataset 13B. Then, the generation unit 15B generates B batches by randomly combining the subsets with pair labels set and the subsets without pair labels set.

[0058] The classification unit 15C has the function of classifying a set of samples from the first modality and a set of samples from the second modality into positive examples. In one embodiment, the classification unit 15C performs the following processing for each B batch. That is, for each M samples of the first modality contained in the b-th batch, the classification unit 15C inputs the sample into the encoder corresponding to the first modality in the MM-based model' that is being further trained. This obtains a first feature vector for each M samples of the first modality. In parallel with this, the classification unit 15C inputs each N samples of the second modality contained in the b-th batch into the encoder corresponding to the second modality in the MM-based model' that is being further trained. This obtains a first feature vector for each N samples of the second modality. Subsequently, the classification unit 15C calculates weights μ to be used for weighting the first and second similarities based on the learning progress. For example, the classification unit 15C can set the initial value of weight μ to "1" and approach "0" as the learning progresses. The parameter representing the learning progress can, for example, be the number of epochs. In this case, the weight μ can be calculated according to equation (7) below. Note that the parameter representing the learning progress is not limited to the number of epochs; it can also be the rate of decrease of the gradient or the rate of decrease of the loss. For example, the weight μ can be brought closer to "1" as the rate of decrease of the gradient or the rate of decrease of the loss increases, and closer to "0" as the rate of decrease of the gradient or the rate of decrease of the loss decreases.

[0059]

[0060] Subsequently, the classification unit 15C calculates a first similarity, such as cosine similarity, from a pair of first feature vectors corresponding to each element of the E sample sets included in the M x N similarity matrix. This provides a first similarity for each of the E sample sets. Furthermore, the classification unit 15C calculates a second similarity, such as cosine similarity, from a pair of second feature vectors corresponding to each element of the E sample sets included in the M x N similarity matrix. The pairs of second feature vectors can be obtained by inputting each of the M samples of the first modality and each of the N samples of the second modality into the encoders of the pre-trained MM base model. Here, since the parameters of the pre-trained MM base model encoders are frozen, the calculation results of the pairs of second feature vectors calculated when the threshold α is set by the setting unit 15A may be referenced. This provides a second similarity for each of the E sample sets.

[0061] Then, the classification unit 15C assigns a first similarity (v) weight (1-μ) to each sample set corresponding to the E elements according to the following formula (8). k ・s b ) and a second similarity (~v) with weight μ assigned to it. k ・~s b The sum of the values ​​from ) and the similarity sim(v) used to calculate the control loss. k ,s b It can be calculated as follows:

[0062]

[0063] Furthermore, the classification unit 15C calculates the similarity sim(v k ・s b Based on the above, the classification unit 15C classifies sample sets corresponding to E elements in the similarity matrix M into positive examples b++ if a pair label is set. The classification unit 15C also classifies sample sets of elements with a similarity of α or higher into pseudo-positive examples b+. All sample sets are used as negative examples.

[0064] The calculation unit 15D has the function of calculating the control loss for the sample set in the batch. In one embodiment, if the ratio of the sum of positive examples and pseudo-positive examples in the batch is greater than or equal to a threshold β, the calculation unit 15D calculates the control loss for each row and column of the M×N similarity matrix. As just one example, the calculation unit 15D calculates the control loss L in the row direction of the M×N similarity matrix according to the following formula (9). I→T The calculation unit 15D can calculate the following: For example, the calculation unit 15D maximizes the similarity of the sample set corresponding to the positive example b++ and the similarity of the sample set corresponding to the pseudo-positive example b+, while minimizing the similarity of the sample set corresponding to the negative example L I→T The following is calculated. At this time, as shown in equation (9) below, a weight λ is assigned to the similarity of the sample sets corresponding to positive examples, while a weight (1-λ) is assigned to the similarity of the sample sets corresponding to pseudo-positive examples. Note that the weight λ may be a hyperparameter that adjusts the weighting of positive examples with pair labels set and pseudo-positive examples classified by threshold α. Then, the calculation unit 15D calculates the row-direction comparison loss L of the M × N similarity matrix according to equation (10) below. I→T and the column-direction contrast loss L of the M×N similarity matrix T→I The control loss L for the entire batch is calculated from this.

[0065]

[0066] The update unit 15E has the function of updating the parameters of the MM base model' during additional training. In one embodiment, the update unit 15E updates the parameters of each encoder of the MM base model' during additional training based on the control loss L of the entire batch calculated by the calculation unit 15D. If the ratio of the sum of positive examples and pseudo-positive examples in a batch is less than the threshold β, the parameter update for that batch is skipped. Here, an example of performing additional training on a pre-trained MM base model has been given, but task training can also be performed on a pre-trained MM base model or an MM base model that has undergone additional training. For example, if a task network, such as a fully connected layer, is added after the encoders of each modality of the MM base model, the parameters of the task network can be updated. Also, if the task network is added inside the encoder, the encoder parameters can be updated.

[0067] <Processing Flow> Next, the processing flow of the information processing device 10 according to this embodiment will be explained. Here, we will first explain (1) the setup process performed by the information processing device 10, and then explain (2) the training process.

[0068] (1) Setting Process Diagram 9 is a flowchart showing the procedure for the setting process. This process is merely an example, and if the first dataset 13A and the second dataset 13B have already been registered in the storage unit 13, it can be started at any time, even before the above training request is received. It should be noted that the flowchart described below is only one example of how to set the threshold α, and it goes without saying that the threshold α can be set by other methods.

[0069] As shown in Figure 9, the setting unit 15A generates B batches' from the datasets corresponding to each of the first and second modalities, from which no pair labels have been set (step S101).

[0070] Subsequently, the setting unit 15A executes a loop process 1, which repeats the processes of step S102 and step S103 below a number of times corresponding to the total number B of batch'.

[0071] In other words, the setting unit 15A executes a loop process 2, which repeats the process in step S102A below a number of times corresponding to the total number M of samples of the first modality included in the b-th batch'. Specifically, the setting unit 15A inputs the m-th sample to the encoder corresponding to the first modality in the pre-trained MM base model (step S102A). As this loop process 2 is repeated, a second feature vector is obtained for each M samples of the first modality.

[0072] In parallel with the above loop process 2, the setting unit 15A executes loop process 3, which repeats the process in step S102B below a number of times corresponding to the total number N of samples of the second modality included in the b-th batch'. That is, the setting unit 15A inputs the nth sample to the encoder corresponding to the second modality in the pre-trained MM base model (step S102B). As this loop process 3 is repeated, a second feature vector is obtained for each N samples of the second modality.

[0073] Then, the setting unit 15A executes a loop process 4, which repeats the process in step S103 below a number of times corresponding to the total number of elements E in the sample set included in the M x N similarity matrix. That is, the setting unit 15A calculates the similarity of the sample set corresponding to the e-th element, for example, the dot product or cosine similarity of the pair of second feature vectors (step S103). By repeating this loop process 4, similarity is obtained for each element of the E sample sets.

[0074] Subsequently, the setting unit 15A sorts the elements of the E sample sets obtained for each B batch' by repeating the above loop process 1 in descending order (step S104). Then, the setting unit 15A sets the similarity of the top k% points among the similarity of the sample sets sorted in descending order in step S104 as a threshold α (step S105), and terminates the process.

[0075] (2) Training process Figures 10 and 11 are flowcharts (1) and (2) showing the procedure for the training process. This process can be initiated as an example when a training request is received from the client terminal 30 requesting the execution of additional training of the MM base model.

[0076] As shown in Figure 10, loop process 1 is executed, which repeats the processes from step S301 to step S315 below until a predetermined number of epochs P is reached.

[0077] In other words, the generation unit 15B obtains a subset of paired data samples with paired labels from the datasets corresponding to each modality of the first dataset 13A and the second dataset 13B (step S301A).

[0078] In parallel with this, the generation unit 15B randomly generates subsets of samples without pair labels from the datasets corresponding to each modality of the first dataset 13A and the second dataset 13B, from which no pair labels have been set (step S301B).

[0079] Then, the generation unit 15B generates B batches by randomly combining the subset with pair labels set obtained in step S301A and the subset without pair labels set obtained in step S301B (step S302).

[0080] Subsequently, loop process 2 is executed, repeating the processes from step S303 to step S315 below a number of times corresponding to the total number of batches B.

[0081] First, the classification unit 15C executes a loop process 3, which repeats the process in step S303A a number of times corresponding to the total number M of samples of the first modality included in the b-th batch.

[0082] In other words, the classification unit 15C inputs the m-th sample to the encoder corresponding to the first modality in the MM-based model' that is being further trained (step S303A). As this loop process 3 is repeated, a first feature vector is obtained for each M samples of the first modality.

[0083] In parallel with the above loop process 3, the classification unit 15C executes a loop process 4 that repeats the process of step S303B a number of times corresponding to the total number N of samples of the second modality included in the b-th batch.

[0084] In other words, the classification unit 15C inputs the nth sample to the encoder corresponding to the second modality in the MM-based model' that is being further trained (step S303B). As this loop process 4 is repeated, a first feature vector is obtained for each N samples of the second modality.

[0085] Subsequently, the classification unit 15C calculates weights μ to be used for weighting the first and second similarities based on the learning progress (step S303C).

[0086] Then, as shown in Figure 11, the classification unit 15C executes a loop process 5 that repeats the processes from step 304 to step S315 below a number of times corresponding to the total number E of elements in the sample set included in the M x N similarity matrix.

[0087] In other words, the classification unit 15C calculates a first similarity, such as cosine similarity, from the first pair of feature vectors corresponding to the elements of the e-th sample set (step S304).

[0088] Furthermore, the classification unit 15C calculates a second similarity, such as cosine similarity, from a second pair of feature vectors corresponding to the elements of the e-th sample set (step S305). The second pair of feature vectors may refer to the calculation result of the second pair of feature vectors calculated in the setting process shown in Figure 9.

[0089] Then, the classification unit 15C assigns a weight (1-μ) to the first similarity calculated in step S304, and assigns a weight μ to the second similarity calculated in step S305, and adds them together to calculate the similarity used for calculating the control loss (step S306).

[0090] If a pair label is set for the sample set corresponding to the e-th element (step S307: Yes), the classification unit 15C classifies the sample set corresponding to the e-th element as a positive example (step S308).

[0091] Furthermore, if no pair label is set for the sample set corresponding to the e-th element, but the similarity of the sample set corresponding to the e-th element is greater than or equal to a threshold α (step S307: No and step S309: Yes), the classification unit 15C classifies the sample set corresponding to the e-th element as a pseudo-positive example (step S310).

[0092] Furthermore, if no pair labels are set for the sample set corresponding to the e-th element, and the similarity of the sample set corresponding to the e-th element is not greater than or equal to the threshold α (step S307: No and step S309: No), the classification unit 15C does not classify the sample set corresponding to the e-th element as either a positive example or a pseudo-positive example.

[0093] As this loop process 5 is repeated, the sample sets corresponding to E elements will be classified as either positive or pseudo-positive examples. Note that all sample sets can also be used as negative examples.

[0094] Subsequently, if the ratio of the sum of positive and pseudo-positive examples in the b-th batch is greater than or equal to the threshold β (step S312: Yes), the calculation unit 15D calculates the control loss for the b-th batch (step S313). Then, the update unit 15E updates the parameters of each encoder in the MM-based model' undergoing additional training based on the control loss L for the entire batch calculated in step S313 (step S314).

[0095] Furthermore, if the ratio of the sum of positive and pseudo-positive values ​​in batch b is not greater than or equal to the threshold β (step S312: No), the parameter update in batch b is skipped (step S315).

[0096] As the above loop process 2 is repeated, the parameters of each encoder in the MM base model' being further trained are trained for every B batches. Furthermore, as the above loop process 1 is repeated, the further trained MM base model can be obtained.

[0097] <Summary> As described above, the information processing device 10 according to this embodiment calculates the comparison loss for each sample set based on the classification result of positive examples corresponding to the similarity of the sample sets of unimodal data, and updates the parameters of the multimodal AI. Therefore, according to the information processing device 10 according to this embodiment, training of a multimodal AI can be achieved even when there is insufficient paired data.

[0098] <Application Example 1> In the above embodiment 1, we gave an example of a first dataset and a second dataset which are a mixture of paired data with paired labels and unimodal data without paired labels. However, it is not necessary for the first dataset and the second dataset to contain paired data.

[0099] Figure 12 is a schematic diagram illustrating Application Example 1. Figure 12 shows an example in which additional training of a specific domain is performed on a pre-trained MM-based model. Furthermore, Figure 12 shows text data as an example of a sample from the first dataset, and image data as an example of a sample from the second dataset. As shown in Figure 12, in Application Example 1, the same processing as shown in Figure 8 is performed, except that paired data is not included. For example, one difference in processing content between Embodiment 1 and Application Example 2 is that the term corresponding to the positive example can be omitted from the formula for calculating the batch control loss. That is, in Application Example 1, the batch control loss can be calculated according to the following formula (11) instead of the above formula (9).

[0100]

[0101] <Application Example 2> In Embodiment 1 described above, an example was given in which additional learning and task learning of the MM base model are achieved by full fine tuning. However, additional learning and task learning of the MM base model may also be achieved by updating the parameters of an additional network added to the MM base model. Such an additional network may include an adapter module, LoRA, prompt, etc. Figure 13 is a schematic diagram illustrating Application Example 2. As shown in Figure 13, by updating the parameters of the additional networks A11 and A12 while keeping the parameters of the encoders E11 and E12 of the MM base model fixed, the difference in the parameters of the MM base model achieved by fine tuning is trained in the additional networks A11 and A12. Additional learning and task learning of the MM base model can also be achieved by training such an additional network.

[0102] <Application Example 3> In Embodiment 1 described above, an MM base model corresponding to two types of modalities, a first modality and a second modality, was given as an example. However, the above training function can be similarly applied to MM base models corresponding to three or more types of modalities. That is, even in MM base models such as AudioCLIP and ImageBind that integrate and process data from three types of modals, if we focus on each combination of modalities, processing for two types of modalities, the first modality and the second modality, is performed. Therefore, the processing shown in Figures 9 to 11 should be performed for each combination of modalities. Figure 14 is a schematic diagram illustrating Application Example 3. As shown in Figure 14, when updating the parameters of an MM base model that integrates and processes data from three types of modalities, language, speech, and vision, the parameters of the text encoder E11, image encoder E13, and speech encoder E13 should be updated according to the following equation (12).

[0103]

[0104] <Exhibition of Creative Ability> The matters described in this embodiment, such as the MM base model and specific examples like equations (5) to (12) above, are merely examples and can be modified. Furthermore, the flowchart described in this embodiment can also be modified within a consistent range, by changing the order of processing or skipping some processes.

[0105] <System> The processing procedures, control procedures, specific names, and information including various data and parameters shown in the above document and drawings can be changed at will unless otherwise specified. For example, one or more of the function units of the information processing device 10, such as the setting unit 15A, generation unit 15B, classification unit 15C, calculation unit 15D, and update unit 15E, may be configured as separate devices.

[0106] Furthermore, the components of each illustrated device are functionally conceptual and do not necessarily need to be physically configured as shown. In other words, the specific forms of distribution and integration of each device are not limited to those shown. That is, all or part of them can be functionally or physically distributed and integrated in any units according to various loads and usage conditions. Note that each configuration may also be a physical configuration.

[0107] Furthermore, the processing performed by the illustrated apparatus can be implemented, in whole or in part, by a program executed by a hardware processor such as an MPU (Micro-Processing Unit) or CPU (Central Processing Unit), or by hardware using wired logic.

[0108] <Hardware> Next, an example of the hardware configuration of the information processing device 10 described in this embodiment will be explained. For example, it can be implemented by installing a program that realizes the functions of the information processing device 10 on a computer. For example, by having the computer run the above program, which is provided as packaged software or online software, the computer can be made to function as the information processing device 10. The computer referred to here includes desktop or notebook personal computers, rack-mounted server computers, etc. In addition, the computer category also includes smartphones, mobile phones and PHS (Personal Handyphone System) and other mobile communication terminals, as well as PDAs (Personal Digital Assistants). Furthermore, the functions of the information processing device 10 may be implemented on a cloud server.

[0109] An example of a computer that executes the above program (training program) will be explained using Figure 15. As shown in Figure 15, the computer 1000 has, for example, a memory 1010, a CPU 1020, a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070. These components are connected by a bus 1080.

[0110] Memory 1010 includes ROM (Read Only Memory) 1011 and RAM (Random Access Memory) 1012. ROM 1011 stores, for example, a boot program such as BIOS (Basic Input Output System). The hard disk drive interface 1030 is connected to the hard disk drive 1090. The disk drive interface 1040 is connected to the disk drive 1100. The disk drive 1100 is used to insert a removable storage medium, such as a magnetic disk or an optical disk. The serial port interface 1050 is used to connect, for example, a mouse 1110 and a keyboard 1120. The video adapter 1060 is used to connect, for example, a display 1130.

[0111] Here, as shown in Figure 15, the hard disk drive 1090 stores, for example, the OS 1091, the application program 1092, the program module 1093, and the program data 1094. The storage unit 13 described in the above embodiment is equipped, for example, in the hard disk drive 1090 or the memory 1010.

[0112] Then, the CPU 1020 reads the program module 1093 and program data 1094 stored in the hard disk drive 1090 into the RAM 1012 as needed and executes the above-described procedures.

[0113] Furthermore, the program module 1093 and program data 1094 related to the above-mentioned information output program are not limited to being stored in the hard disk drive 1090, but may also be stored in a removable storage medium and read by the CPU 1020 via a disk drive 1100 or the like. Alternatively, the program module 1093 and program data 1094 related to the above-mentioned program may be stored in another computer connected via a network such as a LAN or WAN (Wide Area Network) and read by the CPU 1020 via a network interface 1070.

[0114] 10 Information processing device 11 Communication control unit 13 Storage unit 13A First dataset 13B Second dataset 13C Base model data 15 Control unit 15A Setting unit 15B Generation unit 15C Classification unit 15D Calculation unit 15E Update unit 30 Client terminal

Claims

1. A training method performed on an information processing device, comprising the process of: calculating a control loss for each sample set of data of a first type and data of a second type based on a classification result in which the sample set is classified as a positive example according to the relationship between the sample sets; and updating the parameters of a multimodal AI that processes the input data of the first type and the second type based on the control loss.

2. The training method according to claim 1, characterized in that the calculation process involves maximizing the similarity of a set of positive examples whose similarity representing the relationship is equal to or greater than a threshold, thereby calculating the control loss.

3. The training method according to claim 1, characterized in that the multimodal AI is a base model.

4. The training method according to claim 3, characterized in that the relationship is calculated by combining a first similarity obtained between the output of the first encoder when the first type of data is input to the first encoder of the base model in the training process by updating the parameters and the output of the second encoder when the second type of data is input to the second encoder of the base model in the training process and a second similarity obtained between the output of the first encoder when the first type of data is input to the first encoder of the base model and the output of the second encoder when the second type of data is input to the second encoder of the base model.

5. The training method according to claim 1, further comprising the process of generating a batch in which a subset of a first dataset, which is a set of the first type of data, and a subset of a second dataset, which is a set of the second type of data, wherein the calculation process calculates a control loss for each batch of sample sets in the batch, and the updating process excludes from updating the parameters the control losses of batches in which the proportion of sample sets of positive examples whose similarity representing the relationship is equal to or greater than a first threshold is less than a second threshold.

6. The training method according to claim 5, characterized in that the first threshold is set based on a similarity distribution representing the relationship between the sample pairs of data included in the first dataset and the data included in the second dataset.

7. A training program for a computer to perform a process which calculates a control loss for each pair of data of a first type and data of a second type based on the classification result in which the pair is classified as a positive example according to the similarity of the pair, and updates the parameters of a multimodal AI that processes the input data of the first and second types based on the control loss.

8. An information processing apparatus comprising: a calculation unit that calculates a comparison loss for each sample set of data of a first type and data of a second type based on a classification result in which the sample set is classified as a positive example according to the similarity of the sample sets; and an update unit that updates the parameters of a multimodal AI that processes the input data of the first type and the second type based on the comparison loss.