Information processing device, information processing method, and program
Patent Information
- Application Number
- PCT/JP2026/012110
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-25
- Filing Date
- 2026-03-25
- Publication Date
- 2026-10-01
Smart Images

Figure JP2026012110_01102026_PF_FP_ABST
Abstract
Description
Information Processing Apparatus, Information Processing Method, and Program
[0001] The present invention relates to an information processing apparatus, an information processing method, and a program.
[0002] Deep learning models for handling a plurality of modalities such as text and image are known. For example, Patent Document 1 discloses a technique related to control of an interactive agent that uses multimodal input.
[0003] Japanese National Publication of International Patent Application No.2024-540848
[0004] In multimodal learning, a large amount of text and images are used for training a deep learning model, but a difference occurs between the number of texts and the number of images, so it is desired to stabilize the training of the model.
[0005] An object of the present invention is to provide an information processing apparatus, an information processing method, and a program capable of suppressing imbalance in paired data used for training a deep learning model.
[0006] In order to solve the above-mentioned problems and achieve the object, an information processing apparatus according to the present invention is an information processing apparatus that inputs learning data including first data and second data related to the first data to perform machine learning of a plurality of experts, the information processing apparatus comprising: an acquisition unit that acquires paired data of a first image and a first caption; an extraction unit that extracts features of a second image obtained by dividing the first image; a collection unit that collects a second caption indicating the features of the second image; and a selection unit that sends the learning data, in which pairs of the first caption, the second caption, the first image, and the second image are used as the first data and the second data, to the selected expert.
[0007] To solve the above-mentioned problems and achieve the objective, the present invention provides an information processing method for an information processing device that performs machine learning on multiple experts by inputting training data having first data and second data related to the first data, and includes: an acquisition step of acquiring pair data of a first image and a first caption; an extraction step of extracting features of a second image obtained by dividing the first image; a collection step of collecting a second caption that shows the features of the second image; and a selection step of sending the training data, in which pairs of the first caption and the second caption and the first image and the second image are used as the first data and the second data, to the selected experts.
[0008] To solve the above-mentioned problems and achieve the objective, the program according to the present invention causes an information processing device that performs machine learning on multiple experts by inputting training data having first data and second data related to the first data to execute an acquisition step of acquiring pair data of a first image and a first caption; an extraction step of extracting features of a second image obtained by dividing the first image; a collection step of collecting a second caption that shows the features of the second image; and a selection step of sending the training data, in which pairs of the first caption and the second caption and the first image and the second image are used as the first data and second data, to the selected experts.
[0009] According to the present invention, it is possible to suppress imbalances in paired data used for training deep learning models.
[0010] Figure 1 is a schematic diagram showing the system overview of the information processing device according to the embodiment. Figure 2 is a configuration diagram showing an example of the system configuration of the information processing device according to the embodiment. Figure 3 is a diagram for explaining an example of the configuration of the information processing device according to the embodiment. Figure 4 is a flowchart showing a part of the processing procedure of the information processing device according to the embodiment. Figure 5 is a flowchart showing another part of the processing procedure of the information processing device according to the embodiment. Figure 6 is the data flow of the information processing device according to the embodiment.
[0011] Hereinafter, embodiments of the present invention will be described in detail with reference to the attached drawings. Note that the present invention is not limited to these embodiments, and if there are multiple embodiments, they may be constructed by combining each embodiment. Furthermore, in the following embodiments, the same parts are denoted by the same reference numerals to omit redundant explanations.
[0012] (LIMoE) LIMoE (Language-Image Mixture Experts) is a deep learning model for handling multiple modalities such as text and images. This model effectively extracts and learns information from data by utilizing a set of expert networks, each with the best expertise for a particular task. LIMoE has the ability to capture the relationships between different types of data (multimodal) and combine them to generate new knowledge. Multimodal means that multiple modalities such as visual, auditory, and linguistic information are used.
[0013] Multimodal settings present various challenges, one of which is modality imbalance. In real-world environments, one data type may be more prevalent than another. For example, in LIMoE experiments, the number of image tokens was 3 to 17 times greater than the number of text tokens. This embodiment provides a technique to suppress the imbalance of paired data used for training deep learning models.
[0014] [Embodiment] An information processing device that performs image classification by multiple experts according to an embodiment will be described using Figures 1 and 2. Figure 1 is a schematic diagram showing the system overview of the information processing device according to the embodiment. Figure 2 is a configuration diagram showing an example of the system configuration of the information processing device according to the embodiment.
[0015] As shown in Figures 1 and 2, the information processing device 10 is an image classification device that classifies images using multiple experts 200 such as LIMoE. The multiple experts 200 are deep learning models that possess knowledge in specialized fields. In the example shown in Figure 1, the multiple experts 200 are information analyzers with different areas of expertise. For example, the multiple experts 200 are information analyzers specializing in animals (faces), animals (whole bodies), animals, humans, etc. In this embodiment, comparative learning by LIMoE is performed by machine learning the experts 200 using paired data of images and text that describe those images. The purpose of comparative learning is to encourage the feature representations of image and text pairs that show the same subject to be similar, and the feature representations of image and text pairs that show different subjects to be far apart.
[0016] The information processing device 10 performs training (machine learning) on multiple experts 200 by inputting training data 300, which consists of pairs of first data 310 and second data 320 related to the first data 310, into the experts 200. The first data 310 is, for example, text data. The second data 320 is, for example, image data showing an image corresponding to the first data 310.
[0017] The information processing device 10 acquires pair data D10 of multiple first images D11 and first captions D12 using the acquisition unit 30. The first image D11 is image data of an object. The first caption D12 is text data indicating the name of the object shown in the first image D11, and is the correct data for the first image D11. In the example shown in Figure 1, multiple first images D11 are image data showing images corresponding to one first caption D12, "panda". The information processing device 10 acquires the pair data D10, which is large-scale data, via, for example, an input device, a communication device, etc., and outputs it to the extraction unit 32.
[0018] The information processing device 10 includes an extraction unit 32 that extracts features of a second image D13 obtained by dividing the first image D11. The extraction unit 32 includes an image divider 32a, a text divider 32b, an image feature extractor 32c, and a text encoder 32d.
[0019] The image segmenter 32a divides the input first image D11 into second image D13, which are small image fragments called patches, and outputs them to the image feature extractor 32c. For example, if the first image D11 is divided into nine patches, there will be nine second image D13. The text segmenter 32b divides the input first caption D12 into tokens D14 and outputs them to the text encoder 32d. Tokens D14 are basic units that make up text, such as words, subwords, and characters. The image feature extractor 32c extracts features of the second image D13 using image recognition software, machine learning models, etc., and outputs the extraction results to the collection unit 34. Features of the second image D13 include, for example, the image recognition result, the shape of the image, and feature vectors of objects that the image represents. The text encoder 32d extracts features of multiple tokens D14 using analysis software, machine learning models, etc., and outputs the extraction results to the collection unit 34. The features of token D14 are, for example, feature vectors corresponding to words. The features of token D14 may be determined by considering the position of the word in the text and its relationship to other words.
[0020] The information processing device 10 includes a collection unit 34 that collects a second caption D15 that describes the features of the second image D13. The collection unit 34 collects the second caption D15 corresponding to the features of the patch shown in the input second image D13 from a database 400, such as a knowledge base or knowledge graph. For example, a knowledge graph has a graph structure in which nodes representing images and their associated text are connected by edges that show the relationships between the nodes. This knowledge graph also collects highly reliable information from dictionaries edited by humans. Therefore, by using the knowledge graph database 400, the collection unit 34 can collect a second caption D15 suitable for the features of multiple patches. The collection unit 34 can collect a second caption D15 suitable for the features of multiple patches from a database 400, etc., that is suitable for the specialized fields of multiple experts 200.
[0021] In one example shown in Figure 1, the information processing device 10 collects second captions D15 for sentences and words such as "a panda is running" and "a sleeping panda," and generates new captions to create image-caption pairs. This allows the information processing device 10 to process unbalanced first images D11 and first captions D12 and increase the number of second captions D15 associated with each patch. The collection unit 34 outputs the collected first image D11, second image D13, first caption D12, and second caption D15 along with the input data to the adjustment unit 35.
[0022] The information processing device 10 includes an adjustment unit 35 that calculates the correlation between tokens, the correlation between second images D13 (patches), and the correlation between token D14 and second image D13, and adjusts the characteristics of token D14 and patches. The adjustment unit 35 adjusts the first image D11 and second image D13 and the first caption D12 and second caption D15 so that at least one duplicate data remains. The adjustment unit 35 outputs the adjusted first image D11, first caption D12, second image D13, token D14, and second caption D15 to the selection unit 36. The adjustment unit 35 may be included in the acquisition unit 34 or in the selection unit 36.
[0023] The information processing device 10 includes a selection unit 36 that sends training data 300, which consists of pairs of first caption D12 and second caption D15 and first image D11 and second image D13, as first data 310 and second data 320, to a selected expert 200. The selection unit 36 selects an expert 200 to whom the training data 300 will be sent based on the characteristics of the second caption D15 (patch) and token D14. The selection unit 36 then sends the training data 300 to the selected expert 200.
[0024] When the expert 200 receives the learning data 300 sent by the selection unit 36, it analyzes the characteristics of the learning data 300. In this embodiment, the selection unit 36 and the expert 200 are configured as a group, and the output of the expert 200 is input to the next selection unit 36 as input to the next group.
[0025] The output unit 38 outputs pairs of objects and their names, which are the result of the expert 200 analyzing the input image. The output unit 38 outputs the pairs of objects and their names to, for example, a communication device, a display device, etc.
[0026] (Information Processing Device) An example of the configuration of an information processing device according to the embodiment will be explained using Figure 3. Figure 3 is a diagram for explaining an example of the configuration of an information processing device according to the embodiment.
[0027] As shown in Figure 3, the information processing device 10 comprises a communication unit 20, a storage unit 22, and a control unit 24. The information processing device 10 can be implemented as, for example, a server device or an agent device.
[0028] The communication unit 20 is a communication interface that performs communication between the information processing device 10 and an external device. For example, the communication unit 20 performs communication between the information processing device 10 and an external device.
[0029] The memory unit 22 stores various types of information. The memory unit 22 stores the calculation contents of the control unit 24 and information such as programs. The memory unit 22 includes at least one of the following: a main memory device such as RAM (Random Access Memory) and ROM (Read Only Memory), or an external memory device such as an HDD (Hard Disk Drive).
[0030] The storage unit 22 can store data such as program D1, paired data D10, expert 200, and database 400. Program D1 includes a program that causes the control unit 24 of the information processing device 10 to execute information processing methods, etc. Expert 200 is data for a deep learning model. The storage unit 22 can store multiple paired data D10 used for training expert 200. Note that the storage unit 22 may be configured to store expert 200 and database 400 in an external storage device or the like from the information processing device 10.
[0031] The control unit 24 controls each part of the information processing device 10. The control unit 24 includes, for example, an information processing device such as a CPU (Central Processing Unit) or an MPU (Micro Processing Unit), and a storage device such as a RAM (Random Access Memory) or a ROM (Read Only Memory). The control unit 24 executes a program that controls the operation of the information processing device 10 according to the present invention. The control unit 24 may be implemented by an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array). The control unit 24 may be implemented by a combination of hardware and software.
[0032] The control unit 24 includes the acquisition unit 30, extraction unit 32, collection unit 34, selection unit 36, and output unit 38 described above.
[0033] The acquisition unit 30 acquires pair data D10 of multiple first images D11 and first captions D12 and stores them in the storage unit 22. The extraction unit 32 extracts features of the second image D13 obtained by dividing the first image D11. The collection unit 34 collects second captions D15 that show the features of the second image D13. The selection unit 36 sends the training data 300, which consists of pairs of first captions D12 and second captions D15 and first images D11 and second images D13 as first data 310 and second data 320, to the selected expert 200.
[0034] (Processing Procedure of the Information Processing Device) Figure 4 is a flowchart showing a part of the processing procedure of the information processing device according to the embodiment. Figure 5 is a flowchart showing another part of the processing procedure of the information processing device according to the embodiment. The processing procedures shown in Figures 4 and 5 are realized when the control unit 24 executes program D1.
[0035] As shown in Figure 4, the control unit 24 of the information processing device 10 acquires paired data D10 (step S100). The control unit 24 then divides the text (first caption D12) into a plurality of tokens D14 (step S101) and extracts the features of the tokens D14 (step S102). At the same time, the control unit 24 divides the image (first image D11) into a plurality of patches (second image D13) (step S103) and extracts the features of the patches (second image D13) (step S104). In this embodiment, the processing procedure shown in Figure 4 is described as being performed in parallel with steps S101 and S102 and steps S103 and S104, but for example, the processing may be performed in the order of steps S101 to S104.
[0036] The control unit 24 collects a second caption D15 based on token D14 (step S105). Specifically, the control unit 24 compares token D14 with the database 400, calculates the similarity to the features of captions in the database 400, and collects a second caption D15 such as synonyms or similar sentences from the database 400. Then, the control unit 24 collects a second caption D15 based on a patch (second image D13) (step S106). Specifically, the control unit 24 compares the patch (second image D13) with the database 400, calculates the similarity to the features of images in the database 400, and collects a second caption D15 such as words or sentences related to similar images (step S106).
[0037] Next, the control unit 24 deletes any duplicate second captions D15 (step S107). Specifically, the control unit 24 checks whether the second captions D15 collected in steps S105 and S106 are duplicates, and if they are, it deletes all but one of the duplicates.
[0038] Next, the control unit 24 extracts the features of the remaining second caption D15 (step S108). Then, the control unit 24 combines the features of the patch (second image D13) with the features of the second caption D15 to create pairs (step S109). Specifically, the control unit 24 compares the features of the patch (second image D13) in step S104 with the features of the second caption D15 in step S108 and creates pairs based on similar features.
[0039] Next, the control unit 24 calculates the correlation between the patch (second image D13) and the second caption D15 (step S110). Then, based on the correlation, the control unit 24 adjusts the features of the patch (second image D13) and the second caption D15 (step S111). Specifically, the control unit 24 adjusts the features of the patch and the second caption D15 so that they differ from other pairs.
[0040] Next, as shown in Figure 5, the control unit 24 selects and sends an expert 200 based on the characteristics of the patch (second image D13) and the second caption D15 (step S112). Specifically, the control unit 24 sends the patch (second image D13) and the second caption D15 along with the training data 300 to the selected expert 200. Then, the control unit 24 has the expert 200 analyze the characteristics of the sent patch (second image D13) and the second caption D15 (step S113).
[0041] Next, when the processing in step S113 is completed, the control unit 24 determines whether or not there are other experts 200 (step S114). In this embodiment, since there are N groups of selection units 36 and experts 200, the features to be analyzed are input to the next selection unit 36 and the features are analyzed by other experts 200. If the control unit 24 determines that there are other experts 200 (step S114; Yes), it returns to step S112, which has already been described, and continues processing. If the control unit 24 determines that there are no other experts 200 (step S114; No), it determines that the analysis by all experts 200 has been completed, and proceeds to step S115.
[0042] The control unit 24 calculates an average vector of the second caption D15 and the patch (second image D13) (step S115). In the present embodiment, the control unit 24 uses contrast learning of the LIMoE expert 200, and thus adjusts features of text and images to the same dimension before calculating a loss. For this reason, the control unit 24 calculates an average vector (average value) of the representative second caption D15 and the patch.
[0043] Next, the control unit 24 calculates a distance K1 between the second caption D15 and the patch (second image D13), and adjusts the distance K1 with a hyperparameter (step S116). A hyperparameter is a parameter that controls the behavior of a machine learning algorithm, and includes, for example, the number of epochs, learning rate, threshold, mini-batch size, the number of layers, and the number of neurons per layer.
[0044] Next, the control unit 24 calculates the distance between the second caption D15 in step S116 and all patches (second image D13), obtains a total distance K2, and adjusts the total distance K2 with a hyperparameter (step S117). Then, the control unit 24 calculates the distance between the patch in step S116 and all the second captions D15, obtains a total distance K3, and adjusts the total distance K3 with a hyperparameter (step S118).
[0045] Next, the control unit 24 calculates a loss based on the calculation results of steps S116 to S118 (step S119). Specifically, the control unit 24 calculates the loss as a sum of a quotient of the distance K1 and the total K2 and a quotient of the distance K1 and the total K3. Then, the control unit 24 obtains an average of "Importance Loss", "Load Loss", "Z Loss" and "The Mutual-information Loss" based on a calculation formula of a loss function for contrast learning (step S120).
[0046] Next, the control unit 24 calculates the sum of the contrast learning loss and the calculation formula (step S121). Then, the control unit 24 converges the contrast learning loss (step S122). Specifically, in order to converge the feature analysis by the expert 200, the control unit 24 performs training by repeating the processes from step S100 to step S113. When the process of step S122 is completed, the control unit 24 terminates the processing procedure shown in FIG. 4 and FIG. 5.
[0047] (Data Flow of Information Processing Apparatus) FIG. 6 is a data flow of the information processing apparatus according to the embodiment. In step P11 of FIG. 6, the information processing apparatus 10 extracts multi-modality patches (second image D13) from the first image D11, and extracts few-modality tokens D14 from the first caption D12. Then, in step P12, the information processing apparatus 10 pairs the patch (second image D13) with the token D14.
[0048] Next, in step P13, the information processing apparatus 10 collects a plurality of words corresponding to the patch (second image D13) from the database 400, and collects a plurality of synonyms corresponding to the token D14 from the database 400. Then, in step P14, the information processing apparatus 10 uses the collected plurality of words and the plurality of synonyms as a new token D14. Then, in step P15, the information processing apparatus 10 pairs the collected plurality of tokens D14 with the patch (second image D13). Accordingly, the information processing apparatus 10 can increase the number of few-modality tokens D14 based on the plurality of patches (second image D13).
[0049] In this way, the information processing device 10 can collect tokens D14 that represent the features of the patch (second image D13) as second captions D15 by utilizing the database 400. Furthermore, since the information processing device 10 collects tokens D14 that represent the features of the patch (second image D13), the reliability of the second captions D15 can be improved. As a result, the information processing device 10 can suppress imbalances in the paired data D10 used for training the expert 200 (deep learning model). In addition, since the information processing device 10 can collect second captions D15 even when there are few first captions D12 for multiple first images D11, it can prepare training data 300 suitable for training the deep learning model.
[0050] [Other Embodiments] Next, other embodiments will be described. The information processing device 10 has been described in the case where it classifies images using the LIMoE expert 200, but it is not limited to this. For example, the information processing device 10 can be applied to all multimodal related deep learning models.
[0051] In this embodiment, the information processing device 10 has been described in which paired data D10 of a first image D11 and a first caption D12 is used, but it is not limited to this. For example, the first image D11 may be video, audio, or other data, and may be used as paired data D10 with the first caption D12.
[0052] Although embodiments of the present invention have been described above, the present invention is not limited by the content of these embodiments. Furthermore, the aforementioned components include those that can be easily conceived by those skilled in the art, those that are substantially the same, and those that fall within the so-called equivalent range. Moreover, the aforementioned components can be combined as appropriate. Furthermore, various omissions, substitutions, or modifications of the components can be made without departing from the gist of the embodiments described above.
[0053] 10 Information Processing Device 20 Communication Unit 22 Storage Unit 24 Control Unit 30 Acquisition Unit 32 Extraction Unit 32a Image Segmenter 32b Text Segmenter 32c Image Feature Extractor 32d Text Encoder 34 Collection Unit 35 Adjustment Unit 36 Selection Unit 38 Output Unit 200 Expert 300 Training Data 310 First Data 320 Second Data 400 Database D1 Program D10 Pair Data D11 First Image D12 First Caption D13 Second Image D14 Token D15 Second Caption
Claims
1. An information processing device that performs machine learning on multiple experts by inputting training data having first data and second data related to the first data, comprising: an acquisition unit that acquires pair data of a first image and a first caption; an extraction unit that extracts features of a second image obtained by dividing the first image; a collection unit that collects a second caption that shows the features of the second image; and a selection unit that sends the training data, in which pairs of the first caption and the second caption and the first image and the second image are used as the first data and second data, to the selected experts.
2. The information processing apparatus according to claim 1, wherein the collection unit collects second captions suitable for the characteristics of a plurality of experts from a database, and the selection unit selects the experts to whom the learning data is sent based on the characteristics of the first caption and the second caption and the first image and the second image.
3. The information processing apparatus according to claim 2, wherein the extraction unit collects the second caption which is similar to the characteristics of the token obtained by dividing the first caption.
4. An information processing method for an information processing device that performs machine learning on multiple experts by inputting training data having first data and second data related to the first data, the information processing method comprising: an acquisition step of acquiring pair data of a first image and a first caption; an extraction step of extracting features of a second image obtained by dividing the first image; a collection step of collecting a second caption that shows the features of the second image; and a selection step of sending the training data, in which pairs of the first caption and the second caption and the first image and the second image are used as the first data and second data, to the selected experts.
5. A program that causes an information processing device, which takes training data having first data and second data related to the first data as input to perform machine learning on multiple experts, to execute the following steps: an acquisition step of acquiring pair data of a first image and a first caption; an extraction step of extracting features of a second image obtained by dividing the first image; a collection step of collecting a second caption that shows the features of the second image; and a selection step of sending the training data, in which pairs of the first caption and the second caption and the first image and the second image are used as the first data and second data, to the selected experts.