Apparatus and computer implementation method for training multimodal base models
Patent Information
- Application Number
- JP2026030020
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-02-27
- Filing Date
- 2026-02-26
- Publication Date
- 2026-09-08
Smart Images

Figure 2026143378000001_ABST
Abstract
Description
[Technical Field]
[0001] background The present invention relates to an apparatus and computer implementation method for training multimodal base models. [Background technology]
[0002] The multimodal platform model integrates multiple modalities into a single embedding space.
[0003] An example of a multimodal infrastructure model is "ImageBind: One Embedding Space To Bind Them All (arXiv:2305.05665)". [Prior art documents] [Non-patent literature]
[0004] [Non-Patent Document 1] Girdhar, Rohit; El-Nouby, Alaaeldin; Liu, Zhuang; Singh, Mannat; Vasudev Alwala, Kalyan; Joulin, Armand; Misra, Ishan, "ImageBind: One Embedding Space To Bind Them All(arXiv:2305.05665)" [Overview of the project] [Means for solving the problem]
[0005] Disclosure of the invention A computer implementation method for training a multimodal foundational model comprises providing a set of inputs for a first modality and a set of inputs for a second modality, wherein the set of inputs for the first modality includes a sequence of inputs for the first modality representing the motion of an object in the real world, and the set of inputs for the second modality includes a sequence of inputs for the second modality representing the same motion of an object in the real world; determining the output of a multimodal foundational model for inputs to a multimodal foundational model that include different combinations of one input from the set of inputs for the first modality and one input from the set of inputs for the second modality; and training the multimodal foundational model on a loss that includes a loss term specifically weighted for each of the outputs. For the two modalities, the loss is a pairwise loss. Such a pairwise loss may include a contrast loss as a CLIP loss and a SigLIP loss.
[0006] CLIP loss is described, for example, in "Learning Transferable Visual Models From Natural Language Supervision (arXiv:2103.00020v1)".
[0007] Regarding SigLIP loss, see, for example, "Sigmoid Loss for Language Image Pre-Training (arXiv:2303.15343v4)".
[0008] Training requires training data, especially high-quality data across multiple modalities.
[0009] For modalities where data acquired in the real world is not available, providing a set of inputs for a second modality may include determining the set of inputs for the second modality through synthetic data generation.
[0010] Determining the set of inputs for a second modality by generating synthetic data may include determining the inputs for the second modality by, for example, using generative artificial intelligence, depending on the inputs for the first modality, or depending on acquired data representing the motion of an object in the real world, or determining the set of inputs for the second modality in a simulation of the motion or movement of an object, or determining the set of inputs based on rules, or determining the set of inputs by combining these approaches.
[0011] Each input to the first modality may include a different digital image and / or audio sequence. An example of an input to the first modality is video or audio footage of the motion of an object. The video footage may include audio footage.
[0012] The inputs to the second modality may include physical quantities, in particular, values of the object's acceleration, velocity, yaw rate, pitch rate, or roll rate, and one input to the second modality may include at least one value of a physical quantity. The video footage may be accompanied by sensor signals from at least one sensor that measured the physical quantity when capturing the same motion of the object. An example of an input to the second modality is the value of a physical quantity associated with the video footage.
[0013] The input for the second modality can be determined by synthetic data generation, depending on acquired data representing the same motion of an object in the real world. The acquired data may be sensor signals from sensors that measured physical quantities included in the second modality, or it may be sensor signals from sensors that measured physical quantities different from those included in the second modality during the same motion of an object represented by the first modality.
[0014] The input for the second modality may include text, where the text in each input describes the position or part of the motion or movement of the object represented by each input. An example of text is the text description or caption accompanying video footage.
[0015] The method includes providing a signal characterizing the motion of an object, in particular a signal including a first modality; providing a target pattern for the signal; and providing a label associated with the target pattern, wherein determining text includes matching the signal pattern with the target pattern; and determining text to include the label.
[0016] The data structure includes at least one data field for a set of inputs for a first modality and a set of inputs for a second modality, wherein the set of inputs for the first modality includes a sequence of inputs for the first modality representing the motion of an object in the real world, and the set of inputs for the second modality includes a sequence of inputs for the second modality representing the same motion of an object in the real world; the data structure includes at least one data field for the output of the multimodal base model determined for inputs to the multimodal base model, each including different combinations of one input from the set of inputs for the first modality and one input from the set of inputs for the second modality; and the data structure includes at least one data field for the multimodal base model and for the loss, which includes a loss term specifically weighted for each of the outputs.
[0017] A device for training a multimodal platform model comprises at least one processor and at least one memory, the memory of which stores instructions for causing the device to perform a method when executed by at least one processor.
[0018] A computer program can be provided, wherein the computer program comprises computer-readable instructions for causing a computer to carry out the method when executed by the computer.
[0019] Further examples can be derived from the following description and the drawings. Brief Description of the Drawings
[0020] [Figure 1] FIG. 1 is a schematic diagram illustrating an apparatus for training a multimodal foundation model. [Figure 2] FIG. 2 is a diagram showing a flowchart including steps of a method for training a multimodal foundation model. Description of Embodiments
[0021] FIG. 1 schematically shows an apparatus 100 for training a multimodal foundation model, characterized in that the apparatus 100 comprises at least one processor 102 and at least one memory 104.
[0022] The at least one memory 104 stores instructions that, when executed by the at least one processor, cause the apparatus 100 to perform a method for training a multimodal foundation model.
[0023] The apparatus 100 may include at least one sensor 106, or at least one input unit 108 for the at least one sensor 106. An example of the at least one sensor 106 is a sensor for acquiring digital images, particularly a camera. The digital images may be video, radar, LiDAR, ultrasound, motion, or thermal images.
[0024] Examples of the at least one sensor 106 are in particular an IMU or a micro-electro-mechanical system (MEMS) sensor for acquiring acceleration, velocity, yaw rate, pitch rate or roll rate.
[0025] An example of at least one sensor 106 is an accelerometer, a distance meter, a yaw rate sensor, a pitch rate sensor, or a roll rate sensor.
[0026] At least one example of sensor 106 is a microphone for acquiring audio data, particularly audio sequences.
[0027] At least one example of a sensor 106 is a health sensor, specifically a heart rate sensor.
[0028] Figure 2 shows a flowchart that includes the steps of the method.
[0029] The method includes step 202.
[0030] Step 202 includes providing a set of inputs for a first modality and a set of inputs for a second modality.
[0031] The set of inputs for the first modality includes a sequence of inputs for the first modality that represent the motion of an object in the real world.
[0032] The set of inputs for the second modality includes a sequence of inputs for the second modality that represent the same motion of an object in the real world.
[0033] Providing a set of inputs for a second modality may include determining the set of inputs for the second modality by generating synthetic data.
[0034] The set of inputs for the second modality is determined, for example, using generative artificial intelligence, depending on the inputs for the first modality.
[0035] Determining the set of inputs for the second modality by generating synthetic data may include determining the inputs for the second modality depending on the inputs for the first modality.
[0036] Determining the set of inputs for a second modality through synthetic data generation may involve determining the inputs for the second modality based on acquired data representing the motion of objects in the real world.
[0037] Determining the set of inputs for a second modality through synthetic data generation may include determining the set of inputs for a second modality in the simulation of the motion or movement of an object.
[0038] Each input to the first modality contains, for example, different digital images.
[0039] Digital images may be video, radar, LiDAR, ultrasound, motion, or thermal images.
[0040] Each input to the first modality contains, for example, a different audio sequence.
[0041] Each input to the first modality includes, for example, different digital images and audio sequences.
[0042] The input for the second modality includes, for example, physical quantities. Examples of physical quantities include the acceleration, velocity, yaw rate, pitch rate, or roll rate of an object.
[0043] For example, one input to the second modality may contain a single value of a physical quantity. For example, one input to the second modality may contain multiple values of a physical quantity.
[0044] The input for the second modality includes, for example, text. The text in each input describes, for example, the position of the object represented by each input, or part of the motion or movement of the object.
[0045] For example, the text is determined by a signal characterizing the motion of an object. The signal may include a first modality. The signal can be acquired by at least one sensor. At least one sensor can be configured to acquire the object's acceleration, velocity, yaw rate, pitch rate, or roll rate.
[0046] The text is determined, for example, depending on the target pattern for the signal and the label associated with the target pattern.
[0047] Determining text includes, for example, matching a signal pattern with a target pattern and determining text that includes a label associated with the target pattern. For example, a target pattern and a label are provided. Multiple target patterns and a label associated with each of these target patterns may be provided.
[0048] The method includes step 204.
[0049] Step 204 involves determining the output of the multimodal base model depending on various inputs to the multimodal base model.
[0050] Step 204 involves determining inputs to a multimodal base model that include various combinations of one input from the set of inputs for the first modality and one input from the set of inputs for the second modality.
[0051] Step 204 includes the output of the multimodal base model for each input to the multimodal base model.
[0052] The output includes, for example, the embedding of the inputs of the first modality and the embedding of the inputs of the second modality in a common embedding space.
[0053] The embedding of the input for the first modality is determined, for example, using the first encoder E1 of the multimodal base model.
[0054] The embedding of the input for the second modality is determined, for example, using the second encoder E2 of the multimodal base model.
[0055] The first encoder E1 is configured, for example, to output a normalized embedding for the input of a first modality.
[0056] The second encoder E2 is configured, for example, to output a normalized embedding for the input of a second modality.
[0057] Exemplary first modality I Input I1, I2, I3, ..., I N , and the inputs TS1, TS2, TS3, ..., TS of the exemplary second modality TS N Embedding I as element ij of matrix M i TS j An exemplary matrix M is described, which includes the first modality I being image data and the second modality TS being multidimensional IMU data, i.e., data containing accelerometer, gyroscope, and magnetometer values.
[0058]
number
[0059] The input to the multimodal foundation model is not limited to inputs of two mutually different modalities. The input to the multimodal foundation model may include inputs of three or more mutually different modalities. The input to the multimodal foundation model may include two or more inputs of the same modality. Each input of the same modality may originate from mutually different sources. Each input of the same modality can be acquired by mutually different sensors. At least one input of the same modality can be acquired by a sensor, and at least one input can be generated by synthetic data generation.
[0060] The method comprises step 206.
[0061] Step 206 comprises training the multimodal foundation model in dependence on a loss that includes a particularly weighted loss term for each output.
[0062] An exemplary loss for the first modality and the second modality is weight w of elements of a matrix M ij comprising a loss term including. Element ij can each be weighted by weight w ij . The loss is not limited to pairwise loss.
[0063] An exemplary weighted loss L for the first modality and the second modality is, for example, [Formula] where L ij is a loss function, for example, L ij (I i , T_S j ). The loss function is, for example, for a pair I i , T_S j comprises the distance from the embedding of the first modality input I i to the embedding of the second modality input T_S j
[0064] weight wij These may be given as hyperparameters or obtained through training.
[0065] A multimodal base model is, for example, a sensor base model. In a sensor base model, one encoder is trained on IMU time series data, and at least one encoder is trained on another modality. The sensor base model is trained to align the encoder embeddings on the IMU time series data with the embeddings of at least one other modality.
[0066] For example, pairs of image data and IMU data are used to train a multimodal sensor-based model to align the embeddings of two modalities: image and IMU.
[0067] The output of a sensor-based model for the input of one modality is used, for example, for a task.
[0068] An example of a task is to infer an action from the input of one modality.
[0069] An example of a task is anomaly detection, for example, detecting an anomaly when the distance between each embedding, determined for inputs of a first modality and a second modality that characterize the motion of an object simultaneously or within a predetermined time interval, exceeds a threshold.
[0070] A data structure is available, which includes at least one data field for a set of inputs for a first modality and a set of inputs for a second modality.
[0071] The first set of modality inputs includes a sequence of first modality inputs representing the motion of an object in the real world. The second set of modality inputs includes a sequence of second modality inputs representing the same motion of an object in the real world.
[0072] The data structure includes at least one data field for the output of the multimodal base model determined for inputs to the multimodal base model, each containing different combinations of one input from a set of inputs for the first modality and one input from a set of inputs for the second modality.
[0073] The data structure includes at least one data field for the multimodal base model and for the loss, which includes loss terms specifically weighted for each of the outputs.
Claims
1. A computer implementation method for training a multimodal foundational model, The aforementioned method, (202) to provide a set of inputs for a first modality and a set of inputs for a second modality, wherein the set of inputs for the first modality includes a sequence of inputs for the first modality representing the motion of an object in the real world, and the set of inputs for the second modality includes a sequence of inputs for the second modality representing the same motion of the object in the real world. (204) Determining the output of the multimodal base model for inputs to the multimodal base model which includes different combinations of one input from the set of inputs of the first modality and one input from the set of inputs of the second modality, Training the multimodal base model on a loss that includes a loss term that is particularly weighted for each of the outputs (206), A method characterized by including
2. Providing the set of inputs for the second modality (202) includes determining the set of inputs for the second modality by generating synthetic data. The method according to claim 1.
3. Determining the set of inputs for the second modality by generating synthetic data means that Determining the input of the second modality, for example, using generative artificial intelligence, depending on the input of the first modality, or depending on acquired data representing the motion of the object in the real world, Or, In the simulation of the motion of the aforementioned object, the set of inputs for the second modality is determined. including, The method according to claim 2.
4. Each input of the first modality includes a different digital image and / or audio sequence. The method according to any one of claims 1 to 3.
5. The inputs to the second modality include physical quantities, in particular the acceleration, velocity, yaw rate, pitch rate, or roll rate of the object, and one input to the second modality includes at least one value of the physical quantity. The method according to any one of claims 1 to 4.
6. The input for the second modality includes text, the text in each input describing a part of the position or movement of the object represented by each input. The method according to any one of claims 1 to 5.
7. The aforementioned method, To provide a signal characterizing the motion of the object, particularly a signal including the first modality (202), To provide a target pattern for the aforementioned signal, To provide a label associated with the target pattern, wherein determining the text includes matching the signal pattern with the target pattern. Determining the text to include the label, including, The method according to claim 6.
8. A data structure, The data structure includes at least one data field for a set of inputs for a first modality and a set of inputs for a second modality, wherein the set of inputs for the first modality includes a sequence of inputs for the first modality representing the motion of an object in the real world, and the set of inputs for the second modality includes a sequence of inputs for the second modality representing the same motion of the object in the real world. The data structure includes at least one data field for the output of the multimodal base model determined for inputs to the multimodal base model, which include different combinations of one input from the set of inputs of the first modality and one input from the set of inputs of the second modality. The data structure includes at least one data field for the multimodal base model and for the loss, which includes a loss term that is particularly weighted for each of the outputs. A data structure characterized by the following features.
9. A device (100) for training a multimodal base model, The aforementioned device (100) At least one processor (102), At least one memory (104), Equipped with, The at least one memory (104) stores instructions for causing the device (100) to perform the method according to any one of claims 1 to 8 when executed by the at least one processor. A device (100) characterized by the following.
10. A computer program characterized in that, when executed by a computer, it includes a computer-readable instruction causing the computer to perform the method described in any one of claims 1 to 8.