Information processing device, information processing method, and program
The multimodal machine learning model automates the correlation of voice and image data, addressing limitations in existing methods by enabling efficient analysis and inference between sound and image data types, particularly in multimedia applications.
Patent Information
- Application Number
- PCT/JP2024/001260
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-18
- Publication Date
- 2025-07-24
AI Technical Summary
Existing methods for processing the relationship between sound and image data are limited and primarily rely on human analysis, restricting the applications of information processing using this relationship.
An information processing apparatus and method that utilizes a multimodal machine learning model to associate voice-related and image-related data in a predetermined abstract space, enabling the automatic extraction and correlation of data types based on coordinate information, allowing for the generation of voice-related or image-related data from descriptive or non-descriptive inputs.
Enables various information processes using the relationship between voice and image data, facilitating analysis and inference of image data from voice data without direct access, and vice versa, enhancing the automation and efficiency of multimedia processing.
Smart Images

Figure JP2024001260_24072025_PF_FP_ABST
Abstract
Description
Information processing device, information processing method, and program
[0001] The present invention relates to an information processing device, an information processing method, and a program.
[0002] Sounds such as sound effects and music provided along with images (whether still images or moving images) play an important role in videos, movies, games, etc. However, although there are methods such as speech recognition for processing the relationship between images and sounds using computers, the reality is that analysis of, for example, the relationship between sound effects and images has only been carried out by humans.
[0003] Under these circumstances, the applications of information processing that utilizes the relationship between sound and images have been limited.
[0004] The present invention has been made in consideration of the above-mentioned situation, and one of its objects is to provide an information processing device, an information processing method, and a program that enable a variety of information processing using the relationship between audio and images.
[0005] The information processing device is connected to a storage device accessible to the storage device that holds the machine learning model, which uses a descriptive information machine learning model that has been trained to associate image-related data and descriptive information that linguistically describes image data related to the image-related data with coordinate information that is close to each other within a predetermined abstract space, and inputs audio-related data related to audio data and image data provided together with the audio-related data, and has been trained to associate the audio-related data and image-related data, respectively, related to the audio data and image data provided at the same time with coordinate information that is close to each other within the predetermined abstract space. The information processing device includes a processor that accepts input of either non-descriptive information of the audio data or image data, or descriptive information, and uses the machine learning model to obtain, within the abstract space, associated data of one type of the input non-descriptive information or descriptive information, either audio-related data or image-related data, or associated data of the other type associated with coordinate information that is close to coordinate information corresponding to the descriptive information, and outputs the other type of data related to the acquired associated data.
[0006] According to the present invention, various information processes can be performed by utilizing the relationship between audio and images.
[0007] The present invention relates to an information processing device that can process data in a multimodal manner, and an information processing apparatus that can process data in a multimodal manner.
[0008] An information processing device 1 according to an embodiment of the present invention will be described with reference to the drawings. As shown in FIG. 1, the information processing device 1 includes a control unit 11 as a processor, a storage unit 12, an operation control unit 13, and an output control unit 14, and preferably further includes a communication unit 15.
[0009] The information processing device 1 is also connected to an output device 2 and is communicably connected to other computer devices and the like via a communication means such as a network.
[0010] Before describing each unit of the information processing device 1, the machine learning model used by the control unit 11 in its processing will be described below.
[0011] [Machine Learning Model] In one example of the present embodiment, the storage unit 12 corresponds to a storage device that stores a machine learning model (hereinafter referred to as a multimodal machine learning model) used by the control unit 11 in the embodiment of the present invention. Here, the multimodal machine learning model is machine-learned using as input audio-related data related to audio data and image data provided together with the audio data related to the audio-related data. Hereinafter, the audio data and image data will be collectively referred to as non-descriptive information. Specifically, this multimodal machine learning model is in a state where machine learning has been performed so that audio-related data and image-related data (hereinafter collectively referred to as associated data) respectively related to the audio data and image data provided at substantially the same time (e.g., at least a portion of their playback times overlap) are associated with coordinate information that is close to each other in a predetermined abstract space.
[0012] In one example of this embodiment, the multimodal machine learning model uses a model (hereinafter, for the sake of explanation, this machine learning model is referred to as a base model) that has previously learned the relationship between image data and descriptive information (linguistic information, i.e., character string information such as words and sentences). As such a base model, for example, CLIP (Contrastive Language-Image Pre-training: Alec Radford, et al., Learning Transferable Visual Models From Natural Language Supervision, arXiv:2103.00020, 2021) may be used.
[0013] A learning processing device 10 (the information processing device 1 itself may function as this learning processing device 10) that performs machine learning on a basic model such as CLIP accepts multiple pairs of image data and description information describing the image data as training data. For each pair of training data, the learning processing device 10 divides the description information T contained in a given pair into tokens (e.g., words) t and obtains its distributed representation tref (token embedding). The learning processing device 10 then generates description-related data ft(tref) by encoding the tokens using a text encoder. This description-related data ft(tref) is multidimensional vector data and corresponds to one piece of coordinate information in an abstract space.
[0014] The learning processing device 10 also generates image-related data fi(I) by inputting image data I included in the same set as the description information T into an image encoder. The learning processing device 10 sets the image encoder so that this image-related data fi(I) also becomes vector data of the same dimension as the description-related data ft(tref). This image-related data fi(I) also corresponds to one piece of coordinate information in abstract space.
[0015] The learning processing device 10 performs machine learning on the encoders ft and fi so that the description-related data ft(tref) and the image-related data fi(I) are located close to each other in the above-mentioned predetermined abstract space (a vector-dimensional space of the description-related data and the image-related data, known as a multimodal embedding space in CLIP). Conceptually, this results in the image data and the description information describing it being associated with each other in the predetermined abstract space P (S11), as shown in the example of Figure 2. In the above example, the text encoder ft can use a well-known transformer network, and the image encoder fi can use a Vision Transformer (ViT).
[0016] The multimodal machine learning model used by the control unit 11 of this embodiment utilizes an abstract space that associates the image data I thus obtained with the corresponding description information T. Specifically, the learning processing device 10 machine-learns an encoder fa that encodes the voice data A into voice-related data fa(A) (vector data of the same dimension as the description-related data ft(tref)), and associates the voice data A with coordinate information within this abstract space (S12). At this time, there is no need to retrain the encoders corresponding to the description-related data and the image-related data; their parameters may be fixed. Furthermore, this voice-related data fa(A) also corresponds to one piece of related data.
[0017] The learning processing device 10, which associates this audio data with the abstract space of the basic model to obtain the multimodal machine learning model of this embodiment, accepts input of video data I(τ) played back together with audio data A(τ) (where τ is time), which is time-series information. Such video data with audio data may be, for example, audio and video recorded by synchronizing game screens and sound, or may be recorded by a drive recorder installed in a vehicle. Furthermore, this video data with audio data may be a movie, a television program, or a video distributed via the Internet, etc.
[0018] The learning processing device 10 obtains audio-related information fa(A(τ)) related to the audio data at each point in time from the audio data A(τ) played back together with the video data I(τ) using a predetermined encoder (for example, an AST (Audio Spectrogram Transformer) or a SWIN-Transformer).
[0019] The learning processing device 10 also obtains audio-related information fa(A(τ)) obtained in relation to audio data A(τ) at each time point, and image data displayed at that time point (which may be a still image I(τ) constituting video data) by encoding the image data using the same image encoder as the basic model to obtain image-related data fi(I(τ)).The learning processing device 10 then machine-learns the encoder fa corresponding to the audio data so that the obtained audio-related information fa(A(τ)) approximates the image-related data fi(I(τ)).
[0020] In this example of generating a multimodal machine learning model using a basic model, it is possible to learn the relationship between image data and associated audio data through machine learning without label information, and as a result, it is possible to learn the relationship between descriptive information and audio data through machine learning. This also results in the relationship between audio data and image data being learned through machine learning via descriptive information.
[0021] Furthermore, by using video image data, it is possible to perform machine learning to determine the relationship between image data and associated audio data without label information. This means that in various application fields, simply by performing machine learning processing using video image data related to the target field, it is possible to perform processing suitable for analyzing images and audio in the target field.
[0022] [Configuration of Each Unit] The control unit 11 of this embodiment is a program-controlled device such as a CPU, and realizes the processor of the present invention. The control unit 11 is accessible to the storage unit 12 and operates according to a program stored in the storage unit 12. The control unit 11 accepts input of either non-descriptive information (audio data or image data) or descriptive information according to the program. Using a multimodal machine learning model stored in the storage unit 12, the control unit 11 acquires, within the predetermined abstract space, audio-related data or image-related data, which is associated data related to one of the input non-descriptive information and descriptive information, or associated data related to the other type associated with coordinate information close to coordinate information corresponding to the descriptive-related data. The control unit 11 then executes processing to output the other type of data related to the acquired associated data. Detailed processing by the control unit 11 will be described later.
[0023] The storage unit 12 stores a program executed by the control unit 11. This program may be provided by being stored in a computer-readable, non-transitory recording medium and stored in the storage unit 12. The storage unit 12 also operates as a work memory for the control unit 11. Furthermore, in this example of the present embodiment, as described above, the storage unit 12 stores the multimodal machine learning model (including information representing the structure of the machine learning model and parameter information such as corresponding weights and biases) used in the processing of the control unit 11.
[0024] The operation control unit 13 receives information representing the content of a user's operation from the operation device 3, such as a mouse, keyboard, or game controller. The operation control unit 13 outputs the information representing the received content of the operation to the control unit 11. The output control unit 14 outputs image and audio data to be displayed to the output device 2 in accordance with instructions received from the control unit 11.
[0025] The communication unit 15 is a network interface or the like, and according to instructions input from the control unit 11, sends and receives various data to and from other computer devices or the like connected via communication means such as a network.
[0026] [Processing of Control Unit] Next, a description will be given of processing in the control unit 11. The control unit 11 of this embodiment executes a program stored in the storage unit 12, thereby realizing a configuration that functionally includes a data acceptance unit 21, a data acquisition unit 22, and an output processing unit 23, as exemplified in FIG.
[0027] Here, the data receiving unit 21 receives input of either non-descriptive information or descriptive information, such as audio data or image data, in response to an instruction from the user.
[0028] The data acquisition unit 22 uses the multimodal machine learning model stored in the storage unit 12 to obtain associated data related to the type of data accepted by the data acceptance unit 21 (audio-related data for audio data, image-related data for image data, and description-related data for description information) using an encoder corresponding to the type of data accepted by the data acceptance unit 21 within the abstract space used by the multimodal machine learning model. This process is hereinafter referred to as inference processing. For example, if the type of data accepted by the data acceptance unit 21 is audio data, the data acquisition unit 22 generates audio-related data related to the audio data using an audio encoder corresponding to the audio data in the multimodal machine learning model.
[0029] Furthermore, if the type of data accepted by the data accepting unit 21 is descriptive information, the data acquiring unit 22 generates description-related data relating to the descriptive information using an encoder corresponding to the descriptive information in the multimodal machine learning model.
[0030] When the data acquisition unit 22 generates the associated data, it acquires the other type of associated data associated with coordinate information close to the coordinate information corresponding to the generated associated data.
[0031] For example, if the type of data accepted by the data accepting unit 21 is audio data, the data acquiring unit 22 acquires description-related data, which is related data of the other type associated with coordinate information P2 close to coordinate information P1 corresponding to the audio-related data related to the audio data (for example, this coordinate information P2 may be P1 itself or may be coordinate information within a predetermined distance from P1; the same applies hereinafter to the case of close coordinate information).
[0032] Furthermore, if the type of data accepted by the data accepting unit 21 is descriptive information (such as a query that serves as a search key), the data acquiring unit 22 acquires voice-related data, which is another type of related data associated with coordinate information P4 (as already mentioned, for example, this coordinate information P4 may be P3 itself or may be coordinate information within a predetermined distance from P3) that is close to coordinate information P3 corresponding to the descriptive related data related to the descriptive information.
[0033] The output processing unit 23 outputs the other type of data related to the related data acquired by the data acquisition unit 22 for predetermined processing. For example, if the related data acquired by the data acquisition unit 22 is description-related data, the output processing unit 23 outputs description information obtained by decoding the description-related data with a predetermined decoder. Furthermore, if the related data acquired by the data acquisition unit 22 is audio-related data, the output processing unit 23 decodes the audio-related data with a predetermined decoder and outputs audio data related to the audio-related data.
[0034] Here, a decoder is provided corresponding to each type of related data (image related data, audio related data, description related data, etc.). The decoder receives input of related data and outputs data of a type corresponding to the related data (image data, audio data, description information, etc.). Such decoders can be obtained by machine learning or dictionary formation, which will be described later, and can be obtained by well-known methods corresponding to the encoders for each type of data in the multimodal machine learning model, so detailed description thereof will be omitted here.
[0035] In addition, when obtaining a decoder by forming a dictionary, that is, for data on which dictionary information is recorded, the data acquisition unit 22 obtains coordinate information Px corresponding to the related data related to the data accepted by the data acceptance unit 21, and then obtains coordinate information Py that is closest to the obtained coordinate information Px from among the coordinate information corresponding to the related data related to the tokens recorded in the dictionary information (tokens associated with data of a different type from the data accepted by the data acceptance unit 21).
[0036] In this case, the output processing unit 23 performs decoder processing to obtain a token related to associated data corresponding to the acquired coordinate information Py. The output processing unit 23 then refers to dictionary information to obtain and output data associated with the obtained token.
[0037] In this example, if dictionary information is recorded for descriptive information or speech data, the descriptive information or speech data itself that was input during machine learning will be output.
[0038] In another example, for non-descriptive information such as image data or audio data, the associated data may be embedded in an abstract space, modulated using a Q-former or multilayer perceptron to generate tokens, and then embedded before and after the descriptive information, allowing for the generation of sentences associated with the non-descriptive information. This example also makes it possible to obtain captions from audio data.
[0039] [Example of Use] The information processing device 1 of this embodiment configured as described above can be used to realize the following process. That is, in one example of this embodiment, in order to generate a multimodal machine learning model to be used by the information processing device 1, the learning processing device 10 accepts, as training data, video data of a licensed subject sleeping who has given permission to be filmed in advance, and audio data recorded in synchronization with the video data. This training data preferably includes not only video data of a single licensed subject sleeping, but also video data of multiple licensed subjects sleeping.
[0040] In this example, the learning processing device 10 uses, for example, CLIP as a basic model and machine-learns the encoder of the audio data A(τ) so that image-related data fi(I(τ)) obtained by encoding moving image data I(τ) included in the learning data and audio-related data fa(A(τ)) obtained by encoding audio data A(τ) recorded in synchronization with the image-related data fi(I(τ)) are close to each other in terms of coordinate information within the abstract space of CLIP, which is the basic model. In this way, the learning processing device 10 generates a multimodal machine learning model to be used by the information processing device 1.
[0041] The information processing device 1 stores the multimodal machine learning model generated by the learning processing device 10 in the storage unit 12 and executes the following processing: In this example, the information processing device 1 accepts audio data recorded while a non-consenting subject who has not given permission to be photographed is sleeping, divides the accepted audio data into predetermined units (for example, units of a predetermined time or units of continuous audio divided by silent portions), and performs the following processing using the divided audio data (hereinafter referred to as partial audio data) as input data.
[0042] The control unit 11 of the information processing device 1, as processing of the data processing unit 22, uses an audio encoder corresponding to the audio data in the multimodal machine learning model stored in the storage unit 12 to obtain audio-related data corresponding to each of the partial audio data that became input data. The control unit 11 then sequentially selects each partial audio data as partial audio data of interest and obtains coordinate information K in the abstract space used by the multimodal machine learning model corresponding to the audio-related data related to the partial audio data of interest. The control unit 11 then searches for coordinate information T in the abstract space that corresponds to the description-related information and is closest to the coordinate information K obtained corresponding to the audio-related data (or is below a predetermined distance threshold (exceeds a predetermined similarity threshold)). Note that, when a distance threshold or a similarity threshold is used, these thresholds may be arbitrarily set by the user.
[0043] The control unit 11 decodes the description-related information associated with the coordinate information T found by this search to obtain description information, and records the obtained description information in association with the target portion audio data. At this time, the control unit 11 may also associate information specifying the time (corresponding time information) from the beginning of the audio data of the target portion audio data (before division) corresponding to the target portion audio data.
[0044] Furthermore, when searching for coordinate information that is less than a predetermined distance (exceeds a predetermined similarity), if there is no coordinate information T that meets the search conditions, the control unit 11 may record information that there is no corresponding descriptive information, in association with the audio data of the portion of interest.
[0045] The control unit 11 performs the above processing on each partial voice data and obtains corresponding descriptive information for each partial voice data. This descriptive information is expected to be sleep-related, such as "sleep," "wakefulness," "turning over," and "snoring." As processing by the output processing unit 23, the control unit 11 displays a list of the obtained descriptive information, for example, in chronological order of the partial voice data. As a result, even if image data during sleep cannot be obtained due to privacy concerns, the sleeping state of the unauthorized subject can be inferred and obtained based on the voice data recorded while the unauthorized subject was sleeping, and this can be used to analyze the state of breathing, turning over, etc.
[0046] In other words, according to the information processing device 1 of this embodiment, a machine learning model that has learned the relationship between descriptive information and image data can be extended using information that indicates the relationship between audio data and image data (for example, video data with audio), thereby obtaining a machine learning model that has learned the relationship between descriptive information and audio data, which is relatively difficult to generate.
[0047] Similarly, according to the method of this embodiment, when there is data of type X and data of type Y that are associated with each other in some kind of embedding space (corresponding to the predetermined abstract space of this embodiment) using a machine learning model that has already undergone machine learning, if there is interrelationship information indicating the association between the data of type Y and data of type Z (such as a pair of data of types Y and Z that occur simultaneously), this interrelationship information can be used to extend the machine learning model (associating the data of type Z with the embedding space), thereby easily generating a machine learning model in which the relationship between the data of types X and Z has been machine learned. Here, in the above example, type X was descriptive information, type Y was image data, and type Z was audio data, but the method of this embodiment is not limited to these data.
[0048] Alternatively, the control unit 11 may perform the following process: In this example, when the control unit 11 acquires coordinate information in the abstract space used by the multimodal machine learning model corresponding to the speech-related data for each partial speech data, the control unit 11 associates each partial speech data with the corresponding coordinate information (coordinate information in the abstract space) and stores the information as partial speech dictionary information.
[0049] Furthermore, the control unit 11 in this example may record, for each partial voice data, information specifying the corresponding time information corresponding to the partial voice data in association with the partial voice dictionary information.
[0050] In this example, the control unit 11 receives descriptive information such as "sleeping," "awake," and "turning over" as a query from the user. The control unit 11 encodes the descriptive information using a corresponding encoder to obtain description-related information. The control unit 11 also uses a multimodal machine learning model to obtain coordinate information corresponding to the obtained description-related information within the abstract space. The control unit 11 searches for coordinate information recorded in the partial speech dictionary information that is closest to the coordinate information obtained corresponding to the description-related information (or coordinate information that is less than a predetermined distance (more than a predetermined similarity)).
[0051] The control unit 11, operating as the output processing unit 23, outputs partial speech data stored in the partial speech dictionary information in association with the coordinate information found by the search. Alternatively, when searching for coordinate information that is less than a predetermined distance (exceeding a predetermined similarity), if no coordinate information matching the conditions is found, the control unit 11 may display a message indicating that partial speech data matching the conditions cannot be found. Furthermore, if the partial speech dictionary information includes corresponding time information, the control unit 11, operating as the output processing unit 23, may output the partial speech data stored in the partial speech dictionary information in association with the coordinate information found by the search, and the corresponding time information corresponding to the partial speech data.
[0052] With this process, even if image data during sleep cannot be obtained due to privacy concerns, etc., information inferred from audio data recorded while the unauthorized subject is sleeping can be effectively used to search for specific situations during the subject's sleep, and can be used to analyze breathing, turning over, etc.
[0053] Similarly, by using image data from a drive recorder that records audio data together with video data and the audio data recorded in synchronization as learning data to obtain and use a multimodal machine learning model based on CLIP, the information processing device 1 can use image data captured by a drive recorder that does not record audio data as input and infer and obtain corresponding audio data, and can then use the obtained audio data for analysis.
[0054] Furthermore, when a multimodal machine learning model is obtained using a basic model such as CLIP that associates image data and descriptive information with each other, and in the abstract space of the multimodal machine learning model, description-related data related to descriptive information, image-related data related to image data, and audio-related data related to audio data are associated, the information processing device 1 may perform the following processing using this multimodal machine learning model.
[0055] The information processing device 1 in this example executes a program such as a game application or a video playback application to execute a process of outputting image data together with audio data, acquires description-related data associated with coordinate information close to corresponding coordinate information in the abstract space of audio-related data related to the audio data provided together with the image data, and decodes the description information related to the acquired description-related data using a corresponding decoder to obtain and output it.
[0056] In this example, the information processing device 1 can present to the user descriptive information indicating what the audio data provided together with the image data represents.
[0057] In another example, the information processing device 1 divides audio data and image data (called original data) output during execution of a program such as a game application or a video playback application into predetermined units (for example, time units such as one minute or scene units), and encodes the audio data and image data obtained by the division (called partial audio data and partial image data, respectively) to obtain corresponding audio-related data and image-related data.
[0058] The information processing device 1 obtains coordinate information Ap_i, Vp_i (i = 1, 2, ...) within the abstract space related to the above-mentioned multimodal machine learning model, which corresponds to the audio-related data and image-related data obtained for each partial audio data and partial image data, and obtains coordinate information T(Ap_i), T(Vp_i) that is close to each of the coordinate information (for example, closer than a predetermined threshold value or greater than a predetermined similarity) and is associated with any of the description-related data.
[0059] The information processing device 1 records the acquired coordinate information T(Ap_i) and T(Vp_i) as search target information in association with the corresponding partial audio data and partial image data.
[0060] The information processing device 1 then accepts input of descriptive information from the user, which is a query for searching for scenes, encodes the descriptive information input as the query, obtains corresponding description-related information, and obtains coordinate information τ within the above-mentioned abstract space corresponding to the obtained description-related information.
[0061] The information processing device 1 searches for coordinate information T(Ap_i), T(Vp_i) included in the previously recorded search target information, which is closest to coordinate information τ in the abstract space obtained based on the query, or which is within a predetermined distance from coordinate information τ (or is similar to coordinate information τ beyond a predetermined similarity).
[0062] If coordinate information T(Ap_i) and T(Vp_i) is found by this search, the information processing device 1 obtains audio-related data and image-related data related to the coordinate information Ap_i and Vp_i corresponding to the found coordinate information T(Ap_i) and T(Vp_i), and outputs partial audio data and partial image data related to the obtained audio-related data and image-related data, or reproduces a portion of the original data corresponding to the partial audio data or partial image data.
[0063] Furthermore, in this example, description information can be obtained using both image data and audio data. That is, the information processing device 1 obtains image-related data related to image data and audio-related data related to audio data provided together with the image data using a multimodal machine learning model, and acquires coordinate information V and A corresponding to each associated data in the abstract space of the multimodal machine learning model.
[0064] The information processing device 1 then calculates coordinate information F obtained by, for example, the linear combination of these values, F = λ1 × A + λ2 × V, acquires description-related data associated with coordinate information close to this coordinate information F, and outputs description information related to the acquired description-related data.
[0065] Here, λ1 and λ2 (λ1 + λ2 = 1) are, for example, real number parameters that may be determined in advance or set appropriately by the user. Furthermore, these parameters may be determined by a machine learning model that uses, for example, descriptive information serving as a query or coordinate information V and A as input and outputs each parameter.
[0066] When machine learning the relationship between the descriptive information that constitutes the query and these parameters, the learning processing device (which may be the information processing device 1 itself) uses multiple query samples and pre-prepared correct answers for each sample (information indicating whether the query emphasizes audio or video) as training information, and machine-learns and uses a predetermined machine learning model that takes the query as input and outputs λ1 and λ2.
[0067] In another example, a learning processing device (which may be the information processing device 1 itself) uses a set of video data with audio (a series of image data played in synchronization with the audio data) and descriptive information explaining the content of the video data as learning data to machine-learn a machine learning model that expresses the relationship between the parameters λ1 and λ2 and the coordinate information V and A. Such machine learning processing can be performed using the widely known backpropagation method, and therefore a detailed description thereof will be omitted here.
[0068] According to this example of the present embodiment, for video data with audio capturing a scene in which a dog is barking (the audio data contains the sound of the dog), it is expected that the coefficient λ1 obtained using the coordinate information A in the abstract space corresponding to the audio-related data and the coordinate information V in the abstract space corresponding to the image-related data will be set to a relatively large value, and the coefficient λ2 will be set to a relatively small value. Also, for video data with audio capturing a scene in which a dog is running around (the audio data contains almost no sound of the dog), it is expected that the coefficient λ1 of the coordinate information A will be set to a relatively small value, and the coefficient λ2 will be set to a relatively large value.
[0069] In this way, the information processing device 1 of this embodiment uses a multimodal machine learning model that holds information about an abstract space that correlates image-related data related to image data, audio-related data related to audio data, and description-related data related to description information. By inputting any type of data, such as image data, audio data, or description information, and acquiring related data related to other types of data associated with coordinate information close to coordinate information in the abstract space of the related data related to the relevant data, and outputting the acquired data related to the related data related to the other types of data, it is possible to input audio data to acquire related image data or description information, or to input description information to acquire related image data or audio data (or both). The latter example can be applied to processes such as image search (e.g., presenting the time point at which video data similar to output image data or audio data is played).
[0070] In addition, in the example where image data is input and audio data is output, the information processing device 1 may play audio data that matches the image represented by the image data. In this example, if the audio data is audio data for a song, it becomes possible to select background music that matches the image data without any manual operation.
[0071] [Variant Example] In the explanation up to this point, the multimodal machine learning model has been stored in the memory unit 12, but this is just one example, and the control unit 11 of the information processing device 1 may also use a multimodal machine learning model held by another computer device that is communicatively connected via the communication unit 15.
[0072] REFERENCE SIGNS LIST 1 information processing device, 2 output device, 3 operation device, 10 learning processing device, 11 control unit, 12 storage unit, 13 operation control unit, 14 output control unit, 15 communication unit, 21 data reception unit, 22 data acquisition unit, 23 output processing unit.
Claims
1. A machine learning model using a description information machine learning model that has been machine-learned to associate image-related data and description-related data related to description information that linguistically describes the image data with coordinate information that is close to each other within a predetermined abstract space, the machine learning model being an information processing device that is connected to be accessible to a storage device that holds a machine learning model in a state where machine learning has been performed so that voice-related data related to voice data and image data provided together with the voice data related to the voice-related data are associated with coordinate information that is close to each other within the predetermined abstract space, the information processing device including a processor, the processor receiving an input of either non-description information or description information of voice data or image data, using the machine learning model to obtain, within the abstract space, voice-related data or image-related data that is related data of one type among the input non-description information and description information, or related data of the other type that is associated with coordinate information close to the coordinate information corresponding to the description-related data, and outputting the data of the other type related to the obtained related data.
2. The information processing device according to claim 1, wherein the processor executes a process of outputting image data together with voice data, and uses the machine learning model to obtain description-related data that is associated with coordinate information close to the corresponding coordinate information within the predetermined abstract space of the voice-related data related to the voice data output together with the output image data, and outputs the description information related to the obtained description-related data.
3. The information processing device according to claim 2, wherein the processor further obtains description-related data that is associated with coordinate information close to the corresponding coordinate information within the abstract space of the image-related data related to the image data output together with the voice data, and outputs the description information related to the obtained description-related data.
4. An information processing apparatus according to claim 1, wherein the machine learning model is a machine learning model using a description information machine learning model that has been machine learned to associate image-related data and description-related data related to description information that linguistically describes the image data related to the image-related data with coordinate information that is close to each other within a predetermined abstract space, and uses moving image data including image data and audio data as input, and associates audio-related data related to the audio data included in the moving image data and image data provided together with the audio data with coordinate information that is close to each other within the predetermined abstract space for machine learning of the audio-related data and the image-related data respectively related to the audio data and the image data provided at the same time.
5. Using an information processing apparatus that is accessibly connected to a storage device that holds a machine learning model in a state where machine learning has been performed to associate image-related data and description information that linguistically describes the image data related to the image-related data with coordinate information that is close to each other within a predetermined abstract space, and uses, as input, audio-related data related to audio data and image data provided together with the audio data related to the audio-related data, and where the audio-related data and the image-related data respectively related to the audio data and the image data provided at the same time are in a state where machine learning has been performed to associate them with coordinate information that is close to each other within the predetermined abstract space, the information processing apparatus accepts input of either non-description information or description information of audio data or image data, uses the machine learning model to obtain, within the abstract space, related data of the other type associated with coordinate information close to the coordinate information corresponding to audio-related data, image-related data, or description-related data, which is related data of one of the types of the input non-description information and description information, and outputs data of the other type related to the obtained related data.
6. A machine learning model using a description information machine learning model that has been machine-learned to associate image-related data and description information that linguistically describes the image data related to the image-related data with coordinate information that is close to each other within a predetermined abstract space, which receives as input voice-related data related to voice data, voice data related to the voice-related data, and image data provided together therewith, and is connected in an accessible manner to a storage device that holds a machine learning model in a state where the voice-related data and image-related data respectively related to the voice data and image data provided at the same time are machine-learned to be associated with coordinate information that is close to each other within the predetermined abstract space. An information processing device that accepts input of either non-description information of voice data or image data or description information, uses the machine learning model to obtain, within the abstract space, voice-related data or image-related data, which is related data of one type among the input non-description information and description information, or related data of the other type associated with coordinate information close to the coordinate information corresponding to the description-related data, and functions to output the other type of data related to the obtained related data.
Citation Information
Patent Citations
Information processing apparatus, information processing method, and information processing program
JP2023135777A
Machine learning image search
US20210089571A1
Information processing system, information processing method, and recording medium
WO2020255767A1