Information processing device, method for controlling information processing device, and storage medium
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2026-08-13
Smart Images

Figure JP2025004091_13082026_PF_FP_ABST
Abstract
Description
Information Processing Apparatus, Control Method of Information Processing Apparatus, and Storage Medium
[0001] The present invention relates to an information processing apparatus, an information processing method, and a storage medium.
[0002] There is a technology that uses the voice of a speaker as an input to recognize the utterance content of the speaker (see Non-Patent Document 1). Alternatively, there is a technology that uses the uttered voice of a speaker and an image of the speaker as inputs to recognize the utterance content (see Non-Patent Document 2).
[0003] Bowen Shi, et al. 2 authors, "Robust Self-Supervised Audio-Visual Speech Recognition", [online], [searched on January 10, 2025], Internet <https: / / arxiv.org / pdf / 2201.01763> Andrew Rouditchenko, et al. 6 authors, "Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation", [online], [searched on January 10, 2025], Internet <https: / / arxiv.org / pdf / 2406.10082>
[0004] In the technologies disclosed in Non-Patent Documents 1 and 2, the utterance content of a speaker is estimated mainly using voice. Such a speech recognition model has a problem that it is difficult to obtain a correct estimation result in a noisy environment or when multiple people are speaking simultaneously. That is, such a speech recognition model has a problem that the estimation result of the utterance content (utterance text) depends on the environment.
[0005] A main object of the present invention is to provide an information processing apparatus, a control method of the information processing apparatus, and a storage medium that contribute to enabling more accurate estimation of the utterance text of a speaker regardless of the environment in which the speaker is placed.
[0006] According to a first aspect of the present invention, an information processing device is provided, comprising: an image data acquisition means for acquiring at least one image data showing a speaker; an auxiliary information extraction means for extracting auxiliary information from the at least one image data to assist in estimating a spoken sentence by the speaker; a facial feature information extraction means for extracting facial feature information relating to all or part of the speaker's face from the at least one image data; a spoken sentence estimation means for estimating a spoken sentence by the speaker using the extracted auxiliary information and facial feature information; and an estimation result output means for outputting the estimated spoken sentence.
[0007] A control method for an information processing device is provided, comprising: an image data acquisition step of acquiring at least one image data showing a speaker; an auxiliary information extraction step of extracting auxiliary information from the at least one image data to assist in estimating a spoken sentence by the speaker; a facial feature information extraction step of extracting facial feature information relating to all or part of the speaker's face from the at least one image data; a spoken sentence estimation step of estimating a spoken sentence by the speaker using the extracted auxiliary information and facial feature information; and an estimation result output step of outputting the estimated spoken sentence.
[0008] A third aspect of the present invention is provided, a computer-readable storage medium that stores a program for causing a computer to execute: an image data acquisition process for acquiring at least one image data showing a speaker; an auxiliary information extraction process for extracting auxiliary information from the at least one image data to assist in estimating a spoken sentence by the speaker; a facial feature information extraction process for extracting facial feature information relating to all or part of the speaker's face from the at least one image data; a spoken sentence estimation process for estimating a spoken sentence by the speaker using the extracted auxiliary information and facial feature information; and an estimation result output process for outputting the estimated spoken sentence.
[0009] According to each aspect of the present invention, an information processing device, a control method for the information processing device, and a storage medium are provided that contribute to more accurately estimating the spoken text of a speaker regardless of the environment in which the speaker is located. However, the effects of the present invention are not limited to those described above. The present invention may produce other effects in lieu of or in conjunction with the effects described above.
[0010] Figure 1 is a diagram illustrating the outline of one embodiment. Figure 2 is a flowchart illustrating the operation of one embodiment. Figure 3 is a diagram illustrating an example of the schematic configuration of an information processing system according to an embodiment of this disclosure. Figure 4 is a diagram illustrating the operation of an information processing system according to an embodiment of this disclosure. Figure 5 is a diagram illustrating an example of the processing configuration of a recognition device according to an embodiment of this disclosure. Figure 6 is a flowchart illustrating an example of the operation of a recognition device according to an embodiment of this disclosure. Figure 7 is a diagram illustrating the operation of an auxiliary information extraction unit according to a modified example of an embodiment of this disclosure. Figure 8 is a diagram illustrating the operation of an auxiliary information extraction unit according to a modified example of an embodiment of this disclosure. Figure 9 is a diagram illustrating the operation of an auxiliary information extraction unit according to a modified example of an embodiment of this disclosure. Figure 10 is a diagram illustrating the operation of an auxiliary information extraction unit according to a modified example of an embodiment of this disclosure. Figure 11 is a diagram illustrating the operation of a recognition device according to a modified example of an embodiment of this disclosure. Figure 12 is a diagram illustrating an example of the hardware configuration of a recognition device according to this disclosure.
[0011] First, an overview of one embodiment will be described. The reference numerals in the drawings attached to this overview are provided for convenience as examples to aid understanding, and this overview is not intended to be limiting in any way. Furthermore, unless otherwise specified, the blocks shown in each drawing represent functional units, not hardware units. The connecting lines between blocks in each drawing include both bidirectional and unidirectional lines. Unidirectional arrows schematically indicate the flow of the main signal (data) and do not exclude bidirectional flow. In this specification and in the drawings, elements that can be similarly described are given the same reference numerals to avoid redundant explanation.
[0012] An information processing device 100 according to one embodiment includes an image data acquisition means 101, an auxiliary information extraction means 102, a facial feature information extraction means 103, a speech sentence estimation means 104, and an estimation result output means 105 (see Figure 1). The image data acquisition means 101 acquires at least one image data showing the speaker (step S1 in Figure 2). The auxiliary information extraction means 102 extracts auxiliary information from at least one image data to assist in estimating the speech sentence spoken by the speaker (step S2). The facial feature information extraction means 103 extracts facial feature information relating to all or part of the speaker's face from at least one image data (step S3). The speech sentence estimation means 104 estimates the speech sentence spoken by the speaker using the extracted auxiliary information and facial feature information (step S4). The estimation result output means 105 outputs the estimated speech sentence (step S5).
[0013] The information processing device 100 receives image data of a speaker and infers the spoken text based on the image data. The information processing device 100 is an image-based speech recognition device that does not depend on speech. The information processing device 100 can more accurately estimate the spoken text of a speaker even in situations where it is difficult for existing speech recognition technologies to accurately estimate spoken text, such as in noisy environments or when multiple people are speaking at the same time. This is because the information processing device 100 estimates spoken text using image data without relying on speech.
[0014] Specific embodiments will be described in more detail below with reference to the drawings.
[0015] [First Embodiment] The first embodiment will be described in more detail with reference to the drawings.
[0016] [System Configuration and Outline Operation] As shown in Figure 3, the information processing system according to the first embodiment includes a recognition device 10 and a terminal 20.
[0017] The recognition device 10 is an information processing device that recognizes the voice emitted by the user. For example, the recognition device 10 is a server installed on a network (on the cloud).
[0018] Each user possesses a device 20 such as a smartphone, mobile phone, tablet, or game console. Users communicate with other users via the device 20 and the recognition device 10.
[0019] For example, user A and user B communicate via the recognition device 10 and terminal 20 (see Figure 4). At that time, user A's terminal 20 uses its built-in camera to take a picture of user A and transmits the image data (video data) showing user A's face to the recognition device 10.
[0020] The recognition device 10 uses image data to estimate the spoken text (speech content) of user A and transmits the estimated spoken text, including the text data, to user B's terminal 20. User B's terminal 20 displays the received spoken text. For example, terminal 20 displays the received spoken text (speech content) like subtitles.
[0021] For example, if user A says "Good morning" to terminal 20, the string "Good morning" will be displayed on user B's terminal 20.
[0022] As described above, the recognition device 10 according to the embodiment disclosed herein estimates the sentence spoken by the user using image data (at least one image data) of the user's face, and outputs the estimated sentence. In doing so, the recognition device 10 estimates the sentence without using the user's voice.
[0023] Next, we will describe the details of the recognition device 10 according to the first embodiment.
[0024] Figure 5 shows an example of the processing configuration (processing module) of the recognition device 10 according to the embodiment disclosed herein. Referring to Figure 5, the recognition device 10 comprises a communication control unit 201, an image data acquisition unit 202, an auxiliary information extraction unit 203, a facial feature information extraction unit 204, a speech text estimation unit 205, an estimation result output unit 206, and a storage unit 207.
[0025] The communication control unit 201 is a means for controlling communication with other devices. For example, the communication control unit 201 receives data (packets) from the terminal 20. The communication control unit 201 also transmits data to the terminal 20. The communication control unit 201 passes the data received from other devices to other processing modules. The communication control unit 201 transmits the data acquired from other processing modules to other devices. In this way, other processing modules send and receive data with other devices via the communication control unit 201. The communication control unit 201 has the function of a receiving unit that receives data from other devices and the function of a transmitting unit that transmits data to other devices.
[0026] The image data acquisition unit 202 is a means for acquiring at least one image data that shows the speaker. More specifically, the image data acquisition unit 202 acquires at least one image data that can be determined to be temporally continuous. For example, the image data acquisition unit 202 acquires image data (at least one image data; for example, video data) from the terminal 20. The image data acquisition unit 202 temporarily stores the image data received from the terminal 20.
[0027] The image data acquisition unit 202 outputs a series of temporarily stored image data (buffered image data) that are assumed to represent continuous speech by the user to the auxiliary information extraction unit 203 and the facial feature information extraction unit 204. For example, the image data acquisition unit 202 determines that speech by the user has started if the difference between the image data is large, and determines that speech by the user has ended if the difference between the image data remains small for a certain period of time. In this case, the image data acquisition unit 202 may use the difference between image data in the lip region (mouth region) included in each image data as the difference between the image data.
[0028] For example, consider a case where a user says "Good morning," and then, after a predetermined time (for example, several hundred millimeters later), says "It's a nice day today." In this case, the image data acquisition unit 202 outputs a series of image data showing the user's face when they say the first sentence, "Good morning," to the auxiliary information extraction unit 203 and the face feature information extraction unit 204. Subsequently, the image data acquisition unit 202 outputs a series of image data showing the user's face when they say "It's a nice day today," to the auxiliary information extraction unit 203 and the face feature information extraction unit 204.
[0029] In the following explanation, a series of image data that is assumed to represent continuous speech by the user will be referred to as "speech image data group".
[0030] The speech recognition model is comprised of an auxiliary information extraction unit 203, a facial feature information extraction unit 204, and a speech text estimation unit 205. The speech recognition model performs speech recognition using a group of speech image data and outputs text data that includes the speech text (speech content).
[0031] The speech recognition model according to the embodiment disclosed herein extracts (generates) auxiliary information and facial feature information from a group of speech image data. The speech recognition model performs speech recognition using the extracted auxiliary information and facial feature information. For example, the speech recognition model inputs a prompt (query) generated using the extracted auxiliary information into a Large Language Model (LLM) to obtain the spoken text of the person depicted in each image data in the group of speech image data. Alternatively, the speech recognition model may tokenize the extracted auxiliary information and facial feature information, input the tokenized auxiliary information and facial feature information into the Large Language Model, and obtain the spoken content of the person depicted in each image data in the group of speech image data. Note that tokenization of information means representing information such as words and images as numerical vectors.
[0032] The auxiliary information extraction unit 203 is a means for extracting auxiliary information that assists in estimating the spoken sentence from at least one image data. More specifically, the auxiliary information extraction unit 203 extracts auxiliary information from a group of spoken image data.
[0033] Here, auxiliary information refers to information obtained from a set of speech image data that assists the speech recognition model in estimating the spoken text. For example, information about the speaker's language (spoken language; e.g., Japanese, English) is given as an example of auxiliary information.
[0034] Alternatively, vowel sequences (e.g., a, i, u, e, o, n) or phoneme sequences in speech may be used as examples of supplementary information. Alternatively, information regarding the user's emotions (e.g., happy, sad) (emotion estimation results) may be used as examples of supplementary information.
[0035] For example, the auxiliary information extraction unit 203 extracts the language used as auxiliary information based on the race of the person depicted in each image data that makes up the speech image data group. Alternatively, the auxiliary information extraction unit 203 extracts vowel sequences and phoneme sequences (vowel phoneme sequences and consonant phoneme sequences) as auxiliary information from the shape and movement of the lips of the person depicted in the image data.
[0036] Alternatively, the auxiliary information extraction unit 203 may extract information about the lip region (mouth region) contained in each image data as auxiliary information. More specifically, the auxiliary information extraction unit 203 may extract information about the shape of the lips as auxiliary information. For example, the auxiliary information extraction unit 203 may extract information such as "round mouth" (like pursed lips) or "flat mouth" (like lips with sides pulled back) as auxiliary information. Alternatively, the auxiliary information extraction unit 203 may extract the coordinates of each point (feature point) that characterizes the lips as auxiliary information. Alternatively, the auxiliary information extraction unit 203 may extract information such as the width, height, and ratio of width to height of the lips as auxiliary information.
[0037] Alternatively, the auxiliary information extraction unit 203 may extract unique information (for example, a place name) that is assumed to be spoken by the person in each image data, based on the movement of the person's lips.
[0038] Auxiliary information is information obtained from speech image data and is arbitrary information that assists speech recognition. In other words, auxiliary information is not limited to information obtained from the area in which the speaker is pictured. For example, information obtained from image data of buildings, objects (products), etc. that are pictured at the same time as the speaker is also included as auxiliary information. For example, information about a specific place or a specific product can also be given as an example of auxiliary information. For example, the place where the speaker is speaking (office, station, home, etc.) or the name of the product the speaker is talking about can be given as an example of auxiliary information.
[0039] The auxiliary information extraction unit 203 extracts auxiliary information that includes at least one of the following: the speaker's spoken language, the vowel sequence uttered by the speaker, the phoneme sequence uttered by the speaker, information about the shape of the speaker's lips, information about the speaker's emotions, and unique information obtained from at least one image data. Auxiliary information can also be understood as information that a person can explicitly specify as the extraction target for the learning model. In other words, auxiliary information is information whose content a person can recognize.
[0040] For example, the auxiliary information extraction unit 203 extracts auxiliary information from the speech image data set using a pre-implemented learning model. The Transformer model is an example of a learning model used by the auxiliary information extraction unit 203. However, this does not mean that the learning model used by the auxiliary information extraction unit 203 is limited to the Transformer model. The auxiliary information extraction unit 203 may extract auxiliary information using other learning models.
[0041] Furthermore, the administrator of the information processing system generates a learning model for extracting the auxiliary information using a large amount of training data consisting of combinations of image data and auxiliary information extracted from said image data, and implements it in the recognition device 10.
[0042] The auxiliary information extraction unit 203 outputs the auxiliary information extracted from the speech image data group to the speech text estimation unit 205.
[0043] For example, if the user says "Good morning," the auxiliary information extraction unit 203 extracts a string of characters (vowel sequence) consisting of "o, a, o, u" as auxiliary information. The auxiliary information extraction unit 203 then passes the extracted string of characters (vowel sequence) to the speech text estimation unit 205.
[0044] The face feature information extraction unit 204 is a means for extracting face feature information regarding all or part of the speaker's face from at least one or more image data. More specifically, the face feature information extraction unit 204 extracts face feature information from the group of utterance image data.
[0045] For example, the face feature information extraction unit 204 extracts face feature information using a learning model. The face feature information extraction unit 204 extracts, as face feature information, the information that the learning model spontaneously acquires through learning. For example, the face feature information extraction unit 204 treats the information implicitly extracted by a deep learning model as face feature information. Note that the techniques described in Non-Patent Documents 1 and 2 can be applied to the process of extracting face feature information (feature amounts around the lips) using a deep learning model. Therefore, a detailed description of the extraction of face feature information using a deep learning model is omitted.
[0046] The face feature information extraction unit 204 extracts one piece of face feature information from each of the image data that make up the group of utterance image data. For example, when the group of utterance image data includes 10 pieces of image data, the face feature information extraction unit 204 extracts 10 pieces of face feature information.
[0047] For example, the face feature information extraction unit 204 extracts face feature information from the group of utterance image data using a pre-implemented learning model. As a learning model used by the face feature information extraction unit 204, a CNN (Convolutional Neural Network) model is exemplified. However, it is not intended to limit the learning model used by the face feature information extraction unit 204 to a CNN model. The auxiliary information extraction unit 203 may extract face feature information using other learning models.
[0048] Note that an administrator or the like of the information processing system generates the learning model for extracting the face feature information using a large number of teacher data consisting of a combination of the image data group and the face feature information extracted from the image data group, and installs it in the recognition device 10.
[0049] The face feature information extraction unit 204 outputs a plurality of face feature information extracted from the group of utterance image data to the utterance sentence estimation unit 205.
[0050] For example, if the user says "Good morning" as described above, the facial feature information extraction unit 204 passes the facial feature information corresponding to each image data included in the speech image data group to the speech text estimation unit 205.
[0051] The speech text estimation unit 205 is a means for estimating the speech text (speech content) spoken by a speaker using auxiliary information extracted by the auxiliary information extraction unit 203 and facial feature information extracted by the facial feature information extraction unit 204.
[0052] The speech text estimation unit 205 estimates the text spoken by the person captured in the image data included in the speech image data group when the auxiliary information and facial feature information are available.
[0053] For example, the speech text estimation unit 205 uses a pre-implemented learning model to predict the speech text (estimate the content of the speech). A large-scale language model (LLM) is given as an example of a learning model used by the speech text estimation unit 205. However, this does not mean that the learning model used by the speech text estimation unit 205 is limited to a large-scale language model. The auxiliary information extraction unit 203 may use other learning models to estimate the speech text.
[0054] Furthermore, the administrator of the information processing system may generate a new large-scale language model from a vast amount of text data, etc., and implement the generated large-scale language model in the recognition device 10. Alternatively, an existing large-scale language model may be implemented in the recognition device 10.
[0055] Alternatively, a learning model for utterance text estimation may be generated by utilizing an existing (large-scale language model). For example, transfer learning (fine-tuning) may be performed, in which the weights of a previously generated and trained model are trained with new training data. Specifically, a learning model for utterance text estimation may be generated by utilizing an existing large-scale language model and performing additional training with a unique training data set.
[0056] Once the auxiliary information and facial feature information are acquired, the speech text estimation unit 205 inputs the acquired auxiliary information and facial feature information into a large-scale language model using various methods.
[0057] For example, the speech sentence estimation unit 205 generates a prompt using the acquired auxiliary information and inputs the generated prompt and tokenized facial feature information into the large-scale language model. For example, the speech sentence estimation unit 205 generates a prompt such as "Please estimate the speech sentence based on the auxiliary information described below; Auxiliary information: O, A, O, U" and inputs the generated prompt and vectorized facial feature information into the large-scale language model.
[0058] Alternatively, the speech sentence estimation unit 205 may generate a prompt such as "Please estimate the speech sentence based on the following vowel sequence; Vowel sequence: o, a, o, u" and input it into the large-scale language model along with the vectorized facial feature information.
[0059] The speech text estimation unit 205 inputs the generated prompt and vectorized facial feature information into a large-scale language model to obtain the estimation result of the speech text spoken by the person in each image data included in the speech image data set. The speech text estimation unit 205 obtains text data containing the speech content (speech text) from the large-scale language model.
[0060] In this way, the speech text estimation unit 205 generates a prompt that includes the extracted auxiliary information, and by inputting the generated prompt and vectorized facial feature information into a large-scale language model, it can obtain the estimation result of the speaker's speech text.
[0061] Alternatively, the speech text estimation unit 205 may directly input tokenized auxiliary information and facial feature information into the large-scale language model instead of using prompts. In other words, the speech text estimation unit 205 may obtain the speech text from the large-scale language model without using prompts. For example, in order to predict the start of speech by the speaker, the speech text estimation unit 205 may input tokenized auxiliary information and facial feature information into the large-scale language model.
[0062] The speech text estimation unit 205 may also use the prediction results from the large-scale language model as a prompt. More specifically, the speech text estimation unit 205 may estimate the speech text by recursively inputting the already obtained speech text (predicted speech text) as a prompt into the large-scale language model. In this case, the speech text estimation unit 205 can input auxiliary information and facial feature information into the large-scale language model in any form.
[0063] Furthermore, the tokenization of facial feature information and the input of tokenized facial feature information into a large-scale language model are disclosed in Non-Patent Documents 1 and 2, and the tokenization of facial feature information can be performed by applying the technologies disclosed in Non-Patent Documents 1 and 2. Therefore, a detailed explanation regarding the tokenization of facial feature information will be omitted.
[0064] Furthermore, Reference 1 below discloses a technique for treating the features of the entire face as tokens, and it is also possible to tokenize facial feature information by applying the technique disclosed in Reference 1.
[0065] <Reference 1> Arc2Face: A Foundation Model for ID-Consistent Human Faces <URL: https: / / arxiv.org / abs / 2403.11641>
[0066] The speech text estimation unit 205 passes the estimation result (speech text) obtained from the large-scale language model to the estimation result output unit 206.
[0067] The estimation result output unit 206 is a means for outputting the speech text estimated by the speech text estimation unit 205. For example, the estimation result output unit 206 transmits text data containing the speech text (estimation result) to the terminal 20.
[0068] The memory unit 207 is a means for storing information necessary for the operation of the recognition device 10.
[0069] The recognition device 10 may also include a processing module or the like that enables the user to select a conversation partner.
[0070] The operation of the recognition device 10 is summarized in the flowchart shown in Figure 6.
[0071] The image data acquisition unit 202 acquires image data (speech image data group) from the user's terminal 20 (step S01). The image data acquisition unit 202 then passes the acquired image data to the auxiliary information extraction unit 203 and the facial feature information extraction unit 204.
[0072] The auxiliary information extraction unit 203 extracts auxiliary information from the speech image data group (step S02). The auxiliary information extraction unit 203 then passes the extracted auxiliary information to the speech text estimation unit 205.
[0073] The facial feature information extraction unit 204 extracts facial feature information from the speech image data group (step S03). The facial feature information extraction unit 204 then passes the extracted facial feature information to the speech text estimation unit 205.
[0074] The auxiliary information extraction process by the auxiliary information extraction unit 203 and the facial feature information extraction process by the facial feature information extraction unit 204 are executed in parallel. Alternatively, either the auxiliary information extraction process or the facial feature information extraction process may be executed first.
[0075] When the auxiliary information and facial feature information obtained from the same speech image data set are available, the speech text estimation unit 205 performs speech recognition (step S04). For example, the speech text estimation unit 205 generates a prompt using the auxiliary information and inputs the generated prompt and tokenized facial feature information into the large-scale language model. The speech text estimation unit 205 then passes the speech text (estimation result) obtained from the large-scale language model to the estimation result output unit 206.
[0076] The estimation result output unit 206 outputs text data including the spoken sentence (outputs the spoken sentence; step S05).
[0077] [Terminal] A detailed explanation of terminal 20 is omitted. Examples of terminal 20 include mobile devices such as smartphones, mobile phones, game consoles, and tablets, as well as computers (personal computers, laptops), etc. Terminal 20 can be any device or equipment as long as it can receive user input and communicate with the recognition device 10.
[0078] Next, a modified example of the first embodiment will be described.
[0079] <Modification 1> The auxiliary information extraction unit 203 may perform multiple auxiliary information extraction processes. The auxiliary information extraction unit 203 may perform multiple auxiliary information extraction processes and output at least one of the multiple pieces of auxiliary information obtained by performing the multiple auxiliary information extraction processes to the speech text estimation unit 205 as extracted auxiliary information.
[0080] For example, the auxiliary information extraction unit 203 may execute multiple auxiliary information extraction processes in succession (see Figure 7).
[0081] For example, the auxiliary information extraction unit 203 uses the speech image data to extract the user's spoken language (language used) in the first auxiliary information extraction process 211. Subsequently, in the auxiliary information extraction process 212, the auxiliary information extraction unit 203 uses the spoken language and the speech image data to extract a vowel sequence.
[0082] The auxiliary information extraction unit 203 outputs the auxiliary information (vowel sequence) obtained by the execution of the auxiliary information extraction process 212 to the speech text estimation unit 205.
[0083] Thus, the auxiliary information extraction unit 203 may execute multiple auxiliary information extraction processes in succession and output the auxiliary information obtained from the last executed auxiliary information extraction process to the speech text estimation unit 205 as extracted auxiliary information.
[0084] <Modification 2> Alternatively, the auxiliary information extraction unit 203 may execute multiple auxiliary information extraction processes in parallel (see Figure 8).
[0085] The auxiliary information extraction unit 203 may output multiple pieces of auxiliary information obtained by each auxiliary information extraction process to the speech text estimation unit 205. For example, the auxiliary information extraction unit 203 may output the spoken language obtained by the auxiliary information extraction process 213 and the vowel sequence obtained by the auxiliary information extraction process 214 to the speech text estimation unit 205.
[0086] Thus, the auxiliary information extraction unit 203 may execute each of the multiple auxiliary information extraction processes and output the multiple pieces of auxiliary information obtained by each auxiliary information extraction process to the speech text estimation unit 205 as extracted auxiliary information. That is, the auxiliary information extraction unit 203 may extract multiple pieces of auxiliary information (multiple types of auxiliary information) from the speech image data group and output the extracted multiple pieces of auxiliary information to the speech text estimation unit 205.
[0087] In this case, the speech sentence estimation unit 205 can generate a prompt using multiple pieces of auxiliary information. For example, the speech sentence estimation unit 205 can generate a prompt such as, "The speaker is speaking in Japanese. Please estimate the sentence based on the following vowel sequence; Vowel sequence: o, a, o, u".
[0088] <Modification 3> The auxiliary information extraction unit 203 may also have a function to select an auxiliary information extraction process to execute from among a plurality of auxiliary information extraction processes (see Figure 9).
[0089] The auxiliary information extraction unit 203 generates selection basis information for selecting an auxiliary information extraction process to be executed from among multiple auxiliary information extraction processes by performing a selection process 221 using the speech image data group.
[0090] For example, the auxiliary information extraction unit 203 obtains the speaker's "spoken language" as selection basis information by executing the selection process 221. Based on the generated selection basis information, the auxiliary information extraction unit 203 determines which auxiliary information extraction process to execute from among a plurality of auxiliary information extraction processes.
[0091] For example, if the spoken language is "Japanese," the auxiliary information extraction unit 203 selects an auxiliary information extraction process 215 that extracts vowel sequences as auxiliary information. Alternatively, if the spoken language is "English," the auxiliary information extraction unit 203 selects an auxiliary information extraction process 216 that extracts phoneme sequences as auxiliary information.
[0092] The auxiliary information extraction unit 203 outputs the auxiliary information obtained by executing the selected auxiliary information extraction process to the speech text estimation unit 205.
[0093] In this way, the auxiliary information extraction unit 203 generates information for selecting an auxiliary information extraction process to be executed from among a plurality of auxiliary information extraction processes using at least one image data. The auxiliary information extraction unit 203 may select an auxiliary information extraction process to be executed from among the plurality of auxiliary information extraction processes based on the generated information (selection base information).
[0094] Note that each of the auxiliary information extraction processes shown in Figures 7 to 9 may be executed using a different learning model. Using different learning models allows for the efficient generation of learning models specialized for each process. Furthermore, using learning models specialized for each process in the auxiliary information extraction process results in highly accurate auxiliary information.
[0095] <Modification 4> The auxiliary information extraction unit 203 may determine (select) the auxiliary information to output to the speech text estimation unit 205 based on the confidence level obtained when executing multiple auxiliary information extraction processes (the confidence level of the results output by the learning model corresponding to each auxiliary information extraction process). For example, the auxiliary information extraction unit 203 may output to the speech text estimation unit 205 auxiliary information with a high confidence level among the spoken language and vowel sequences.
[0096] Thus, each of the multiple auxiliary information extraction processes outputs auxiliary information using a learning model, and the auxiliary information extraction unit 203 may select the auxiliary information to output to the speech text estimation unit 205 based on the confidence level of the auxiliary information output by each of the multiple auxiliary information extraction processes.
[0097] <Variation 5> The speech image data set may include multiple people. The recognition device 10 may identify the speaker by biometric authentication (facial recognition) and estimate the spoken text for each speaker.
[0098] In this case, the image data acquisition unit 202 extracts each of the multiple face regions (face images) contained in the acquired image data.
[0099] Note that existing technologies can be used for the face region extraction process by the image data acquisition unit 202, so a detailed explanation will be omitted. For example, the image data acquisition unit 202 may extract face regions from image data using a learning model trained by a CNN (Convolutional Neural Network). Alternatively, the image data acquisition unit 202 may extract face regions using methods such as template matching.
[0100] The image data acquisition unit 202 assigns an ID to the extracted face image (speaker) and stores the extracted face image. The image data acquisition unit 202 extracts a face image from the image data acquired from the terminal 20 and performs a matching process (authentication process) using the extracted face image and the stored face image.
[0101] If the image data acquisition unit 202 succeeds in the matching process (i.e., if the extracted face image is one that has been previously extracted), it manages the extracted face image along with its ID. If the matching process fails (i.e., a new face image is extracted), the image data acquisition unit 202 assigns an ID to the new face image and manages (stores) it.
[0102] Furthermore, the image data acquisition unit 202 processes the image data acquired from the terminal 20 (each image data included in the speech image data group) so that the speech recognition model can identify the speaker. For example, the image data acquisition unit 202 writes an ID near the extracted face region. Alternatively, the image data acquisition unit 202 may surround the face region with a dotted line and write an ID near it (see Figure 10).
[0103] The speech recognition model (auxiliary information extraction unit 203, facial feature information extraction unit 204, and speech text estimation unit 205) executes the auxiliary information extraction process, facial feature information extraction process, and speech text estimation process described above for each ID attached to the image data.
[0104] The speech text estimation unit 205 passes the estimation result for each ID to the estimation result output unit 206. For example, the speech text estimation unit 205 outputs text data such as "ID 11: Good morning." and "ID 12: It's a nice day today." to the terminal 20. The terminal 20 should then display the content of the acquired text data, treating different IDs as utterances from different speakers (see Figure 11).
[0105] The auxiliary information extraction unit 203 may extract auxiliary information applicable to the estimation of the spoken sentences of each of the multiple people captured in the image data. In this case, the auxiliary information extraction unit 203 may determine the auxiliary information that is commonly applied to the estimation of each person's speech based on the confidence level of the auxiliary information corresponding to each person. For example, if auxiliary information such as "Spoken language: Japanese, Confidence level: 40%" is extracted for user A, and auxiliary information such as "Spoken language: English, Confidence level: 80%" is extracted for user B, the auxiliary information extraction unit 203 may adopt the latter auxiliary information.
[0106] Thus, the image data may contain multiple people, and the recognition device 10 may associate the speaker with the spoken text by applying a speech recognition model to each person and output the result.
[0107] As described above, the recognition device 10 according to the first embodiment estimates spoken text using image data of the speaker without using the speaker's voice. Therefore, the recognition device 10 can solve various problems that may arise in speech-based speech recognition models.
[0108] For example, speech recognition models that rely primarily on speech data have difficulty accurately estimating spoken text in noisy environments. However, the recognition device 10 disclosed in this application does not use speech data, so it can estimate spoken text regardless of the amount of noise.
[0109] Alternatively, in systems that utilize speech-based speech recognition models, it is necessary to prepare highly directional microphones or similar devices to accurately collect the speaker's voice. However, since the recognition device 10 disclosed in this application does not use voice data, such special devices are unnecessary.
[0110] Furthermore, speech recognition models that rely primarily on voice have limitations, such as being unusable by people with speech impairments. However, the recognition device 10 disclosed in this application does not use voice data, so there are no restrictions regarding the user.
[0111] Some speech recognition models that use speech also utilize images. However, even in such speech recognition models, images only play a role in making the estimation results robust to ambient noise. After diligent research by the inventors, it has been found that if speech data is not input into such speech recognition models, the accuracy of speech recognition decreases significantly. In other words, with existing speech recognition models, speech prediction without using speech is difficult, and speech assistance is indispensable.
[0112] The recognition device 10 disclosed in this application utilizes auxiliary information obtained from image data (information obtained from image data other than facial feature information such as face shape) as a substitute for spoken speech, thereby achieving speech prediction without using speech data. Here, the input to the auxiliary information extraction unit 203 is image data, and the auxiliary information is extracted internally (in the learning model) of the auxiliary information extraction unit 203, so the recognition device 10 can perform speech recognition without requiring additional external information such as speech data.
[0113] Furthermore, in speech recognition, auxiliary information acts as a powerful "inductive bias," enabling accurate estimation of spoken text even without the use of audio data. In other words, auxiliary information acts as prior knowledge that a person provides to the learning model as a point of focus for the information contained in the image data. For example, if spoken language is provided as auxiliary information, the learning model (large-scale language model) can interpret facial feature information (e.g., changes in lip shape) on the premise that the speaker is speaking in the acquired language. Therefore, auxiliary information functions as a substitute for audio data, and the recognition device 10 can provide highly accurate estimation results of spoken text.
[0114] Furthermore, the recognition device 10 performs auxiliary information extraction, facial feature information extraction, and speech text estimation using learning models specialized for each of these processes. Therefore, learning to obtain each learning model is performed efficiently. In addition, since a learning model is prepared for each process, updating the learning models is easily done.
[0115] Next, we will describe the hardware of each device that makes up the information processing system. Figure 12 shows an example of the hardware configuration of the recognition device 10.
[0116] The recognition device 10 can be configured using an information processing device (a so-called computer), and has the configuration illustrated in Figure 12. For example, the recognition device 10 includes a processor 311, a memory 312, an input / output interface 313, and a communication interface 314, etc. The components of the processor 311, etc., are connected by an internal bus or the like and are configured to communicate with each other.
[0117] However, the configuration shown in Figure 12 is not intended to limit the hardware configuration of the recognition device 10. The recognition device 10 may include hardware not shown, and it may not have to include the input / output interface 313 if necessary. Also, the number of processors 311 etc. included in the recognition device 10 is not limited to the example in Figure 12; for example, multiple processors 311 may be included in the recognition device 10.
[0118] The processor 311 is a programmable device such as a CPU (Central Processing Unit), MPU (Micro Processing Unit), DSP (Digital Signal Processor), TPU (Tensor Processing Unit), or GPU (Graphics Processing Unit). Alternatively, the processor 311 may be a device such as an FPGA (Field Programmable Gate Array) or ASIC (Application Specific Integrated Circuit). The processor 311 executes various programs, including an operating system (OS).
[0119] Memory 312 can be RAM (Random Access Memory), ROM (Read Only Memory), HDD (Hard Disk Drive), SSD (Solid State Drive), etc. Memory 312 stores the OS program, application programs, and various data.
[0120] The input / output interface 313 is an interface for a display device or input device (not shown). The display device is, for example, a liquid crystal display. The input device is, for example, a device that accepts user input such as a keyboard or mouse.
[0121] The communication interface 314 is a circuit, module, etc., that communicates with other devices. For example, the communication interface 314 may include a NIC (Network Interface Card).
[0122] The functions of the recognition device 10 are realized by various processing modules. These processing modules are realized, for example, by the processor 311 executing a program stored in the memory 312. The program can be recorded on a computer-readable storage medium. The storage medium can be a non-transitory material such as a semiconductor memory, hard disk, magnetic recording medium, or optical recording medium. In other words, the present invention can also be embodied as a computer program product. Furthermore, the program can be downloaded via a network or updated using the storage medium on which the program is stored. Moreover, the processing module may be realized by a semiconductor chip.
[0123] In addition, the terminal 20 can also be configured using an information processing device, similar to the recognition device 10, and its basic hardware configuration is no different from that of the recognition device 10, so its explanation will be omitted.
[0124] The recognition device 10, which is an information processing device, is equipped with a computer, and its functions can be realized by having the computer execute a program. Furthermore, the recognition device 10 executes a control method for the recognition device 10 (information processing device) through this program.
[0125] [Modification] Note that the configuration and operation of the information processing system described in the above embodiment are illustrative examples and are not intended to limit the system configuration.
[0126] In the above embodiment, the case in which the recognition device 10 estimates the speaker's utterance was described. However, the estimation of the utterance may be performed by the terminal 20. That is, the terminal 20 may include the functions of the recognition device 10. In this case, users converse with other users via the terminals 20 that they each possess.
[0127] In the above embodiment, the case was described in which the recognition device 10 transmits the spoken text (estimated result of the spoken content) to a terminal 20 different from the terminal 20 that photographs the user's face. However, the recognition device 10 may also transmit the spoken text to the terminal 20 that photographs the user's face.
[0128] The recognition device 10 may use multiple types of facial feature information to estimate the spoken text. For example, the facial feature information extraction unit 204 may extract facial feature information related to the shape of the eyes and facial feature information related to the shape of the mouth (lips) from the speech image data group. The spoken text estimation unit 205 may use this multiple types of facial feature information to estimate the spoken text.
[0129] The recognition device 10 may feed back the utterance estimation result to the extraction of auxiliary information or facial feature information. For example, if the utterance estimation unit 205 determines that the confidence level of the spoken utterance estimated using a large-scale language model is lower than a predetermined threshold, it may instruct the auxiliary information extraction unit 203 to provide different auxiliary information or to add auxiliary information. For example, if it is determined that the confidence level of the estimation result using a vowel sequence is low, the auxiliary information extraction unit 203 may provide the utterance language to the utterance estimation unit 205 as additional auxiliary information.
[0130] The speech-to-text estimation unit 205 may store the results of previous speech estimation and use the stored results to estimate the speech-to-text. For example, consider a case where the speaker says "Good morning" followed by "It's nice weather." In this case, the speech-to-text estimation unit 205 may input a prompt such as "Please estimate the speech-to-text following 'Good morning' based on the following vowel sequence; Vowel sequence: i, i, e, ..." along with tokenized facial feature information to a large-scale language model.
[0131] The recognition device 10 may transmit not only the estimated result of the spoken text but also the confidence level of the estimated result to the terminal 20. In this case, the terminal 20 may change the manner in which it presents the estimated result (spoken text) to the user according to the acquired confidence level. For example, the terminal 20 may change the size, color, font, etc. of the characters according to the confidence level.
[0132] The recognition device 10 may transmit the estimated speech text and image data of the speaker to the terminal 20. In the example in Figure 4, the recognition device 10 may transmit image data of user A and the speech text of user A to user B's terminal 20.
[0133] In the above embodiment, the recognition device 10 was described in a case where it estimates the spoken text using the speaker's voice (voice data). However, the recognition device 10 may also estimate the spoken text using the speaker's voice data.
[0134] The recognition device 10 or terminal 20 may convert the speaker's spoken text (text data) into speech data and output it.
[0135] The above embodiment described a case in which multiple learning models are implemented in the recognition device 10. However, all or part of these multiple learning models may be implemented on an external server or the like, different from the recognition device 10. In this case, the recognition device 10 only needs to send the data to be input to the learning model to the external server and obtain the output results of the learning model from the external server.
[0136] In the flowcharts (sequence diagrams) used in the above description, multiple processes are shown in order, but the execution order of the processes performed in the embodiment is not limited to the order in which they are shown. In the embodiment, the order of the illustrated processes can be changed to the extent that it does not impede the content, for example, by executing each process in parallel.
[0137] The embodiments described above are explained in detail to facilitate understanding of the disclosure, and it is not intended that all the configurations described above are necessary. Furthermore, when multiple embodiments are described, each embodiment may be used individually or in combination. For example, it is possible to replace parts of the configuration of one embodiment with those of another embodiment, or to add configurations from other embodiments to the configuration of one embodiment. In addition, it is possible to add, delete, or replace parts of the configuration of one embodiment with those of another.
[0138] As described above, the industrial applicability of the present invention is clear, and it is particularly applicable to information processing systems that estimate the spoken text of a speaker without using the speaker's voice.
[0139] Some or all of the above embodiments may also be described as follows, but are not limited to the following:
[0140] [Note 1] An information processing device comprising: an image data acquisition means for acquiring at least one image data showing a speaker; an auxiliary information extraction means for extracting auxiliary information from the at least one image data to assist in estimating the spoken sentence by the speaker; a facial feature information extraction means for extracting facial feature information relating to all or part of the speaker's face from the at least one image data; a spoken sentence estimation means for estimating the spoken sentence by the speaker using the extracted auxiliary information and facial feature information; and an estimation result output means for outputting the estimated spoken sentence.
[0141] [Note 2] The information processing device described in Note 1, wherein the speech sentence estimation means obtains the estimation result of the speech sentence spoken by the speaker by inputting the extracted auxiliary information and facial feature information into a large-scale language model.
[0142] [Note 3] The information processing apparatus according to Note 1, wherein the auxiliary information extraction means performs a plurality of auxiliary information extraction processes and outputs at least one of the plurality of auxiliary information obtained by performing the plurality of auxiliary information extraction processes as the extracted auxiliary information.
[0143] [Note 4] The information processing apparatus according to Note 3, wherein the auxiliary information extraction means performs the plurality of auxiliary information extraction processes by extracting first auxiliary information and then extracting second auxiliary information using the first auxiliary information, and outputs the second auxiliary information as the extracted auxiliary information.
[0144] [Note 5] The information processing apparatus according to Note 3, wherein the auxiliary information extraction means executes each of the plurality of auxiliary information extraction processes and outputs the plurality of auxiliary information obtained by each auxiliary information extraction process as extracted auxiliary information.
[0145] [Appendix 6] The information processing apparatus according to Appendix 3, wherein each of the plurality of auxiliary information extraction processes outputs the auxiliary information using a learning model, and the auxiliary information extraction means selects the auxiliary information to be output based on the confidence level of the auxiliary information output by each of the plurality of auxiliary information extraction processes.
[0146] [Note 7] The information processing apparatus according to Note 1, wherein the auxiliary information extraction means generates information for selecting an auxiliary information extraction process to be executed from among a plurality of auxiliary information extraction processes using the at least one image data, and selects an auxiliary information extraction process to be executed from among the plurality of auxiliary information extraction processes based on the generated information.
[0147] [Note 8] The information processing apparatus according to any one of Notes 1 to 7, wherein the auxiliary information includes at least one of the following: the spoken language of the speaker, the vowel sequence uttered by the speaker, the phoneme sequence uttered by the speaker, information regarding the shape of the speaker's lips, information regarding the speaker's emotions, and unique information obtained from the at least one image data.
[0148] [Note 9] A control method for an information processing device comprising: an image data acquisition step of acquiring at least one image data showing a speaker; an auxiliary information extraction step of extracting auxiliary information from the at least one image data to assist in estimating the spoken sentence by the speaker; a face feature information extraction step of extracting face feature information relating to all or part of the speaker's face from the at least one image data; a spoken sentence estimation step of estimating the spoken sentence by the speaker using the extracted auxiliary information and face feature information; and an estimation result output step of outputting the estimated spoken sentence.
[0149] [Note 10] The control method for the information processing device described in Note 9, wherein the speech content estimation step involves inputting the extracted auxiliary information and facial feature information into a large-scale language model to obtain the estimation result of the spoken text by the speaker.
[0150] [Note 11] The control method for the information processing device described in Note 9, wherein the auxiliary information extraction step performs a plurality of auxiliary information extraction processes, and outputs at least one of the plurality of auxiliary information obtained by the execution of the plurality of auxiliary information extraction processes as the extracted auxiliary information.
[0151] [Note 12] The control method for the information processing device described in Note 11, wherein the auxiliary information extraction step performs the multiple auxiliary information extraction processes by extracting first auxiliary information and then extracting second auxiliary information using the first auxiliary information, and outputs the second auxiliary information as the extracted auxiliary information.
[0152] [Note 13] The control method for the information processing apparatus described in Note 11, wherein the auxiliary information extraction step executes each of the plurality of auxiliary information extraction processes and outputs the plurality of auxiliary information obtained by each auxiliary information extraction process as extracted auxiliary information.
[0153] [Note 14] The control method for the information processing device described in Note 11, wherein each of the plurality of auxiliary information extraction processes outputs the auxiliary information using a learning model, and the auxiliary information extraction step selects the auxiliary information to be output based on the confidence level of the auxiliary information output by each of the plurality of auxiliary information extraction processes.
[0154] [Note 15] The control method for the information processing device described in Note 9, wherein the auxiliary information extraction step generates information for selecting an auxiliary information extraction process to be executed from among a plurality of auxiliary information extraction processes using the at least one image data, and selects an auxiliary information extraction process to be executed from among the plurality of auxiliary information extraction processes based on the generated information.
[0155] [Note 16] The control method for an information processing apparatus according to any one of Notes 9 to 15, wherein the auxiliary information includes at least one of the following: the spoken language of the speaker, the vowel sequence uttered by the speaker, the phoneme sequence uttered by the speaker, information regarding the shape of the speaker's lips, information regarding the speaker's emotions, and unique information obtained from the at least one image data.
[0156] [Note 17] A computer-readable storage medium that stores a program for causing a computer to execute: an image data acquisition process for acquiring at least one image data of a speaker; an auxiliary information extraction process for extracting auxiliary information from the at least one image data to assist in estimating the spoken sentence by the speaker; a facial feature information extraction process for extracting facial feature information relating to all or part of the speaker's face from the at least one image data; a spoken sentence estimation process for estimating the spoken sentence by the speaker using the extracted auxiliary information and facial feature information; and an estimation result output process for outputting the estimated spoken sentence.
[0157] [Note 18] The speech content estimation process uses the storage medium described in Note 17, which inputs the extracted auxiliary information and facial feature information into a large-scale language model to obtain the estimation result of the spoken text by the speaker.
[0158] [Note 19] The storage medium according to Note 17, wherein the auxiliary information extraction process performs a plurality of auxiliary information extraction processes and outputs at least one of the plurality of auxiliary information obtained by the execution of the plurality of auxiliary information extraction processes as the extracted auxiliary information.
[0159] [Note 20] The storage medium described in Note 19, wherein the auxiliary information extraction process performs the multiple auxiliary information extraction processes by extracting first auxiliary information and then using the first auxiliary information to extract second auxiliary information, and outputs the second auxiliary information as the extracted auxiliary information.
[0160] [Note 21] The storage medium described in Note 19, wherein the auxiliary information extraction process executes each of the plurality of auxiliary information extraction processes and outputs the plurality of auxiliary information obtained by each auxiliary information extraction process as extracted auxiliary information.
[0161] [Note 22] The storage medium according to Note 19, wherein each of the plurality of auxiliary information extraction processes outputs the auxiliary information using a learning model, and the auxiliary information extraction process selects the auxiliary information to output based on the confidence level of the auxiliary information output by each of the plurality of auxiliary information extraction processes.
[0162] [Note 23] The storage medium according to Note 17, wherein the auxiliary information extraction process generates information for selecting an auxiliary information extraction process to be executed from among a plurality of auxiliary information extraction processes using at least one or more image data, and selects an auxiliary information extraction process to be executed from among the plurality of auxiliary information extraction processes based on the generated information.
[0163] [Note 24] The storage medium according to any one of Notes 17 to 23, wherein the auxiliary information includes at least one of the following: the spoken language of the speaker, the vowel sequence uttered by the speaker, the phoneme sequence uttered by the speaker, information regarding the shape of the speaker's lips, information regarding the speaker's emotions, and unique information obtained from the at least one image data.
[0164] Furthermore, some or all of the configurations described in Appendices 2 to 8, which are subordinate to Appendice 1 above, may also be subordinate to Appendices 9 and 17 in the same way as those described in Appendices 2 to 8. Moreover, not limited to Appendices 1, 9 and 17, some or all of the configurations described as appendices may also be subordinate to various hardware, software, various recording means for recording software, or systems, without departing from the embodiments described above.
[0165] Furthermore, each disclosure of the above-mentioned prior art documents cited herein is incorporated herein by reference. Although embodiments of the present invention have been described above, the present invention is not limited to these embodiments. It will be understood by those skilled in the art that these embodiments are merely illustrative and that various modifications are possible without departing from the scope and spirit of the present invention. That is, the present invention naturally includes the entire disclosure, including the claims, and various modifications and alterations that can be made by those skilled in the art in accordance with the technical idea.
[0166] 10 Recognition device 20 Terminal 100 Information processing device 101 Image data acquisition means 102 Auxiliary information extraction means 103 Face feature information extraction means 104 Speech content estimation means 105 Estimation result output means 201 Communication control unit 202 Image data acquisition unit 203 Auxiliary information extraction unit 204 Face feature information extraction unit 205 Speech text estimation unit 206 Estimation result output unit 207 Storage unit 211 Auxiliary information extraction process 212 Auxiliary information extraction process 213 Auxiliary information extraction process 214 Auxiliary information extraction process 215 Auxiliary information extraction process 216 Auxiliary information extraction process 221 Selection process 311 Processor 312 Memory 313 Input / output interface 314 Communication interface
Claims
1. An information processing device comprising: an image data acquisition means for acquiring at least one image data showing a speaker; an auxiliary information extraction means for extracting auxiliary information from the at least one image data to assist in estimating a spoken sentence by the speaker; a facial feature information extraction means for extracting facial feature information relating to all or part of the speaker's face from the at least one image data; a spoken sentence estimation means for estimating a spoken sentence by the speaker using the extracted auxiliary information and facial feature information; and an estimation result output means for outputting the estimated spoken sentence.
2. The information processing apparatus according to claim 1, wherein the speech text estimation means inputs the extracted auxiliary information and facial feature information into a large-scale language model to obtain an estimation result of a speech text uttered by the speaker.
3. The information processing apparatus according to claim 1, wherein the auxiliary information extraction means performs a plurality of auxiliary information extraction processes and outputs at least one of the plurality of auxiliary information obtained by performing the plurality of auxiliary information extraction processes as the extracted auxiliary information.
4. The information processing apparatus according to claim 3, wherein the auxiliary information extraction means performs the plurality of auxiliary information extraction processes by extracting first auxiliary information and then extracting second auxiliary information using the first auxiliary information, and outputs the second auxiliary information as the extracted auxiliary information.
5. The information processing apparatus according to claim 3, wherein the auxiliary information extraction means executes each of the plurality of auxiliary information extraction processes and outputs the plurality of auxiliary information obtained by each auxiliary information extraction process as extracted auxiliary information.
6. The information processing apparatus according to claim 3, wherein each of the plurality of auxiliary information extraction processes outputs the auxiliary information using a learning model, and the auxiliary information extraction means selects the auxiliary information to be output based on the confidence level of the auxiliary information output by each of the plurality of auxiliary information extraction processes.
7. The information processing apparatus according to claim 1, wherein the auxiliary information extraction means generates information for selecting an auxiliary information extraction process to be executed from among a plurality of auxiliary information extraction processes using the at least one image data, and selects an auxiliary information extraction process to be executed from among the plurality of auxiliary information extraction processes based on the generated information.
8. The information processing apparatus according to any one of claims 1 to 7, wherein the auxiliary information includes at least one of the following: the speaker's spoken language, the vowel sequence uttered by the speaker, the phoneme sequence uttered by the speaker, information regarding the shape of the speaker's lips, information regarding the speaker's emotions, and unique information obtained from the at least one image data.
9. A control method for an information processing device, comprising: an image data acquisition step of acquiring at least one image data showing a speaker; an auxiliary information extraction step of extracting auxiliary information from the at least one image data to assist in estimating the spoken sentence by the speaker; a facial feature information extraction step of extracting facial feature information relating to all or part of the speaker's face from the at least one image data; a spoken sentence estimation step of estimating the spoken sentence by the speaker using the extracted auxiliary information and facial feature information; and an estimation result output step of outputting the estimated spoken sentence.
10. A computer-readable storage medium that stores a program for causing a computer to execute: an image data acquisition process for acquiring at least one image data of a speaker; an auxiliary information extraction process for extracting auxiliary information from the at least one image data to assist in estimating the spoken sentence by the speaker; a facial feature information extraction process for extracting facial feature information relating to all or part of the speaker's face from the at least one image data; a spoken sentence estimation process for estimating the spoken sentence by the speaker using the extracted auxiliary information and facial feature information; and an estimation result output process for outputting the estimated spoken sentence.