Speech recognition device, speech recognition system, and speech recognition method

The speech recognition device uses a camera and microphone to enhance accuracy by correlating lip-reading with audio processing, addressing noise interference in public settings.

JP2026058433APending Publication Date: 2026-04-06ALSOK INC
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-25
Publication Date
2026-04-06

AI Technical Summary

Technical Problem

Existing speech recognition devices struggle to accurately recognize user speech in noisy environments due to variations in speech frequencies and noise levels, often filtering out valuable speech components.

Method used

A speech recognition device that combines a microphone and a camera to perform lip-reading and speech recognition, using lip-reading to estimate speech content from image data and filtering audio data to enhance speech recognition accuracy.

Benefits of technology

The device achieves high-accuracy speech recognition by correlating lip-reading results with audio data processing, effectively separating and identifying user speech amidst environmental noise.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026058433000001_ABST
    Figure 2026058433000001_ABST
Patent Text Reader

Abstract

This invention provides a speech recognition device and a speech recognition method that accurately recognize the voice of a user. [Solution] A speech recognition device that recognizes speech in order to interact with a user, comprising: a microphone for acquiring the user's voice; a camera for capturing images of the user; a control unit that identifies the user's speech content based on speech content obtained by performing lip-reading processing to estimate the content of a person's speech from the movement of the person's lips contained in the image data captured by the camera, and speech content obtained by performing speech recognition processing on the voice data acquired by the microphone.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a speech recognition device, a speech recognition system, and a speech recognition method for recognizing speech content from the speech of a user.

Background Art

[0002] Conventionally, a device that can recognize the speech uttered by a device user and interact with the user has been used. For example, Patent Document 1 discloses a security monitoring display device that recognizes the speech of a user and interacts with the user. The monitoring display device detects a user in front of the device, recognizes the speech, and realizes the interaction between the security guard character displayed on the display unit and the user.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] When performing speech recognition with a device installed in a public place as in the above prior art, speech processing for suppressing the influence of environmental noise may be performed. For example, a filtering process is performed to reduce or cut noise components included in the speech data collected by the microphone and extract the speech of the interlocutor who is the user of the device. However, due to differences in the speech frequencies of the interlocutors and changes in the noise generated around the device, simply filtering sounds in a predetermined frequency range may cut off part of the speech of the interlocutor or may not be able to reduce the noise. [[ID=�7]]

[0005] The present disclosure has been made in view of the prior art including the above problems, and one of its objects is to provide a speech recognition device, a speech recognition system, and a speech recognition method capable of accurately recognizing the speech of a user.

Means for Solving the Problems

[0006] The speech recognition device relating to this disclosure is a speech recognition device that recognizes speech in order to interact with a user, and comprises a microphone for acquiring the user's voice, a camera for capturing images of the user, and a control unit that identifies the user's speech content based on speech content obtained by performing lip-reading processing to estimate the content of a person's speech from the movement of the person's lips included in the image data captured by the camera, and speech content obtained by performing speech recognition processing on the voice data acquired by the microphone.

[0007] In the above configuration, the control unit may identify the user's voice included in the audio data acquired by the microphone based on the content of the person's speech obtained by performing the lip-reading process, and identify the content of the speech obtained by performing speech recognition processing on the identified voice as the content of the user's speech.

[0008] In the above configuration, the control unit may identify the user's speech content based on the correlation between the speech content of the person obtained by performing the lip-reading process and the speech content corresponding to each of the target voices obtained by performing a speech recognition process on one or more voices separated from the voice data acquired by the microphone.

[0009] In the above configuration, the control unit may set a filter that can extract audio corresponding to the speech content of the person obtained by performing the lip-reading process from the audio data acquired by the microphone, filter the audio data with the set filter to extract the audio, and then perform speech recognition processing on the extracted audio to identify the speech content obtained as the speech content of the user.

[0010] In the above configuration, the control unit may perform lip-reading processing to estimate the speech content of each person from the movement of the lips of each of the multiple people included in the image data captured by the camera, set a filter for each person that can extract the sound corresponding to the speech content of that person obtained by performing the lip-reading processing from the audio data acquired by the microphone, filter the audio data using the set filters to extract the sound corresponding to the speech content of each person, and perform speech recognition processing on each of the obtained sounds to identify the speech content of each person.

[0011] In the above configuration, the control unit may separate one or more voices from the voice data acquired by the microphone, perform voice recognition processing on each separated voice to obtain the speech content corresponding to each target voice, and then, based on the correlation between the speech content of the person obtained by performing the lip-reading processing, identify the user's voice from among the multiple voices and identify the speech content obtained from that voice as the user's speech content.

[0012] In the above configuration, the control unit may compare the correlation between the speech content of each person obtained by performing lip-reading processing to estimate the speech content of one or more people included in the image data captured by the camera, and the speech content corresponding to each voice obtained by performing speech recognition processing on one or more voices separated from the audio data acquired by the microphone, and for each person, determine that the voice in the combination with the highest correlation is the voice of that person, and identify the speech content obtained by performing speech recognition processing on that voice as the speech content of the user.

[0013] In the above configuration, the control unit may identify the user's voice based on the speech content obtained by performing lip-reading on a portion of image data extracted from image data captured by the camera at multiple time points in time, and the speech content obtained by performing speech recognition on a portion of audio data extracted from audio data acquired by the microphone, corresponding to the portion of the image data.

[0014] The speech recognition system relating to this disclosure is a speech recognition system that recognizes speech in order to interact with a user, and comprises a microphone for acquiring the user's voice, a camera for capturing images of the user, and an information processing device that identifies the user's speech content based on speech content obtained by performing lip-reading processing to estimate the content of a person's speech from the movement of the person's lips included in the image data captured by the camera, and speech content obtained by performing speech recognition processing on the voice data acquired by the microphone.

[0015] The speech recognition method relating to this disclosure is a speech recognition method for recognizing speech in order to interact with a user, and includes the steps of: performing lip-reading processing to estimate the content of a person's speech from the movement of the person's lips included in image data captured by a camera; performing speech recognition processing on speech data acquired by a microphone; and identifying the content of the user's speech based on the content of speech obtained by performing the lip-reading processing and the content of speech obtained by performing the speech recognition processing. [Effects of the Invention]

[0016] According to the speech recognition device, speech recognition system, and speech recognition method described herein, the voice of a user of a device including the speech recognition device can be recognized with high accuracy. [Brief explanation of the drawing]

[0017] [Figure 1] Figure 1 is a diagram illustrating the overview of the speech recognition device according to this embodiment. [Figure 2] Figure 2 is a block diagram showing an example of the functional configuration of a speech recognition device. [Figure 3] Figure 3 is a flowchart showing an example of the process performed by a speech recognition device. [Figure 4] Figure 4 is a diagram illustrating an example of the processing performed by a speech recognition device. [Figure 5] Figure 5 shows an example of a response displayed on the display unit by the voice recognition device. [Figure 6]FIG. 6 is a diagram for explaining another example of the process executed by the voice recognition device. [Figure 7] FIG. 7 is a diagram for explaining an example in the case where there are a plurality of interlocutors. MODE FOR CARRYING OUT THE INVENTION

[0018] Hereinafter, embodiments of a voice recognition device, a voice recognition system, and a voice recognition method according to the present disclosure will be described with reference to the accompanying drawings. The voice recognition device is an information processing device that executes a process of recognizing voice uttered by a device user. The usage method of the voice recognition device is not particularly limited. For example, the voice recognition device may output the voice recognition result to an external device that executes various processes using the voice recognition result, or the voice recognition device may execute another process using the voice recognition result. In the present embodiment, a case where the voice recognition device operates as a dialogue device that realizes a dialogue with a user by executing a voice recognition process to specify the utterance content of the user and executing a process of replying to the user will be described as an example.

[0019] FIG. 1 is a diagram for explaining the outline of the voice recognition device 10 according to the present embodiment. The voice recognition device 10 may be used alone, or a mode in which the voice recognition device 10 and the server device 1 connected to the voice recognition device 10 via the network 2 constitute a system may be used. For example, a mode in which the voice recognition device 10 cooperates with the server device 1 to realize voice recognition of the user 100 and dialogue with the user 100 may be used. For example, the voice recognition device 10 may specify the utterance content of the user 100, and the server device 1 may determine the reply content to the user 100. A mode in which a part of the process for specifying the utterance content is executed by the server device 1 may be used. The reply to the user 100 may be made by the voice recognition device 10 reproducing voice from a speaker, or may be made by displaying information indicating the reply on the display unit 40.

[0020] The voice recognition device 10 includes a microphone 20 and a camera 30. As shown in FIG. 1, the voice recognition device 10 uses the microphone 20 to acquire voice data 120 including the voice of the user 100 (A). The voice data 120 may include, in addition to the voice of the user 100, noise generated around the voice recognition device 10, the voice of another person speaking around the voice recognition device 10, and the like.

[0021] The voice recognition device 10 uses the camera 30 to acquire image data 130 obtained by imaging the user 100 (B). The camera 30 acquires image data 130 that images at least the oral cavity portion including the lips of the user 100. The image data 130 may be data including a moving image or data including still images captured continuously.

[0022] The voice recognition device 10 analyzes the image data 130 and executes a lip-reading process for estimating the speech content from the movement of the oral cavity portion of the user 100 (C). In the lip-reading process, the speech content of the user 100 is estimated by recognizing the movement of the lips of the user 100 from the images obtained by continuously imaging the oral cavity portion of the user 100.

[0023] For example, the voice recognition device 10 sets feature points that can specify the contour shape and opening degree of the lips, and based on the feature points, acquires a feature amount indicating the movement of the lips that changes in time series. The voice recognition device 10 estimates phonemes corresponding to the movement of the lips of the user 100 from the feature amounts. Although there may be a plurality of phonemes corresponding to the same lip shape, the voice recognition device 10 can estimate the speech content of the user 100 based on the dictionary data prepared in advance from the combination of phonemes estimated in time series. The processing content of the lip-reading process executed by the voice recognition device 10 is not particularly limited, and a conventionally known technique may be used. For example, the lip-reading process may be executed using an AI (artificial intelligence) technique such as a machine learning model or deep learning.

[0024] The speech recognition device 10 uses the content of the user's speech obtained from the image data 130 through lip-reading to extract the user's voice from the voice data 120 (D). In other words, the speech recognition device 10 performs a process to reduce or remove noise other than the user's voice from the voice data 120.

[0025] For example, a speech recognition device 10 that obtains the content of a user's speech from image data 130 through lip-reading processing sets a filter that can extract the speech that yields the speech recognition result closest to the content of the speech from the speech data 120, and performs filtering of the speech data 120 using the filter.

[0026] For example, the speech recognition device 10 filters the audio data 120 using multiple frequency filters with different frequency ranges to cut or reduce, performs speech recognition processing on the extracted audio, and selects the frequency filter that yields the speech recognition result closest to the utterance obtained by lip-reading processing. Multiple frequency filters can be pre-prepared. For example, multiple frequency filters with modified filtering frequency ranges can be prepared to cover the frequency range in which human voices can be obtained (e.g., 50-300 Hz). The process of selecting a frequency filter may be performed on a portion of the time interval of the audio data 120 and image data 130. For example, the speech content obtained by lip-reading processing on the first few seconds of data from the audio data 120 and image data 130 can be compared with the speech content obtained by applying each of the multiple frequency filters to the extracted audio data and performing speech recognition processing, and the frequency filter that yields the same or closest utterance can be selected. After selecting a frequency filter, the user's voice 100 can be extracted from the entire audio data 120 using the selected frequency filter.

[0027] Speech recognition processing is performed, for example, by recognizing phonemes from data containing speech and using dictionary data to identify the content of an utterance from a sequence of phonemes. While technologies such as DNN (Deep Neural Network) and HMM (Hidden Markov Model) are widely used in speech recognition processing, the processing content of the speech recognition device 10 is not particularly limited, and it is acceptable to perform speech recognition processing using AI technologies such as machine learning models and deep learning.

[0028] The method by which the speech recognition device 10 extracts the voice of user 100 is not limited to the filtering method described above. For example, the speech recognition device 10 may extract the voice of user 100 by performing a speech separation process that separates the voice data 120 into sounds from multiple sound sources, using AI technologies such as independent component analysis (ICA), machine learning, and deep learning. It should be noted that the sounds separated from the voice data 120 by filtering or speech separation processing may include sounds that do not contain human voices, in addition to sounds that do contain human voices. In this embodiment, sounds that do not contain human voices may also be described as voice or voice data.

[0029] When performing speech separation processing, the speech recognition device 10 can perform speech recognition processing on each speech separated from the speech data 120 and select the speech that yields the speech recognition result closest to the utterance of the user 100 obtained by lip-reading processing as the voice of the user 100. The process of selecting the speech may be performed on a portion of the speech separated from the speech data 120 and a portion of the image data 130. For example, the speech recognition result for the first few seconds of each speech separated from the speech data 120 can be compared with the lip-reading processing result for the corresponding few seconds of the image data 130 to extract the voice of the user 100 from among the multiple speeches separated from the speech data 120. In this embodiment, the speech recognition result is the utterance obtained by performing speech recognition processing on the speech data 120, and the lip-reading processing result is the utterance obtained by performing lip-reading processing on the image data 130.

[0030] The speech recognition device 10 performs speech recognition processing on the voice of user 100 to identify the content of user 100's speech (E).

[0031] When the speech recognition device 10 extracts the voice of user 100 from the voice data 120 by performing filtering using a frequency filter selected based on the lip-reading results, the speech recognition device 10 performs speech recognition processing on the extracted voice to identify the content of user 100's speech.

[0032] When extracting user 100's voice based on the correlation between the speech recognition results of multiple voices separated from voice data 120 by performing a voice separation process and the lip-reading results, the speech recognition device 10 identifies the speech recognition results of the extracted voice as the content of user 100's speech.

[0033] The speech recognition device 10, having identified the content of the user 100's speech, generates a response to the user 100 and displays it on the display unit 40. The generation of the response to the user 100's speech can be performed using conventional technologies such as AI. The speech recognition device 10 sequentially recognizes the content of the user 100's speech and displays the response on the display unit 40, thereby realizing a dialogue with the user 100.

[0034] In this way, the speech recognition device 10 extracts the voice of the user 100 from the voice data 120 and performs speech recognition processing using the content of the user's speech obtained by performing lip-reading processing on the image data 130. By using the image data 130 in addition to the voice data 120, the speech recognition device 10 can perform speech recognition processing with higher accuracy compared to when only the voice data 120 is used.

[0035] Figure 2 is a block diagram showing an example of the functional configuration of the speech recognition device 10. In addition to the microphone 20, camera 30, and display unit 40 shown in Figure 1, the speech recognition device 10 includes a control unit 50, a storage unit 60, and a communication unit 70.

[0036] The display unit 40 may only output and display information, or it may also function as an operation unit that receives information input from the user 100. For example, the display unit 40 may function as an operation display unit consisting of a touch panel type liquid crystal display device. For example, the voice recognition device 10 installed in a store or facility may function as a visitor reception device, and in addition to interacting with the user 100, it may receive instructions from the user 100 made by operating a touch panel and display the information requested by the user 100.

[0037] The memory unit 60 is a non-volatile memory device. Various data necessary for the operation of the speech recognition device 10 are stored in the memory unit 60. The memory unit 60 is used to store the voice data 120 obtained by the microphone 20 and the image data 130 obtained by the camera 30.

[0038] The communication unit 70 is used to send and receive data with external devices via the network 2. External devices include the server device 1 shown in Figure 1. When the speech recognition device 10 receives a question from the user 100, it may obtain the information necessary to answer the question from the external device via the communication unit 70 and display it on the display unit 40.

[0039] The control unit 50 controls each part of the speech recognition device 10. For example, a program corresponding to the control unit 50 is pre-stored in the storage unit 60 or a dedicated storage device, and the functions and operations of the control unit 50 are realized when this program is executed by hardware such as a CPU. The control unit 50 can control each part based on information acquired by the microphone 20 and camera 30, information transmitted and received via the communication unit 70, and information stored in the storage unit 60. This realizes the functions and operations of the speech recognition device 10 described in this embodiment. In this embodiment, the functions and operations of the speech recognition device 10 realized by the control unit 50 may be simply described as the functions and operations of the speech recognition device 10.

[0040] The control unit 50 includes an audio data processing unit 51, an image data processing unit 52, a lip-reading processing unit 53, an audio extraction unit 54, an audio recognition unit 55, and a dialogue processing unit 56. The audio data processing unit 51 acquires audio data 120 containing the voice of the user 100 using the microphone 20. The image data processing unit 52 acquires image data 130 of the user 100 captured by the camera 30. The lip-reading processing unit 53 performs lip-reading processing based on the image data 130 to acquire the content of the user 100's speech. The audio extraction unit 54 uses the content of the speech obtained by the lip-reading processing unit 53 to extract the user's voice from the audio data 120. The audio recognition unit 55 recognizes the speech extracted by the audio extraction unit 54. The dialogue processing unit 56 identifies the content of the user 100's speech based on the audio recognition result obtained by the audio recognition unit 55, generates a response to the user 100, and displays it on the display unit 40.

[0041] Figure 3 is a flowchart showing an example of speech recognition processing performed by the speech recognition device 10. The speech recognition device 10 acquires speech data 120 and image data 130 (step S1). The speech data processing unit 51 acquires the speech data 120 using the microphone 20 and stores it in the storage unit 60. The image data processing unit 52 acquires the image data 130 using the camera 30 and stores it in the storage unit 60.

[0042] The speech recognition device 10 performs processing to identify the voice of the user 100, i.e., the person speaking (step S2). Specifically, the speech recognition device 10 performs lip-reading processing and processing to select a frequency filter suitable for extracting the voice of the user 100. Based on the content of the user 100's speech obtained from the image data 130 by performing lip-reading processing and the content of the speech obtained from the voice data 120 by performing speech recognition processing, the speech recognition device 10 selects a frequency filter suitable for extracting the voice of the user 100.

[0043] The lip-reading processing unit 53 performs lip-reading on a portion of the data extracted from the image data 130 to estimate the content of the user's speech 100. The voice data processing unit 51 obtains multiple voices filtered with different frequency filters on a portion of the voice data 120 corresponding to the data to be lip-read. The voice recognition unit 55 performs voice recognition on each voice obtained through filtering to obtain the content of the speech. The voice extraction unit 54 compares the content of the speech obtained by the lip-reading processing unit 53 with the content of the speech obtained by the voice recognition unit 55 and selects the frequency filter that yields the voice with the highest correlation as the frequency filter suitable for extracting the user's voice 100. In this way, a frequency filter suitable for extracting the user's voice 100 is selected from among the multiple frequency filters, and the user's voice 100 is identified by this frequency filter.

[0044] The speech recognition device 10 extracts the voice of the user 100 from the voice data 120 obtained by the microphone 20 (step S3). The voice extraction unit 54 extracts the voice of the user 100 from the voice data 120 using the frequency filter selected in step S2. The selection of the frequency filter is performed by extracting a portion of the data from the voice data 120, such as the first few seconds, but the extraction of the user 100's voice using the selected frequency filter is performed on the entire voice data 120.

[0045] The speech recognition device 10 recognizes the voice of the user 100 (step S4). The speech recognition unit 55 performs speech recognition processing on the voice of the user 100 extracted by the speech extraction unit 54. Speech recognition processing is performed on the entire voice data 120. Based on the results of the speech recognition processing, the content of the user 100's speech is identified. After the content of the user 100's speech is identified by the speech recognition processing, the dialogue processing unit 56 performs dialogue processing to generate a response based on the content of the user 100's speech and display it on the display unit 40.

[0046] After selecting the optimal frequency filter for extracting the user 100's voice, the speech recognition device 10 applies the selected frequency filter to extract the user 100's voice each time the user 100 speaks, performs speech recognition, and displays the response to the identified utterance on the display unit 40. This enables the speech recognition device 10 to interact with the user 100.

[0047] Figure 3 shows an example of speech recognition processing performed by the speech recognition device 10, and does not limit the processing performed by the speech recognition device 10. For example, if the speech recognition device 10 performs speech separation processing to identify the voice of user 100, a different process will be performed than that shown in Figure 3.

[0048] Specifically, for example, audio data 120 and image data 130 are acquired (step S1), and the voice of user 100 is identified based on the speech recognition results of each voice obtained by the speech separation process and speech recognition process, and the lip-reading results (step S2). At this point, the voice of user 100 has already been separated from the audio data 120.

[0049] Therefore, steps S3 and S4 may not be executed, and the speech recognition result of the voice identified as belonging to user 100 may be identified as the content of user 100's speech. For example, if the entire voice data 120 is targeted at voice separation, and the speech recognition result obtained is compared with the lip-reading result to identify user 100's voice, steps S3 and S4 are not executed because user 100's voice has already been extracted and the speech recognition result has been obtained. On the other hand, for example, if voice separation is performed on the entire voice data 120, but only a portion of the obtained voice (e.g., a few seconds) is targeted at speech recognition and lip-reading to identify user 100's voice, step S3 is not executed because user 100's voice has already been extracted, but the speech recognition process in step S4 is executed on the entire voice of user 100, and the content of user 100's speech is identified.

[0050] Figure 4 is a diagram illustrating an example of the processing performed by the speech recognition device 10. The speech recognition device 10 performs lip-reading on the device user, i.e., the person speaking 101, using the image data 131 obtained by the camera 30. The speech recognition device 10 uses the obtained lip-reading results to extract the voice data 121a of the person speaking 101 from the voice data 121 obtained by the microphone 20.

[0051] As described above, filtering is performed using a frequency filter selected based on the lip-reading results, and the voice data 121a of the interlocutor 101 is extracted from the voice data 121 obtained by the microphone 20. If the voice data 121 contains the voice data 121a of the interlocutor 101 and other noise data 121b, these are separated by filtering, and the voice data 121a is extracted. As shown in Figure 4, in an environment where a car is driving on the road behind the interlocutor 101, the noise from the car is reduced from the voice data 121 obtained by the microphone 20, and the voice of the interlocutor 101 is extracted.

[0052] For example, if a person 101 asking for the location of a store says "mise no (of the store)...", lip-reading processing will recognize each phoneme from the movement of the lips, such as / miseno / (consonant / m / and vowel / i / for "mi", consonant / s / and vowel / e / for "se", and consonant / n / and vowel / o / for "no"), and obtain a lip-reading result that estimates the utterance to be "mise no". The speech recognition device 10 selects a frequency filter from the speech data 121 that will produce the same speech recognition result as the lip-reading result "mise no". The speech recognition device 10 filters the speech data 121 using the selected frequency filter to extract the speech data 121a of the person 101.

[0053] In lip-reading, the utterance "mise no" (store) of the interlocutor 101 may be misrepresented as a different sound, such as "ie no" ( / ieno / ), but the speech recognition device 10 is designed to handle this as well. For example, when comparing each sound in the lip-reading result and the speech recognition result, the speech recognition device 10 assigns an evaluation score to each sound according to its similarity: 1 point if both are "mi" ( / mi / ), 0.5 points if one is "mi" ( / mi / ) and the other is an "i" ( / i / ), or any other sound in the "i" row like "mi", and 0 points if one is "mi" ( / mi / ) and the other is a sound other than the "i" row. The speech recognition device 10 calculates the total evaluation score for each sound and evaluates the correlation between the lip-reading result and the speech recognition result based on the total value. This allows the speech of the interlocutor 101 to be extracted using the lip-reading result, even if the lip-reading result does not perfectly match the actual utterance of the interlocutor 101. In addition to filtering, when speech separation processing is performed and the speech recognition results of the separated speech are compared with the lip-reading results, the degree of sound similarity is evaluated according to whether the sounds are the same, similar, or different. Note that the method of changing the evaluation score according to the degree of sound similarity is merely an example and does not limit the evaluation method. The evaluation of sound similarity may also be carried out using other technologies, such as other correlation evaluation techniques or AI technologies.

[0054] As shown in Figure 4, the speech recognition device 10 performs speech recognition processing on the speech data 121a of the interlocutor 101 extracted from the speech data 121 to recognize the content of the interlocutor 101's speech. The speech recognition device 10 realizes a dialogue with the interlocutor 101 by displaying a response to the content of the interlocutor 101's speech on the display unit 40.

[0055] Figure 5 shows an example of a response displayed on the display unit 40 by the speech recognition device 10. For example, as shown in Figure 5, the screen of the display unit 40 displays the character 200 that interacts with the interlocutor 101 and the response 201 to the interlocutor 101. Additional information 202 related to the response 201 to the interlocutor 101 may also be displayed on the screen of the display unit 40.

[0056] For example, in the example shown in Figure 4, if the interlocutor 101 says "I'm looking for a store" and then says the name of a store, then, as shown in Figure 5, a response 201 that provides the store's location in text, and additional information 202 that includes a map showing the store's location will be displayed.

[0057] Furthermore, the character 200 displayed on the display unit 40 is not limited to a person; it may be an animal or other living thing, or a robot or a building or other inanimate object. Part or all of the character 200 may be composed of images (photographs) of real living or inanimate objects, or it may be composed of images (pictures) such as illustrations or diagrams drawn by a person or AI.

[0058] Figure 6 illustrates another example of the processing performed by the speech recognition device 10. In the example shown in Figure 6, the image data 132 obtained by the camera 30 shows the person speaking 102 and other people 300 (300a, 300b). In such a case, the speech recognition device 10 can perform a person identification process to identify the person speaking 102 from among the multiple people shown in the image data 132.

[0059] For example, the speech recognition device 10 identifies the person who appears largest in the image as the conversational partner 102. That is, it identifies the person closest to the speech recognition device 10 as the conversational partner 102. Alternatively, for example, the speech recognition device 10 may recognize the faces of each person in the image and identify the person whose face is directly facing the speech recognition device 10 as the conversational partner 102.

[0060] The speech recognition device 10 may identify the voices of each person captured by the microphone 20 based on the timing of each person's speech in the image, compare their volumes, and identify the person with the highest volume as the conversationalist 102.

[0061] The speech recognition device 10 may perform lip-reading on each person in the image, identify the voice of each person obtained by the microphone 20 based on the lip-reading results, and compare the volume of each voice to identify the person with the highest volume as the conversationalist. For example, if lip-reading reveals that there is a person who said "store" and a person who said "today", the speech recognition device 10 will compare the volume of the voice "store" and the voice "today" contained in the voice data 122 and identify the person with the higher volume as the conversationalist 102.

[0062] The speech recognition device 10 performs a speech separation process to separate multiple voices contained in the speech data 122, and obtains multiple voice data 122a and 122b as shown in Figure 6. The speech recognition device 10 performs a speech recognition process on each of the obtained voice data 122a and 122b, compares the obtained speech recognition results with the lip-reading results, and determines that the voice data 122a with a high correlation to the lip-reading results of the interlocutor 102 is the voice of the interlocutor 102.

[0063] For example, suppose the speech recognition result of voice data 122a is "shop," the speech recognition result of voice data 122b is "today," and the lip-reading result of the interlocutor 102 is "shop" or "house," etc. In this case, the speech recognition result of voice data 122a will have a higher correlation with the lip-reading result than the speech recognition result of voice data 122b. Therefore, the speech recognition device 10 determines that the voice contained in voice data 122a is the voice of the interlocutor 102 and identifies the content of the interlocutor 102's utterance based on the speech recognition result of voice data 122a. The speech recognition device 10 performs dialogue processing based on the content of the interlocutor 102's utterance and displays the response to the interlocutor 102 on the display unit 40 as shown in the example in Figure 5.

[0064] Figure 4 shows an example where the voice of interlocutor 101 is extracted by filtering, and Figure 6 shows an example where the voice of interlocutor 102 is identified from multiple voices obtained by voice separation processing. However, the voice recognition device 10 may also perform only one of filtering or voice separation processing.

[0065] The speech recognition device 10 may be configured to select and perform filtering or speech separation processing depending on whether the image captured by the camera 30 contains multiple people. For example, as shown in Figure 4, filtering may be performed if there is only one person in the image data 131, and as shown in Figure 6, speech separation processing may be performed if there are multiple people in the image data 132. If there is only one person, filtering can be used to reduce environmental noise, so filtering should be performed. If there are multiple people, filtering may not be able to separate the voices of each person, so speech separation processing using AI technologies such as independent component analysis, machine learning, or deep learning should be performed.

[0066] The examples described above show the case where there is one interlocutor, but the speech recognition device 10 can also interact with multiple interlocutors. Figure 7 is a diagram illustrating an example where there are multiple interlocutors 103 (103a, 103b).

[0067] For example, as shown in Figure 7(a), if the image data 133 obtained by the camera 30 contains multiple people 103a and 103b, the speech recognition device 10 identifies each person 103a and 103b included in the image data 133 and performs lip-reading processing.

[0068] The speech recognition device 10 performs a speech separation process to separate the speech data 123 obtained by the microphone 20 into multiple speech data 123a and 123b, similar to the example described in Figure 6. The speech recognition device 10 then performs a speech recognition process on each of the speech data 123a and 123b.

[0069] The speech recognition device 10 compares the lip-reading results for each person 103a and 103b obtained from the image data 133 with the speech recognition results obtained from each voice data 123a and 123b, and identifies the voice data 123a and 123b for each person 103a and 103b based on the correlation.

[0070] Specifically, when lip-reading results are obtained indicating that person 103a uttered "mise o" or "ie o," the speech recognition device 10 performs speech recognition processing on the audio portion of each audio data 123a and 123b that corresponds to the utterance timing in the lip-reading process. The speech recognition device 10 compares the speech recognition results with the lip-reading results and identifies the audio data 123a containing the word "mise o" as the audio data of person 103a based on the correlation.

[0071] Similarly, if the lip-reading results indicate that another person 103b has spoken "chizu o" or "inu o," the speech recognition device 10 performs speech recognition processing on each of the speech data 123a and 123 at the same timing as the utterance in the lip-reading process. The speech recognition device 10 compares the speech recognition results with the lip-reading results and identifies the speech data 123b containing the word "chizu o" as the speech data of person 103a based on the correlation. However, if there are two people and two sets of speech data, the speech recognition device 10 may compare the lip-reading results and speech recognition results for only one person to identify the speech data, and identify the remaining speech data as the speech data of the other person.

[0072] Thus, even when multiple people 103a and 103b speak almost simultaneously, the speech recognition device 10 can estimate the content of each person's speech by lip-reading processing if their speech content is different, and can identify the correspondence between each person 103a and 103b and each speech data 123a and 123b.

[0073] The speech recognition device 10, having identified the voice data 123a and 123b of each person 103a and 103b, performs dialogue processing based on the content of each person's speech obtained by speech recognition of each voice data 123a and 123b. The speech recognition device 10 can respond to multiple interlocutors 103a and 103b almost simultaneously and realize dialogue with each interlocutor 103a and 103b.

[0074] In the dialogue process, the speech recognition device 10 displays responses to each of the interlocutors 103a and 103b. For example, as shown in Figure 7(b), the display unit 40 of the speech recognition device 10 displays the character 200 corresponding to the speech recognition device 10, the characters 203a and 203b corresponding to each of the interlocutors 103a and 103b, and the responses 202a and 202b to each of the interlocutors 103a and 103b.

[0075] In the example shown in Figure 7(b), a human illustration is displayed as character 200 corresponding to the voice recognition device 10. In addition, partial images of the faces of each interlocutor 103a and 103b, extracted from the images captured by the camera 30, are displayed as characters 203a and 203b corresponding to each interlocutor 103a and 103b.

[0076] One of the interlocutors, 103a, said, "Tell me where the store is," and then said the name of the store. Therefore, corresponding to the character 203a of interlocutor 103a, the response 202a, which provides the location of the store in text, is displayed. The other interlocutor, 103b, said, "Show me a map." Therefore, corresponding to the character 203b of interlocutor 103b, the response 202b, which includes a map in addition to the text indicating the response, is displayed. When multiple character interactions are taking place in combination, and these interactions are displayed together on the display unit 40, for example, as shown in Figure 7(b), arrows are added next to the display of each utterance to indicate which character is speaking to which character. The speech recognition device 10 can display interactions taking place in combination in a way that distinguishes each interaction from the others.

[0077] Thus, the speech recognition device 10 not only separates and extracts the voices of each person contained in the image data from the speech data, but also recognizes the correspondence between each person and each voice through lip-reading processing, enabling it to establish a dialogue with each person.

[0078] Although Figure 7 illustrates an example of voice separation processing, the voice recognition device 10 may also perform filtering to extract voice data from each person. If lip-reading is performed and it is determined that one person 103a said "shop," then voice data can be extracted using a frequency filter capable of extracting the sound "shop," and the voice of this person 103a can be recognized. Similarly, if it is determined that the other person 103b said "map," then voice data can be extracted using a frequency filter capable of extracting the sound "map," and the voice of this person 103a can be recognized.

[0079] The voice recognition device 10 according to this embodiment is not particularly limited in type, location of use, or manner of use, as long as it comprises at least one microphone 20, a camera 30, and a display unit 40. For example, in addition to the configuration consisting of a self-standing stand as in the example described above, the voice recognition device 10 may be used in a configuration that is embedded in a wall or pillar. The voice recognition device 10 may be installed inside or outside a building in various facilities, transportation facilities such as stations, event venues, etc., and may be used as digital signage for displaying advertisements, or it may be used to provide various kinds of information. The voice recognition device 10 may also be installed on a robot used to guide users.

[0080] In this embodiment, an example was described in which the lip-reading result obtained by extracting the first portion of data from image data obtained by a camera is compared with the speech recognition result of the corresponding portion of audio data obtained by a microphone. However, the data to be subjected to lip-reading and speech recognition is not particularly limited. For example, the lip-reading and speech recognition may be performed on data extracted from the middle portion of the image data and audio data, or on the entire image data and audio data.

[0081] In this embodiment, a configuration example in which the speech recognition device includes a microphone and a camera has been mainly described. However, it is also possible that at least one of the microphone and the camera is provided independently of the speech recognition device, thereby constituting a speech recognition system.

[0082] The configurations of the devices and systems illustrated in this embodiment are functionally schematic, and the physical configurations of the devices and systems are not limited to these configurations. The forms of distribution and integration of each device are not limited to the examples described above, and all or part of them can be configured by functionally or physically distributing and integrating them in any unit according to various loads and usage conditions.

[0083] As described above, the speech recognition device according to this embodiment can perform lip-reading processing to estimate speech content from the movement of a person's lips contained in image data captured by a camera, and speech recognition processing of voice data acquired by a microphone. The speech recognition device can set a filter to obtain the same or similar speech recognition result as the lip-reading processing result from the voice data obtained by the microphone, and perform filtering of the voice data using the filter to acquire the user's voice. The speech recognition device can also perform voice separation processing to separate and extract the voice contained in the voice data obtained by the microphone, and acquire the voice from which the same or similar speech recognition result as the lip-reading processing result is obtained as the user's voice. The speech recognition device performs processing to identify the user's voice contained in the voice data, targeting part or all of the voice data. The speech recognition device extracts and acquires the user's voice from the voice data by performing filtering and voice separation processing, at least one of the two. The speech recognition device performs speech recognition processing on the acquired user's voice to identify the user's utterances, and can realize dialogue with the user by returning a response to the identified utterances to the user by displaying information indicating the response and playing back the voice. This allows the speech recognition device to acquire clear user speech, perform speech recognition, and accurately identify the content of the utterance. As a result, dialogue with the user can be made smoother.

[0084] While embodiments of the speech recognition device according to this disclosure have been described above with reference to the drawings, the configuration and operation of the speech recognition device are not limited to the above embodiments, and may be implemented in various ways with improvements, changes, and modifications based on the knowledge of those skilled in the art, without departing from the spirit of the invention. [Industrial applicability]

[0085] As described above, the speech recognition device, speech recognition system, and speech recognition method related to this disclosure are useful for accurately recognizing the user's voice. [Explanation of Symbols]

[0086] 1.11 Server equipment 2 Network 10. Voice recognition device 20, 510 (510a, 510b) Microphone 30, 520 (520a, 520b) Camera 40 Display section 50 Control Unit 60 Storage section 70 Communications Department 500 (500a, 500b) Computer Equipment

Claims

1. A speech recognition device that recognizes speech in order to interact with users, A microphone for acquiring the voice of the user, A camera for capturing images of the aforementioned user, A control unit identifies the user's speech content based on the speech content obtained by performing lip-reading processing to estimate the content of a person's speech from the movement of the person's lips included in the image data captured by the camera, and the speech content obtained by performing speech recognition processing on the voice data acquired by the microphone. A voice recognition device characterized by being equipped with the following features.

2. The control unit, Based on the content of the person's speech obtained by performing the lip-reading process, the user's voice included in the audio data acquired by the microphone is identified. The speech content obtained by performing speech recognition processing on the identified voice is identified as the user's speech content. The speech recognition device according to feature 1.

3. The control unit, Based on the correlation between the spoken content of the person obtained by performing the lip-reading process and the spoken content corresponding to each of the target voices obtained by performing speech recognition processing on one or more voices separated from the audio data acquired by the microphone, the user's spoken content is identified. The speech recognition device according to feature 1.

4. The control unit, A filter is set to extract audio corresponding to the content of the person's speech obtained by performing the lip-reading process from the audio data acquired by the microphone, The audio data is filtered using the set filter to extract the audio. The extracted audio is subjected to speech recognition processing, and the resulting utterance is identified as the utterance of the user. The speech recognition device according to feature 1.

5. The control unit, Lip-reading processing is performed to estimate the content of speech from each person based on the movement of the lips of each person included in the image data captured by the aforementioned camera. For each of the aforementioned individuals, a filter is set that can extract audio corresponding to the content of that person's speech obtained by performing the lip-reading process from the audio data acquired by the microphone. The audio data is filtered using the configured filters to extract audio corresponding to the utterances of each person, and speech recognition processing is performed on each obtained audio to identify the utterances of each person. The speech recognition device according to feature 4.

6. The control unit, One or more voices are separated from the audio data acquired by the microphone, Based on the correlation between the speech content corresponding to each of the target speeches obtained by performing speech recognition processing on each separated speech and the speech content of the person obtained by performing the lip-reading processing, the user's speech is identified from the multiple speeches, and the speech content obtained from that speech is identified as the user's speech. The speech recognition device according to feature 1.

7. The control unit, By performing lip-reading processing to estimate the speech content of one or more people contained in the image data captured by the camera, and comparing the correlation between the speech content of each person obtained by performing speech recognition processing on one or more voices separated from the audio data acquired by the microphone and the speech content corresponding to each voice of the target, the correlation is compared. For each of the aforementioned individuals, the voice in the combination with the highest correlation is determined to be the voice of that individual, and the speech content obtained by performing speech recognition processing on that voice is identified as the speech content of the user, that individual. The speech recognition device according to feature 6.

8. The control unit identifies the user's voice based on the speech content obtained by performing lip-reading on a portion of image data extracted from image data captured by the camera at multiple points in time, and the speech content obtained by performing speech recognition on a portion of audio data extracted from audio data acquired by the microphone, corresponding to the portion of the image data. The speech recognition device according to feature 1.

9. A speech recognition system that recognizes speech in order to interact with users, A microphone for acquiring the voice of the user, A camera for capturing images of the aforementioned user, An information processing device that identifies the user's speech content based on speech content obtained by performing lip-reading processing to estimate the content of a person's speech from the movement of the person's lips contained in the image data captured by the camera, and speech content obtained by performing speech recognition processing on the voice data acquired by the microphone. A speech recognition system characterized by having the following features.

10. A speech recognition method for recognizing speech in order to interact with users, A process of performing lip-reading to estimate the content of a person's speech from the movement of the person's lips contained in image data captured by a camera, The process involves performing speech recognition processing on the audio data acquired by the microphone, A step of identifying the user's speech content based on the speech content obtained by performing the lip-reading process and the speech content obtained by performing the speech recognition process. A speech recognition method characterized by including [a certain feature].

Citation Information

Patent Citations

  • Security system and monitoring display device

    JP7074716B2