Speech recognition device, speech recognition method, speech recognition program, speech recognition system

The voice recognition system addresses the challenge of identifying and displaying the speech of a specific speaker by using face and speech analysis to ensure accurate and clutter-free display of the intended speaker's content.

JP7838292B2Active Publication Date: 2026-04-01RICOH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-02-10
Publication Date
2026-04-01

AI Technical Summary

Technical Problem

Conventional voice recognition systems struggle to accurately identify and display the speech content of a specific speaker when multiple individuals are present in an image, leading to potential misidentification of the speaker whose speech is displayed.

Method used

A voice recognition system that includes a target speaker determination unit to identify the speaker based on face images and speech content recognition, utilizing speaker embedding information to ensure accurate display of the intended speaker's speech, even when the face is not visible.

Benefits of technology

The system effectively displays the speech content of the intended speaker, improving accuracy and reducing visual clutter by focusing on the specific speaker of interest, even when face detection is unavailable.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007838292000001
    Figure 0007838292000001
  • Figure 0007838292000002
    Figure 0007838292000002
  • Figure 0007838292000003
    Figure 0007838292000003
Patent Text Reader

Abstract

To provide a speech recognition device, a speech recognition method, a speech recognition program, and a speech recognition system that display speech contents of a speaker who is a sound source identified from images contained in video data.SOLUTION: In a speech recognition system in which an information processing device, an image capture device, and a display device are connected via a network, etc., a speech recognition processing unit 230 of an information processing terminal 200A, which is the information processing device, includes: a focus speaker determination unit 247 that determines a focus speaker based on a facial image of a person detected from an image shown by image data included in video data; and a speech content recognition result output unit 233 that displays text data converted from speech data of the focus speaker among the speech data included in the video data on the display device.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a voice recognition device, a voice recognition method, a voice recognition program, and a voice recognition system.

Background Art

[0002] In recent years, a technique has been known for identifying a speaker who is the sound source from an image and converting the voice uttered by the identified speaker into a character image and displaying it on a display unit. Specifically, for example, there is known a system that converts the voice of a person whose mouth is moving in an image into characters and displays it when the person's mouth is moving.

Summary of the Invention

Problems to be Solved by the Invention

[0003] In the above-described conventional technique, when there are a plurality of persons whose mouths are moving in an image, etc., it is not possible to select a speaker to be noted. For this reason, in the conventional technique, when paying attention to a specific person, there is a possibility that the speech content of the person being noted is not appropriately displayed.

[0004] The disclosed technique is in view of the above circumstances, and aims to display the speech content of a specific speaker.

Means for Solving the Problems

[0005] The disclosed technique includes a target speaker determination unit that determines a target speaker based on a face image of a person detected from an image shown by image data included in video data, and a speech content recognition result output unit that causes a display device to display text data converted from the speech data of the target speaker among the speech data included in the video data. A speaker embedding information calculation unit calculates speaker embedding information to identify the person who spoke the audio data, and an off-screen speaker estimation unit estimates the speaker of the audio data based on the speaker embedding information calculated from the audio data when the face image of the speaker of interest is not detected from the image shown in the image data included in the video data. It is a voice recognition device having the above.

Effects of the Invention

[0006] The speech content of a specific speaker can be displayed.

Brief Description of the Drawings

[0007] [Figure 1] This figure shows an example of a speech recognition system according to the first embodiment. [Figure 2] This is the first figure illustrating the overview of speech recognition in the first embodiment. [Figure 3] This diagram illustrates the application of a voice recognition system to smart glasses. [Figure 4] The second figure illustrates the overview of speech recognition in the first embodiment. [Figure 5] This is a diagram explaining the functions of smart glasses. [Figure 6] This is the first flowchart illustrating the operation of the smart glasses of the first embodiment. [Figure 7] This is a second flowchart illustrating the operation of the smart glasses in the first embodiment. [Figure 8] This is a diagram illustrating the determination of the speaker of interest in the first embodiment. [Figure 9] This is a third flowchart illustrating the operation of the smart glasses in the first embodiment. [Figure 10] This is the first diagram illustrating an example of how smart glasses work. [Figure 11] This is the second diagram illustrating an example of how smart glasses work. [Figure 12] This is the third diagram illustrating an example of how smart glasses work. [Figure 13] This is the fourth diagram illustrating an example of how smart glasses work. [Figure 14] This is the fifth diagram illustrating an example of how smart glasses work. [Figure 15] This is the sixth diagram illustrating an example of how smart glasses work. [Figure 16] This is the seventh diagram illustrating an example of how smart glasses work. [Figure 17] This is a diagram illustrating the recognition of speech content in the first embodiment. [Figure 18] It is a diagram showing an example of a translation system according to a second embodiment. [Figure 19] It is a diagram showing an example of a speech content recording system according to a third embodiment.

Embodiments for Carrying Out the Invention

[0008] (First Embodiment) The first embodiment will be described below with reference to the drawings. FIG. 1 is a diagram showing an example of an audio recognition system according to the first embodiment.

[0009] The audio recognition system 100 of the present embodiment includes an information processing apparatus 200, an imaging apparatus 300, and a display apparatus 400, and the information processing apparatus 200, the imaging apparatus 300, and the display apparatus 400 are connected via a network or the like.

[0010] In the audio recognition system 100 of the present embodiment, the information processing apparatus 200 has an audio recognition processing unit 230.

[0011] The imaging apparatus 300 of the present embodiment acquires video data and transmits it to the information processing apparatus 200. The video data includes image data (moving image data) and audio data.

[0012] In the present embodiment, the video data may be shot by the user of the audio recognition system 100 using the imaging apparatus 300 and transmitted to the information processing apparatus 200. Therefore, the image data included in the video data includes the speaker that the user himself / herself is paying attention to.

[0013] The information processing apparatus 200 of the present embodiment identifies the speaker that the user of the audio recognition system 100 is paying attention to based on the image data included in the video data acquired from the imaging apparatus 300 by the audio recognition processing unit 230. Then, the information processing apparatus 200 converts only the audio data of the speaker that the user has paid attention to into text data by the audio recognition processing unit 230 and displays it on the display apparatus 400. Note that the display apparatus 400 may be, for example, a display or the like included in the information processing apparatus 200.

[0014] Thus, in this embodiment, from the image data, the speaker who the user is paying attention to is identified, and only the voice data of the identified speaker is converted into text data and output. Therefore, according to this embodiment, the speech content of a specific speaker who the user is paying attention to can be displayed.

[0015] Hereinafter, referring to FIG. 2, an overview of the speech recognition by the information processing apparatus 200 of this embodiment will be described. FIG. 2 is a first diagram for explaining an overview of the speech recognition of the first embodiment.

[0016] The image 21 shown in FIG. 2 is an example of an image shown by the image data acquired by the imaging device 300.

[0017] The image 21 includes an image of person A and an image of person B. The information processing apparatus 200 of this embodiment identifies, as the person to be noted, the image of the person among the images of the persons included in the image 21 whose face image is located at a position closer to the center of the image 21.

[0018] In the example of FIG. 2, the face image of person A is closer to the center of the image 21 than the face image of person B. Therefore, in FIG. 2, person A is identified as the speaker to be noted.

[0019] Note that the center of the image may be the position where the diagonals of the rectangle shown by the image intersect. In the following description, the speaker to be noted may be expressed as the noted speaker. The speaker to be noted is, in other words, a specific speaker that the user of the speech recognition system 100 is paying attention to.

[0020] In this embodiment, when person A is identified as the noted speaker, using the moving image showing the movement of the lip portion of person A and the voice data acquired by the imaging device 300, etc., the voice data of person A who is the noted speaker is converted into text data 23 and displayed. Therefore, according to this embodiment, the speech content of the noted speaker can be converted into text data with high accuracy.

[0021] In this embodiment, the image 21 and the text data 23 may be displayed superimposed on each other.

[0022] Next, referring to Figure 3, we will describe the case where the voice recognition system 100 of this embodiment is applied to smart glasses.

[0023] Figure 3 illustrates the application of a voice recognition system to smart glasses. In Figure 3, the voice recognition system 100 is described as smart glasses 100A.

[0024] Figure 3 shows an example of the hardware configuration of the smart glasses. The smart glasses 100A in this embodiment are a glasses-type wearable terminal that includes a glasses-type display device 300A, an information processing terminal 200A, and a cable 150. In the example in Figure 3, the glasses-type display device 300A and the information processing terminal 200A are connected by the cable 150, but this is not limited to this. The glasses-type display device 300A and the information processing terminal 200A may communicate wirelessly.

[0025] The glasses-type display device 300A includes a camera (imaging device) 110, a microphone (sound collecting device) 120, a display (display device) 130, and an operating member 140. In other words, the glasses-type display device 300A includes an imaging device and a display device.

[0026] The camera 110 acquires image data in the direction of the wearer's gaze while wearing the smart glasses 100A. The microphone 120 acquires audio data from the area around the smart glasses 100A. The display 130 displays text data output from the information processing terminal 200A. In this embodiment, the display 130 may be an optical see-through type display. The operating member 140 may be a physical button or the like, and various operations are performed on the glasses-type display device 300A.

[0027] Furthermore, although the camera 110 and microphone 120 are provided separately in this embodiment, the system is not limited to this configuration. The microphone 120 in this embodiment may be built into the camera 110. In this case, the camera 110 will acquire video data including image data and audio data.

[0028] Cable 150 transmits image data acquired by camera 110 and audio data acquired by microphone 120 to information processing terminal 200A. Cable 150 also transmits various information from information processing terminal 200A to glasses-type display device 300A.

[0029] The information processing terminal 200A includes an information input / output interface (I / F) 201, memory 202, operating device 203, storage 204, power supply 205, CPU (Central Processing Unit) 206, and network interface (I / F) 207.

[0030] The information input / output interface (I / F) 201 is an interface for sending and receiving various types of data between the information processing terminal 200A and the glasses-type display device 300A. The memory 202 stores temporary information such as audio data and image data (video data). The operating device 203 allows the wearer of the smart glasses 100A to perform various operations such as running applications and turning the power on / off. The operating device 203 may be implemented, for example, by a touch panel.

[0031] Storage 204 stores various models and other items, which will be described later. Power supply 205 supplies power to each device of the smart glasses 100A. CPU 206 performs various processes and controls the overall operation of the smart glasses 100A.

[0032] The information processing terminal 200A realizes the functions of the speech recognition processing unit 230 by having the CPU 206 read and execute a program stored in the storage 204 or the like.

[0033] Network interface 207 is an interface for accessing the communication network.

[0034] Although the smart glasses 100A shown in Figure 3 include a glasses-type display device 300A and an information processing terminal 200A, it is not limited to this configuration. In the smart glasses 100A, the glasses-type display device 300A may have all the components of the information processing terminal 200A.

[0035] Next, with reference to Figure 4, an overview of speech recognition using smart glasses 100A will be described. Figure 4 is the second diagram illustrating the overview of speech recognition in the first embodiment.

[0036] In the smart glasses 100A of this embodiment, the information processing terminal 200A identifies the speaker of interest based on image data acquired from the camera 110. The information processing terminal 200A then displays on the display 130 text data converted from the voice data of the speaker of interest acquired from the microphone 120.

[0037] In the example shown in Figure 4, the wearer P of the smart glasses 100A is focusing on person A. Also, in the example shown in Figure 4, both person A and person B are in the wearer P's line of sight. In this case, the image captured by the camera 110 of the smart glasses 100A will be an image with person A, the person P is focusing on, positioned in the center.

[0038] Therefore, the smart glasses 100A identify person A as the speaker of interest, convert only the speaker's voice data into text data, and display only the text data on the display 130.

[0039] In this embodiment, by applying the voice recognition system 100 to the smart glasses 100A, the person the wearer P of the smart glasses 100A is paying attention to is identified as the speaker of interest simply by the wearer P turning in the direction of the person P is paying attention to.

[0040] Furthermore, in this embodiment, the display 130 of the smart glasses 100A is of the optical see-through type. Therefore, in this embodiment, the text data 23 can be viewed by the wearer P without obstructing the wearer P's field of vision.

[0041] The display 130 does not necessarily have to be an optical see-through type; the image shown by the image data acquired by the camera 110 and the text data 23 may be superimposed and displayed.

[0042] Furthermore, the smart glasses 100A may be, for example, a retinal scanning type spectacle projection device. In this case, a display 130 is unnecessary, and the text data 23 can be projected directly onto the wearer P's retina by the optical system.

[0043] Next, the functions of the smart glasses 100A of this embodiment will be described with reference to Figure 5. Figure 5 is a diagram illustrating the functions of the smart glasses. Specifically, Figure 5 shows the functions of the information processing terminal 200A of the smart glasses 100A.

[0044] The information processing terminal 200A of this embodiment has a speech recognition processing unit 230. The speech recognition processing unit 230 includes a video input unit 231, a voice input unit 232, a speaker of interest identification unit 240, a lip feature acquisition unit 250, an acoustic feature acquisition unit 260, a person identification unit 270, a multimodal recognition unit 280 (first speech recognition unit), a speech recognition unit 290 (second speech recognition unit), and a speech content recognition result output unit 233.

[0045] The video input unit 231 acquires image data (video data) captured by the camera 110. The audio input unit 232 acquires audio data collected by the microphone 120. At this time, the audio input unit 232 may acquire the audio data as uncompressed monaural data sampled under predetermined conditions.

[0046] The speech content recognition result output unit 233 displays on the display 130 the text data which is the result of speech content recognition by the multimodal recognition unit 280 and the text data which is the result of speech content recognition by the speech recognition unit 290.

[0047] The speaker identification unit 240 identifies the speaker of interest from the video data acquired by the video input unit 231. The speaker identification unit 240 includes an image conversion unit 241, a face region recognition unit 242, a face region detection model 243, a face position determination unit 244, a lip region extraction unit 245, a face feature point estimation model 246, and a speaker of interest determination unit 247.

[0048] The image conversion unit 241 converts the video data into a series of frame images. To speed up processing, the image conversion unit 241 may also convert RGB image data to grayscale image data or convert the number of pixels.

[0049] The face region recognition unit 242 uses the face region detection model 243 to recognize regions containing face images (face regions) in the acquired time-series frame images. The face region detection model 243 is a model that detects face images from images, and is a model that has been trained on a neural network using a large amount of data in advance. The face images detected here are face images of people who are candidates for the speaker of interest.

[0050] The face position determination unit 244 determines the position of the face region in the image shown by the image data acquired by the camera 110 and obtains information indicating the position of the face region.

[0051] The lip region extraction unit 245 uses a facial feature point estimation model 246 to detect the lip region, which includes an image of the lip area, from the facial image within the lip region, and extracts the image within the lip region from the facial image within the lip region.

[0052] The facial feature point estimation model 246 obtains the coordinates of the contours of the eyes, nose, and lips from a facial image, and detects the coordinates around the lips.

[0053] In this embodiment, the lip region extraction unit 245 extracts images of the lip region, but it is not limited to this. For example, if the lip region in a face image is obscured by a person's hand or the like, the lip region extraction unit 245 may extract images of regions corresponding to facial features such as the eyes or nose. Regions corresponding to facial features such as the eyes or nose may be detected by the facial feature point estimation model 246.

[0054] The speaker of interest determination unit 247 determines the speaker of interest based on the position of the person's face region in the image and the mouth movements shown in the image (video) within the lip region. Details of the processing of the speaker of interest identification unit 240 will be described later.

[0055] The lip feature acquisition unit 250 of this embodiment acquires lip features indicating mouth movement from the image within the lip region of the image acquired by the camera 110.

[0056] The lip feature acquisition unit 250 includes a lip pixel count conversion unit 251, a lip feature calculation unit 252, and a lip feature calculation model 253.

[0057] The lip pixel count conversion unit 251 converts the extracted image within the lip region into an image of a predetermined size. In other words, the lip pixel count conversion unit 251 enlarges or reduces the image within the lip region, which varies in size depending on the distance between the camera 110 and the person being photographed, so that it becomes an image of a uniform size.

[0058] The lip feature calculation unit 252 calculates lip features using the lip feature calculation model 253. Specifically, the lip feature calculation unit 252 inputs video data showing images within the lip region in a time series with changed size into the lip feature calculation model 253 and calculates lip features that are effective for recognizing speech content. Lip features are multidimensional vectors output from the lip feature calculation model 253 after inputting video data within the lip region into the lip feature calculation model 253.

[0059] The acoustic feature acquisition unit 260 detects speech segments, which are periods in which a person is speaking, from the audio data acquired by the audio input unit 232, and acquires the acoustic features of the audio data of the speech segments.

[0060] The acoustic feature acquisition unit 260 includes a speech utterance interval detection unit 261, a speech utterance interval detection model 262, and an acoustic feature calculation unit 263.

[0061] The speech utterance segment detection unit 261 uses the speech utterance segment detection model 262 to detect speech segments from the input speech data.

[0062] The acoustic feature calculation unit 263 calculates acoustic features from the speech waveform of the section detected as a speech interval. The acoustic features may be, for example, Mel-frequency cepstrum coefficients (MFCCs), log-Mel filter bank features (FBANK), or log-Mel filters.

[0063] The person identification unit 270 acquires information for identifying the person who spoke from acoustic features. The person identification unit 270 includes a speaker embedding information calculation unit 271, a speaker embedding information calculation model 272, and an off-screen speaker estimation unit 273.

[0064] The speaker embedding information calculation unit 271 calculates speaker embedding information that represents the voice quality of the speaker using the speaker embedding information calculation model 272. Speaker embedding information is information used to identify the speaker and may be a feature quantity of a certain number of dimensions extracted using methods such as i-vector, d-vector, or x-vector.

[0065] The off-screen speaker estimation unit 273 estimates the speaker using speaker embedding information when the orientation of the wearer P of the smart glasses 100A changes and the image of the speaker of interest is no longer included in the image acquired by the camera 110. Details of the processing of the off-screen speaker estimation unit 273 will be described later.

[0066] The multimodal recognition unit 280 recognizes the speech content of a speaker of interest using lip features and acoustic features. The multimodal recognition unit 280 includes a feature integration unit 281, a multimodal speech content recognition unit 282, and a multimodal speech content recognition model 283.

[0067] The feature integration unit 281 integrates the acoustic features acquired by the acoustic feature acquisition unit 260 and the lip features acquired by the lip feature acquisition unit 250 to create a multimodal feature. A multimodal feature is a feature that includes multiple types of features. More specifically, a multimodal feature includes acoustic features and lip features.

[0068] The multimodal speech content recognition unit 282 recognizes speech content using the multimodal speech content recognition model 283. More specifically, the multimodal speech content recognition unit 282 in this embodiment recognizes speech content using acoustic features extracted from audio data and lip features extracted from video data.

[0069] The speech recognition unit 290 recognizes the content of speech from the speech data acquired by the speech input unit 232, based on the acoustic features acquired by the acoustic feature acquisition unit 260. The speech recognition unit 290 includes a speech content recognition unit 291 and a speech content recognition model 292.

[0070] The speech content recognition unit 291 performs speech content recognition using acoustic features if lip features of the person designated as the speaker of interest cannot be obtained. Specifically, the speech content recognition unit 291 uses the speech content recognition model 292 to recognize speech content based on speech data and passes the recognition result to the speech content recognition result output unit 233.

[0071] In this embodiment, if the image extracted by the lip region extraction unit 245 corresponds to an image of a facial part other than the mouth, the lip feature quantity may be considered not to have been calculated.

[0072] In this embodiment, the face region detection model 243, the face feature point estimation model 246, the speech utterance segment detection model 262, and the speaker embedding information calculation model 272 may be models using publicly known technologies.

[0073] Next, the operation of the smart glasses 100A of this embodiment will be described with reference to Figure 6. Figure 6 is a first flowchart illustrating the operation of the smart glasses of the first embodiment.

[0074] The process shown in Figure 6 is executed, for example, when the wearer P of the smart glasses 100A issues an operation to initiate the recognition process of the speech content of the speaker of interest.

[0075] In the smart glasses 100A of this embodiment, the information processing terminal 200A acquires image data (video data) and audio data using the video input unit 231 and the audio input unit 232 (step S601).

[0076] Next, the information processing terminal 200A performs a process to detect a speech segment using the speech segment detection unit 261 (step S602).

[0077] If no speech segment is detected in step S602, the information processing terminal 200A returns to step S601.

[0078] In step S602, if a speech segment is detected, the information processing terminal 200A repeats the processing from step S605 to step S607 for the number of people whose face images were detected (step S604).

[0079] The information processing terminal 200A uses the face region recognition unit 242 to detect the face region containing the face image in the image data acquired by the video input unit 231 (step S605).

[0080] Next, the information processing terminal 200A uses the lip region extraction unit 245 to detect the lip region from within the face region (step S606). If the lip region is not detected within the face region, the lip region extraction unit 245 only needs to detect a region corresponding to an image of another facial feature (such as the eyes or nose). In other words, the lip region extraction unit 245 only needs to detect a region from the face region that corresponds to an image of a part of the face.

[0081] Furthermore, in this embodiment, detecting the lip region may be synonymous with extracting a lip image from the face image within the face region.

[0082] Next, the information processing terminal 200A selects a speaker of interest using the speaker selection unit 247 (step S607). Details of the process in step S607 will be described later.

[0083] The information processing terminal 200A repeats the process from step S605 to step S607 for each person (step S608). In this embodiment, the speaker of interest is determined by repeating this process.

[0084] Next, the information processing terminal 200A calculates acoustic features from the voice data of the person identified as the speaker of interest using the acoustic feature acquisition unit 260 (step S609).

[0085] Next, the information processing terminal 200A calculates the speaker embedding information of the speaker of interest using the speaker embedding information calculation unit 271 (step S610). The speaker embedding information calculation unit 271 may retain the speaker embedding information of the speaker of interest after the speaker of interest has been determined. The speaker embedding information calculation unit 271 may also delete the retained speaker embedding information when the speaker of interest is no longer the speaker of interest.

[0086] Next, the information processing terminal 200A determines whether or not a lip region has been detected by the lip feature acquisition unit 250 (step S611). In other words, the information processing terminal 200A determines whether or not the image extracted by the lip region extraction unit 245 is an image within the lip region.

[0087] If the lip area is not detected in step S611, the information processing terminal 200A uses the speech recognition unit 290 to recognize the content of the speech based on the speech data (step S612), and then proceeds to step S615, which will be described later.

[0088] In step S611, if a lip region is detected, the information processing terminal 200A calculates lip features from the image within the lip region using the lip feature acquisition unit 250 (step S613).

[0089] Next, the information processing terminal 200A uses the multimodal recognition unit 280 to recognize the content of the speech using the acoustic features calculated in step S609 and the lip features calculated in step S613 (step S614).

[0090] Next, the information processing terminal 200A outputs the recognition result text data using the speech content recognition result output unit 233 (step S615), and terminates processing. In other words, the speech content recognition result output unit 233 displays the recognition result text data on the display 130 and terminates processing.

[0091] Thus, in this embodiment, only the voice data of the person designated as the speaker of interest is used as the voice data for recognizing the content of the speech.

[0092] Next, the processing of the speaker of interest determination unit 247 in this embodiment will be described with reference to Figure 7. Figure 7 is a second flowchart illustrating the operation of the smart glasses in the first embodiment. Figure 7 shows the details of the processing in step S607 of Figure 6.

[0093] In the information processing terminal 200A of this embodiment, the speaker of interest determination unit 247 determines in step S605 whether or not multiple facial regions have been detected (step S701).

[0094] If no multiple face regions are detected in step S701, that is, if only one face region is detected, the speaker of interest determination unit 247 proceeds to step S704, which will be described later.

[0095] In step S701, if multiple regions are detected, the speaker of interest determination unit 247 calculates the distance between the x-coordinate of the center of the lip region in the 1 face image and the x-coordinate of the center point of the image indicated by the image data acquired by the video input unit 231 (step S702).

[0096] If the lip region extraction unit 245 extracts an image of a part of the face and a corresponding region instead of the lip region, then the x-coordinate of the center point of this region can be used instead of the x-coordinate of the center of the lip region.

[0097] Next, the speaker of interest determination unit 247 determines whether the calculated distance is the minimum among the distances calculated for multiple face regions (step S703). In other words, the speaker of interest determination unit 247 determines whether the calculated distance is smaller than the previously calculated distance. That is, here the person closest to the center of the image captured by the camera 110 is detected.

[0098] In step S703, if the distance is not the minimum, the speaker of interest determination unit 247 determines that the person corresponding to this face region is not the speaker of interest (step S705), and terminates the process.

[0099] In step S703, if the distance is minimal, the speaker of interest determination unit 247 determines whether or not the lips are moving from the image within the region extracted by the lip region extraction unit 245 (step S704). In other words, here the speaker of interest determination unit 247 determines whether or not the person corresponding to the face region is speaking.

[0100] If the lips are not moving in step S704, the speaker determination unit 247 proceeds to step S705. If the lips are not moving, it indicates that the speaker is not speaking.

[0101] In step S704, if the lips are moving, the speaker of interest determination unit 247 selects this facial region as the facial region of the speaker of interest (step S706) and terminates the process.

[0102] The determination of the speaker of interest by the speaker of interest determination unit 247 will be further explained below with reference to Figure 8. Figure 8 is a diagram illustrating the determination of the speaker of interest in the first embodiment.

[0103] The image 81 shown in Figure 8 is the image data acquired by the video input unit 231. Point o in image 81 is the center point of image 81, and its coordinates are (x1, y1). In this embodiment, the coordinates of the center point o may be, for example, the coordinates when the upper left vertex of image 81 is taken as the origin.

[0104] Figure 8 shows the case where the face region of person A and the face region of person B are detected in step S605 of Figure 6. In this case, the information processing terminal 200A detects the lip region from each face region in step S606 of Figure 6. In the example in Figure 8, the lip region Ra is extracted from the face region of person A, and the lip region Rb is extracted from the face region of person B.

[0105] Here, the speaker of interest determination unit 247, for example, first selects the face region of person B and calculates the distance Lb between the x-coordinate of the center point of the lip region Rb and the x-coordinate of the center point o. At this time, since the distance Lb is the minimum, if person B's lips are moving, person B is selected as the speaker of interest.

[0106] Next, the speaker of interest determination unit 247 selects the face region of person A and calculates the distance La between the x-coordinate of the center point Ra of the lip region and the x-coordinate of the center point o. At this time, the distance La is smaller than the distance Lb. Therefore, the speaker of interest determination unit 247 excludes person B from being the speaker of interest, and if person A's lips are moving, it determines person A to be the speaker of interest.

[0107] Thus, in this embodiment, the person whose face image is detected closest to the center of the image captured by the camera 110 is identified as the speaker of interest. The center of the image captured by the camera 110 is, in other words, the direction of the wearer's gaze on the smart glasses 100A. In other words, in this embodiment, the person closest to the wearer's gaze on the smart glasses 100A is determined to be the speaker of interest. Then, in this embodiment, only the speech of the speaker of interest is converted into text data.

[0108] Therefore, according to this embodiment, even if the image captured by the camera 110 includes multiple people, the smart glasses 100A can identify the person the user is focusing on and display only the speech content of the identified person on the display 130. In other words, the speech content of a specific speaker that the user of the speech recognition system 100 is focusing on can be displayed appropriately without obstructing the user's field of view.

[0109] Furthermore, in this embodiment, since only the speech content of the speaker of interest is displayed on the display 130, it is possible to suppress the amount of information displayed on the display 130 from becoming excessive.

[0110] Furthermore, in this embodiment, since speech content is recognized using both the lip features and acoustic features of the speaker of interest, the accuracy of speech content recognition can be improved.

[0111] Furthermore, in this embodiment, if the lip region is not extracted, a part of the face image and a corresponding region are used as a substitute. Therefore, even if the lip region is not extracted from the face region, the speaker of interest can still be identified.

[0112] Next, with reference to Figure 9, the operation of the smart glasses 100A after the speaker of interest has been determined will be described. Figure 9 is a third flowchart illustrating the operation of the smart glasses of the first embodiment. The process shown in Figure 9 is a process that is executed periodically after the speaker of interest has been determined by the process in Figure 6.

[0113] In the smart glasses 100A of this embodiment, the information processing terminal 200A performs a process to detect a speech interval using the speech utterance interval detection unit 261 (step S901). Subsequently, the information processing terminal 200A calculates acoustic features from the voice data acquired in the speech interval using the acoustic feature calculation unit 263 (step S902).

[0114] Next, the information processing terminal 200A calculates the speaker embedding information of the person who spoke in the speech interval using the speaker embedding information calculation unit 271 (step S903).

[0115] Next, the information processing terminal 200A uses the off-screen speaker estimation unit 273 to determine whether the image of the current speaker of interest is included in the image data obtained by the video input unit 231 (step S904). In other words, it determines whether the speaker of interest remains in the wearer's line of sight.

[0116] In this case, the off-screen speaker estimation unit 273 may, for example, perform face recognition processing on the image data obtained by the video input unit 231 to determine whether or not the face image of the speaker of interest is included.

[0117] If the image of the speaker of interest is not included in step S904, the information processing terminal 200A proceeds to step S909, which will be described later.

[0118] If the image of the speaker in focus is not included, it indicates that the speaker in focus has moved, or the wearer has turned their head, causing the speaker to disappear from the wearer's field of view or move to the edge of their field of view.

[0119] In step S904, if an image of the speaker of interest is included, the information processing terminal 200A determines whether or not the lip region of the speaker of interest has been detected by the lip region extraction unit 245 (step S905). In step S905, if the lip region is not detected, the information processing terminal 200A proceeds to step S911, which will be described later.

[0120] In step S905, if a lip region is detected, the information processing terminal 200A calculates lip features from the image extracted from the lip region using the lip feature acquisition unit 250 (step S906), and proceeds to step S907.

[0121] The processes in steps S907 and S908 in Figure 9 are the same as those in steps S614 and S615 in Figure 6, so their explanation will be omitted.

[0122] In step S904, if the speaker of interest is not included in the image, the information processing terminal 200A uses the off-screen speaker estimation unit 273 of the person identification unit 270 to determine whether less than 10 seconds have passed since the speaker of interest was removed from the image (step S909). Note that 10 seconds is just one example of a preset time and is not limited to this.

[0123] In step S909, if the time is less than 10 seconds, the off-screen speaker estimation unit 273 determines whether the speaker embedding information calculated in step S903 matches the speaker embedding information of the speaker of interest (step S910). The speaker embedding information of the speaker of interest is the speaker embedding information calculated in step S610 in Figure 6.

[0124] Here, the off-screen speaker estimation unit 273 may, for example, calculate the cosine similarity of two speaker embeddings, and if the calculated value is greater than or equal to a predetermined threshold, it may be considered that the two speakers are a match.

[0125] If both match in step S910, the information processing terminal 200A performs speech recognition using the acoustic features calculated in step S902 with the speech recognition unit 290 (step S911), and proceeds to step S908.

[0126] If the two do not match in step S910, the information processing terminal 200A proceeds to step S912, which will be described later.

[0127] In step S909, if it is not less than 10 seconds, that is, if 10 seconds or more have passed since the speaker of interest moved out of the line of sight of the wearer of the smart glasses 100A, the information processing terminal 200A cancels the determination of the speaker of interest (step S912) and terminates the process.

[0128] In other words, in this embodiment, if the face image of the speaker of interest is not detected from the image data acquired by the video input unit 231 for a set period of time or longer, the determination of the speaker of interest is canceled.

[0129] Releasing the focus speaker selection means, in other words, returning from a state where a focus speaker has been selected to the initial state where no focus speaker has been selected.

[0130] In this embodiment, even if the speaker of interest temporarily moves out of the line of sight of the wearer of the smart glasses 100A, it is possible to determine from the audio whether or not it is the speaker of interest speaking and to display the recognition result of the speech content on the display 130.

[0131] The following describes examples of the operation of the smart glasses 100A with reference to Figures 10 to 17.

[0132] Figure 10 is the first diagram illustrating an example of smart glasses operation. In Figure 10, person A is positioned in the line of sight of person P wearing smart glasses 100A, and the image captured by camera 110 shows that only the face region of person A is detected.

[0133] In this case, the smart glasses 100A detect one face region from the image captured by the camera 110 and identify person A, who corresponds to this face region, as the speaker of interest. The smart glasses 100A then detect the lip region 22, perform multimodal speech recognition processing using acoustic features and lip features, and display the recognition result text data 23 on the display 130.

[0134] Figure 11 is a second diagram illustrating an example of smart glasses operation. Figure 11 shows a state where person A, designated as the speaker of interest, has moved out of the line of sight of the wearer P of the smart glasses 100A for a predetermined set time (e.g., 10 seconds).

[0135] In this case, the smart glasses 100A determine that person A is the speaker of interest based solely on person A's voice data, perform speech recognition processing using acoustic features calculated from the voice data, and display the recognition result text data 23a on the display 130.

[0136] Figure 12 is a third diagram illustrating an example of smart glasses operation. Figure 12 shows a state where a predetermined set time has elapsed since person A, designated as the speaker of interest, moved out of the line of sight of the wearer P of the smart glasses 100A.

[0137] In this case, the smart glasses 100A cancel the designation of person A as the focus speaker and return to the initial state where no focus speaker is designated. Therefore, nothing is displayed on the display 130.

[0138] Figure 13 is the fourth figure illustrating an example of smart glasses operation. In Figure 13, person A is positioned in the line of sight of person P wearing smart glasses 100A, and the image captured by camera 110 shows that only the face region of person A is detected, while the lip region of person A is not detected.

[0139] In this case, the smart glasses 100A detect one face region from the image captured by the camera 110 and identify person A, who corresponds to this face region, as the speaker of interest. Furthermore, since the lip region of person A cannot be detected, the smart glasses 100A performs speech recognition processing using acoustic features calculated from the audio data and displays the recognition result text data 23a on the display 130.

[0140] Figure 14 is the fifth figure illustrating an example of smart glasses operation. Figure 14 shows the state in which the face regions of person A and person B are detected in the image captured by camera 110.

[0141] In this case, the smart glasses 100A calculate the distance between the x-coordinate of the center point of the lip region of person A and the x-coordinate of the center point of the image captured by camera 110, and the distance between the x-coordinate of the center point of the lip region of person B and the x-coordinate of the center point of the image captured by camera 110.

[0142] Next, the smart glasses 100A select the person who is closer as the speaker of interest. In Figure 14, person A is selected as the speaker of interest. The smart glasses 100A then detect the lip region 22 of person A, perform multimodal speech recognition processing using acoustic features and lip features, and display the recognition result text data 23 on the display 130.

[0143] Figure 15 is the sixth figure illustrating an example of smart glasses operation. Figure 15 shows a case where, in the image captured by camera 110, the person closer to the center point of the image captured by camera 110 changes from person A to person B due to the wearer P's head movement, etc.

[0144] In this case, the smart glasses 100A detect the lip region 22B of person B, perform multimodal speech recognition processing using acoustic features and lip features, and display the recognition result text data 23B on the display 130.

[0145] Figure 16 is the seventh diagram illustrating an example of smart glasses operation. In Figure 16, the mouth of person A, who is closer to the center point of the image captured by camera 110, is obscured in the image captured by camera 110.

[0146] In this case, instead of using the x-coordinate of the center point of the lip region of person A, the smart glasses 100A find the x-coordinate of the center point of a portion of the facial image of person A, and calculate the distance between this x-coordinate and the x-coordinate of the center point of the image captured by the camera 110. Next, based on this distance, the smart glasses 100A determine that person A is the speaker of interest.

[0147] The smart glasses 100A then perform speech recognition processing using acoustic features calculated from the voice data of person A, and display the recognition result text data 23a on the display 130.

[0148] Thus, in this embodiment, even when there are multiple people in the line of sight of the wearer of the smart glasses 100A, or when the mouth of the person the wearer is focusing on is hidden, text data indicating the content of the speaker's speech can be displayed on the display 130.

[0149] Next, with reference to Figure 17, the speech content recognition capabilities of the smart glasses 100A in this embodiment will be described. Figure 17 is a diagram illustrating the speech content recognition of the first embodiment.

[0150] In this embodiment, the speech recognition processing unit 230 of the smart glasses 100A inputs a video extracted from the lip region of a person identified as a speaker of interest to the lip feature calculation unit 252 to obtain lip features 171. In this embodiment, the speech waveform of the person identified as a speaker of interest is input to the acoustic feature calculation unit 263 to obtain acoustic features 172.

[0151] Then, the speech recognition processing unit 230 combines the lip feature quantity 171 and the acoustic feature quantity 172 in the feature quantity integration unit 281 to obtain a multimodal feature quantity 173.

[0152] Next, the speech recognition processing unit 230 inputs the multimodal feature quantities 173 to the multimodal speech content recognition unit 282, generates text data representing the speech content using the multimodal speech content recognition model 283, and outputs the text data to the speech content recognition result output unit 233.

[0153] Furthermore, in this embodiment, if the lips are hidden or the image of the speaker in question is not included in the image captured by the camera 110 and the image within the lip region cannot be used, the acoustic feature quantity 172 extracted by the acoustic feature quantity calculation unit 263 is input to the speech content recognition unit 291. The speech content recognition unit 291 uses the speech content recognition model 292 to generate text data indicating the speech content and outputs the text data to the speech content recognition result output unit 233.

[0154] Thus, in this embodiment, the method of recognizing the speech content is switched depending on whether or not the lip region of the speaker of interest can be detected, thereby improving the accuracy of speech recognition.

[0155] Furthermore, in this embodiment, the lip feature calculation model 253, the multimodal speech content recognition model 283, and the speech content recognition model 292 are pre-trained models in which a neural network has been trained using video data of the lip region, audio data, and correct text data as training data.

[0156] Furthermore, in this embodiment, voice data is acquired for each utterance segment and the content of the utterance is recognized, but this is not limited to this. In this embodiment, for example, if voice data from multiple people is acquired simultaneously, the voice data of the speaker of interest may be selected based on the facial image of the person detected from the image data.

[0157] (Second Embodiment) A second embodiment will be described below with reference to the drawings. The second embodiment is a translation system that applies the smart glasses 100A of the first embodiment.

[0158] Figure 18 shows an example of the system configuration of a translation system according to the second embodiment. The translation system 500 of this embodiment includes smart glasses 100A and an automatic translation device 700. The smart glasses 100A and the automatic translation device 700 are connected, for example, via a network.

[0159] The automatic translation device 700 of this embodiment, upon receiving text data in a first language and a language selection, translates the text data in the first language into the selected language (second language) and outputs text data in the second language.

[0160] In the translation system 500 shown in Figure 18, the smart glasses 100A transmit text data resulting from the recognition of the speaker's utterance based on image data and audio data to the automatic translation device 700 as text data in the first language. At this time, the smart glasses 100A may accept the selection of a second language in advance. In that case, the smart glasses 100A transmits information indicating the second language along with the text data in the first language to the automatic translation device 700.

[0161] The automatic translation device 700 receives text data in the first language and information indicating the second language, converts the text data in the first language into text data in the second language, and transmits it to the smart glasses 100A.

[0162] The smart glasses 100A display the text data in the second language received from the automatic translation device 700 on the display 130.

[0163] In this embodiment, by linking the smart glasses 100A and the automatic translation device 700, the content of the speaker's speech can be displayed to the wearer of the smart glasses 100A in a second language different from the first language used by the speaker.

[0164] (Third embodiment) A third embodiment will be described below with reference to the drawings. The third embodiment is a meeting minutes creation system that applies the smart glasses 100A of the first embodiment.

[0165] Figure 19 shows an example of the system configuration of a speech content recording system according to the third embodiment. The speech content recording system 600 of this embodiment includes smart glasses 100A and a speech content recording device 700A. The smart glasses 100A and the speech content recording device 700A are connected, for example, via a network.

[0166] In this embodiment, the smart glasses 100A are used, for example, in educational institutions to store the content of a teacher's speech as text data during lectures. In this case, the wearer of the smart glasses 100A can store the content of the teacher's speech as text data in the storage unit of the speech content recording device 700A simply by directing their gaze toward the teacher T giving the lecture.

[0167] Furthermore, in this embodiment, it can also be used, for example, to save the speech content of a specific person as text data when multiple people are present on a stage set up in a space such as an auditorium.

[0168] In this embodiment, by linking the smart glasses 100A with the speech content recording device 700A, it is possible to save only the speech content of the speaker of interest as text data, even in situations where multiple people speak in a random order.

[0169] Furthermore, the smart glasses 100A can be applied to embodiments other than those described above. For example, the smart glasses 100A are useful when the wearer P has a hearing impairment.

[0170] Each of the functions of the embodiments described above can be realized by one or more processing circuits. Hereinafter, "processing circuit" as used herein includes processors programmed to execute each function by software, such as processors implemented by electronic circuits, as well as devices such as ASICs (Application Specific Integrated Circuits), DSPs (digital signal processors), FPGAs (field programmable gate arrays), and conventional circuit modules designed to execute each of the functions described above.

[0171] Furthermore, the apparatus described in the embodiments represents only one of several computing environments for carrying out the embodiments disclosed herein.

[0172] In one embodiment, the information processing device 200 (information processing terminal 200A) includes a plurality of computing devices, such as a server cluster. The plurality of computing devices are configured to communicate with each other via any type of communication link, including a network or shared memory, and perform the processing disclosed herein. Similarly, the information processing device 200 may include a plurality of computing devices configured to communicate with each other.

[0173] Furthermore, the information processing device 200 can be configured to share the disclosed processing steps in various combinations. For example, a process performed by the information processing device 200 may be performed by other information processing devices. Similarly, the functions of the information processing device 200 may be performed by other information processing devices. Also, each element of the information processing device and other information processing devices may be combined into a single information processing device or divided among multiple devices.

[0174] Although the present invention has been described above based on various embodiments, the present invention is not limited to the requirements shown in the above embodiments. These points can be modified as long as they do not impair the spirit of the present invention, and can be appropriately determined according to their application. [Explanation of symbols]

[0175] 100 Voice Recognition Systems 100A Smart Glasses 110 Camera 120 microphones 130 displays 200 Information Processing Devices 200A Information Processing Terminal 230 Speech Recognition Processing Unit 231 Video Input Section 232 Voice Input Section 233 Speech content recognition result output unit 240 Special Speaker Identification Department 250 Lip Feature Acquisition Unit 260 Acoustic Feature Acquisition Unit 270 Person Identification Unit 280 Multimodal Recognition Unit 290 Voice Recognition Unit [Prior art documents] [Patent Documents]

[0176] [Patent Document 1] Japanese Patent Publication No. 2021-026050

Claims

1. A speaker selection unit determines a speaker of interest based on the facial image of a person detected from the image data contained in the video data, A speech content recognition result output unit that displays text data converted from the speech data of the speaker of interest among the audio data included in the video data on a display device, A speaker embedding information calculation unit calculates speaker embedding information to identify the person who spoke the aforementioned audio data, A speech recognition device having an off-screen speaker estimation unit that, when the face image of the speaker of interest is not detected from the image shown in the image data included in the video data, estimates the speaker of the audio data based on speaker embedding information calculated from the audio data.

2. A lip region extraction unit detects a lip region, including an image of the lips, from the facial image of the person identified as the speaker of interest, The speech recognition device according to claim 1, further comprising a first speech recognition unit that converts the speech data of the speaker of interest into text data using image data showing an image within the lip region and the speech data of the speaker of interest.

3. If the lip region is not detected from the facial image, The speech recognition device according to claim 2, further comprising a second speech recognition unit that converts the speech data of the speaker of interest into text data using the speech data of the speaker of interest.

4. The display device is a spectacle-type display device, and the spectacle-type display device is equipped with an imaging device that captures an image in the direction of the wearer's line of sight when wearing the spectacle-type display device. The speech recognition device according to any one of claims 1 to 3, wherein the video data is video data acquired by the imaging device.

5. The aforementioned speaker selection unit is: The speech recognition device according to any one of claims 1 to 4, wherein when multiple face images are detected from the image data contained in the video data, the face image with the smallest distance between the center point of a portion of the face image and the center point of the image data contained in the video data is selected as the face image of the speaker of interest.

6. The off-screen speaker estimation unit, The speech recognition device according to claim 1, wherein the period during which the face image of the speaker of interest is not detected from the image data included in the video data is less than a predetermined set time, and the speaker embedding information calculated from the audio data acquired during the period in which the face image of the speaker of interest is not detected matches the speaker embedding information of the speaker of interest.

7. The off-screen speaker estimation unit, The speech recognition device according to claim 6, wherein if the state in which the face image of the speaker of interest is not detected from the image shown in the image data included in the video data continues for a predetermined set time or longer, the determination of the speaker of interest is canceled.

8. A speech recognition system comprising an information processing device, an imaging device capable of communicating with the information processing device, and a display device capable of communicating with the information processing device, The aforementioned information processing device is A speaker of interest determination unit determines a speaker of interest based on a facial image of a person detected from an image shown in the image data contained in the video data acquired by the aforementioned imaging device, A speech content recognition result output unit that displays text data converted from the speech data of the speaker of interest among the audio data included in the video data on the display device, A speaker embedding information calculation unit calculates speaker embedding information to identify the person who spoke the aforementioned audio data, A speech recognition system comprising: an off-screen speaker estimation unit that, when the face image of the speaker of interest is not detected from the image shown in the image data included in the video data, estimates the speaker of the audio data based on speaker embedding information calculated from the audio data.

9. A computer-based speech recognition method, wherein the computer, Based on the facial images of individuals detected from the image data contained in the video data, the speaker of interest is determined. The text data converted from the audio data of the speaker of interest, which is included in the video data, is displayed on the display device. Speaker embedding information is calculated to identify the person who spoke the aforementioned audio data. A speech recognition method that, when the face image of the speaker of interest is not detected from the image shown in the image data included in the video data, estimates the speaker of the audio data based on speaker embedding information calculated from the audio data.

10. Based on the facial images of individuals detected from the image data contained in the video data, the speaker of interest is determined. The text data converted from the audio data of the speaker of interest, which is included in the video data, is displayed on the display device. Speaker embedding information is calculated to identify the person who spoke the aforementioned audio data. A speech recognition program that, if the face image of the speaker of interest cannot be detected from the image data contained in the video data, causes a computer to perform a process to estimate the speaker of the audio data based on speaker embedding information calculated from the audio data.

11. A smart glasses comprising an information processing terminal, an imaging device connected to the information processing terminal, and a display device connected to the information processing terminal, The aforementioned information processing terminal is A speaker of interest determination unit determines a speaker of interest based on a facial image of a person detected from an image shown in the image data contained in the video data acquired by the aforementioned imaging device, A speech content recognition result output unit that displays text data converted from the speech data of the speaker of interest among the audio data included in the video data on the display device, A speaker embedding information calculation unit calculates speaker embedding information to identify the person who spoke the aforementioned audio data, Smart glasses comprising: an off-screen speaker estimation unit that, when the face image of the speaker of interest is not detected from the image shown in the image data included in the video data, estimates the speaker of the audio data based on speaker embedding information calculated from the audio data.

12. A translation system comprising: smart glasses having an information processing terminal, an imaging device connected to the information processing terminal, a display device connected to the information processing terminal, and a translation device capable of communicating with the smart glasses. The information processing terminal of the smart glasses is A speaker of interest determination unit determines a speaker of interest based on a facial image of a person detected from an image shown in the image data contained in the video data acquired by the aforementioned imaging device, A speech content recognition result output unit outputs to the translation device text data of a first language converted from the speech data of the speaker of interest among the audio data included in the video data, A speaker embedding information calculation unit calculates speaker embedding information to identify the person who spoke the aforementioned audio data, The system includes an off-screen speaker estimation unit that, if the face image of the speaker of interest is not detected from the image data included in the video data, estimates the speaker of the audio data based on speaker embedding information calculated from the audio data. A translation system comprising a translation device that displays text data in a second language, translated from text data in a first language, on a display device.

13. A speech content recording system comprising: an information processing terminal; an imaging device connected to the information processing terminal; a smart glasses having a display device connected to the information processing terminal; and a speech content recording device capable of communicating with the smart glasses. The information processing terminal of the smart glasses is A speaker of interest determination unit determines a speaker of interest based on a facial image of a person detected from an image shown in the image data contained in the video data acquired by the aforementioned imaging device, A speech content recognition result output unit that displays text data converted from the audio data of the speaker of interest among the audio data included in the video data on the display device and outputs the text data to the speech content recording device, A speaker embedding information calculation unit calculates speaker embedding information to identify the person who spoke the aforementioned audio data, The system includes an off-screen speaker estimation unit that, if the face image of the speaker of interest is not detected from the image data included in the video data, estimates the speaker of the audio data based on speaker embedding information calculated from the audio data. The aforementioned speech content recording device is A speech content recording system having a storage unit for storing the text data output from the information processing terminal.

Citation Information

Patent Citations

  • Imaging apparatus with voice input function and its voice recording method

    JP2009141555A

  • Person retrieval and registration system

    JP2010061265A

  • Speech section detecting device and speech recognition device, program and recording medium

    JP2011059186A

  • Eyeglass-type display device

    JP2012059121A

  • Image information retrieval server, image information retrieval method and user terminal

    JP2017033382A