Speech recognition system, communication system, speech recognition device, mobile control system, speech recognition method and program

The utterance recognition system addresses the challenge of recognizing specific speakers and speech content in multimodal systems by integrating lip and sound features, enhancing accuracy through error correction using mouth shape recognition.

JP7771590B2Active Publication Date: 2025-11-18RICOH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2021154862
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-09-22
Publication Date
2025-11-18
Estimated Expiration
2041-09-22

AI Technical Summary

Technical Problem

Conventional multimodal speech recognition systems struggle to recognize a specific speaker and specific speech content in a single process.

Method used

An utterance recognition system that records images and sounds accompanying utterances by one or more speakers, utilizing an utterance recognition device to perform multimodal recognition processing, recognizing a specific speaker and specific speech content by combining lip features and sound features, and correcting speech recognition errors using mouth shape recognition results.

Benefits of technology

Enables simultaneous recognition of a specific speaker and specific speech content, improving accuracy in noisy environments by correcting speech recognition errors with mouth shape recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007771590000001
    Figure 0007771590000001
  • Figure 0007771590000002
    Figure 0007771590000002
  • Figure 0007771590000003
    Figure 0007771590000003
Patent Text Reader

Abstract

To provide an utterance recognition system, a communication system, an utterance recognition device, a moving body control system, and an utterance recognition method and program that utilize multi-modal recognition to recognize an utterer and also recognize specific utterance details through single processing.SOLUTION: In an utterance recognition device of an utterance recognition system, a processing part comprises a lips feature quantity calculation part which uses a previously obtained lips feature quantity calculation model to calculate a lips feature quantity; a speech feature quantity calculation part which extracts a speech feature quantity from a speech waveform input at a speech input part; a feature quantity integration part which combines the speech feature quantity and lips feature quantity together to obtain a multi-modal feature quantity; and a multi-modal recognition part which uses the obtained multi-modal recognition model to perform multi-modal recognition.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a speech recognition system, a communication system, a speech recognition device, a mobile object control system, an utterance recognition method, and a program. [Background technology]

[0002] To reduce the effects of acoustic noise, there is a speech recognition technology that converts speech signals into text using multimodal speech recognition that uses not only the spoken speech signal but also lip movement images during speech. In this multimodal speech recognition, a multimodal voice activity detection technology is known that identifies the actual speech interval by comprehensively using acquired speech information and image information (see, for example, Patent Document 1). Summary of the Invention [Problem to be solved by the invention]

[0003] However, conventional techniques have had a problem in that speech recognition processing using multimodal recognition cannot recognize a specific speaker and a specific speech content in a single process. [Means for solving the problem]

[0004] In order to solve the above-mentioned problem, the invention of claim 1 provides an utterance recognition system including a recording device that records images and sounds accompanying utterances by one or more speakers, and an utterance recognition device that receives image information related to the images and sound information related to the sounds transmitted by the recording device and recognizes the content of the utterances, wherein the utterance recognition device has: processing means that recognizes a specific speaker from the one or more speakers and recognizes specific content of the utterance uttered by the specific speaker by performing multimodal recognition processing in parallel using lip features obtained based on each lip image that changes when the one or more speakers speak and sound features obtained based on each sound uttered by the one or more speakers; and transmission means that transmits the specific content of the utterance identified by the processing means to a display device.the processing means recognizes a mouth shape based on the lip feature amount to obtain a mouth shape recognition result, recognizes a voice based on the voice feature amount to obtain a voice recognition result, adopts a voice recognition result corresponding to the mouth shape recognition result of the specific speaker from among the voice recognition results of the one or more speakers, and recognizes the specific speech content by deleting an insertion error included in the adopted voice recognition result using the mouth shape recognition result of the specific speaker. The present invention provides a speech recognition system. [Effects of the Invention]

[0005] As described above, according to the present invention, in a speech recognition process using multimodal recognition, it is possible to perform recognition of a specific speaker and recognition of a specific speech content in a single process. [Brief explanation of the drawings]

[0006] [Figure 1] FIG. 1 is a diagram illustrating an example of an application scene of this embodiment. [Figure 2] FIG. 1 is a diagram illustrating an example of the overall configuration of a communication system. [Figure 3] FIG. 2 is a diagram illustrating an example of a hardware configuration of a data acquisition device. [Figure 4] FIG. 2 is a diagram illustrating an example of the hardware configuration of an utterance recognition device, a display device, and an utterance content management server. [Figure 5] FIG. 1 illustrates an example of a functional configuration of a communication system. [Figure 6] FIG. 10 is a diagram illustrating an example of a functional configuration at the time of pre-stage combination in multimodal speech recognition processing. [Figure 7] FIG. 10 is a conceptual diagram illustrating an example of a mouth shape pattern management table. [Figure 8] FIG. 10 is a conceptual diagram illustrating an example of an utterance recognition result management table. [Figure 9] FIG. 4 is a sequence diagram illustrating an example of an overall process according to the first embodiment. [Figure 10] FIG. 2 is a schematic diagram showing an example of processing at the time of pre-stage combination in the multimodal speaker recognition system according to the first embodiment. [Figure 11] 10 is an example of a screen display on a display device showing a speech recognition result. [Figure 12] FIG. 10 is a diagram illustrating an example of a functional configuration at the time of post-stage combination in multimodal speech recognition processing. [Figure 13] FIG. 10 is a sequence diagram illustrating an example of an overall process according to the second embodiment. [Figure 14] 10 is an overall flowchart showing an example of multimodal processing according to the second embodiment. [Figure 15] 10 is a flowchart showing a process of outputting a multimodal recognition result according to the second embodiment. [Figure 16] FIG. 10 is a schematic diagram showing an example of processing at the time of post-stage coupling in the multimodal speaker recognition system according to the second embodiment. [Figure 17] 1 is a diagram illustrating an example of the overall configuration of a mobile object control system. DETAILED DESCRIPTION OF THE INVENTION

[0007] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. In the description of the drawings, the same elements are given the same reference numerals, and if there are overlapping parts, their description will be omitted.

[0008] [First embodiment] The first embodiment will be described with reference to FIGS.

[0009] [System Overview] <Example of application scenario> FIG. 1 is a diagram showing an example of an application scenario of this embodiment. FIG. 1 shows a state in which a recording device 2 and an utterance recognition device 3, which is an example of an information processing device, are placed on a conference table in a predetermined conference attended by, for example, four participants A, B, C, and D. Here, the recording device 2 and the utterance recognition device 3 are connected to each other by wire or wirelessly. Furthermore, the utterance recognition device 3 is connected to a display device 5 by a wired cable or the like. The display device 5 is, for example, a display device such as an interactive whiteboard (IWB), and is capable of displaying the contents of utterances, images, minutes, etc., in the predetermined conference. Note that the speech recognition device 3 and the display device 5 may be connected wirelessly.

[0010] [Overall configuration of communication system] <System configuration example> Fig. 2 is a diagram showing an example of the overall configuration of a communication system. As shown in Fig. 2, the communication system 1 includes a recording device 2, an utterance recognition device 3, a display device 5, and an utterance content management server 6, and each device and server are connected to each other via a communication network 100. The communication system 1 includes an utterance recognition system 4 configured by the recording device 2 and the utterance recognition device 3.

[0011] The communication network 100 is a communication network through which an unspecified number of communications are conducted, and is constructed using the Internet, an intranet, a LAN (Local Area Network), etc. The communication network 100 may include not only wired communication but also communication networks using wireless communication such as 3G (3rd Generation), 4G (4th Generation), 5G (5th Generation), WiMAX (Worldwide Interoperability for Microwave Access), and LTE (Long Term Evolution).

[0012] The speech recognition process using the multimodal recognition of this embodiment is performed using the above-described communication system 1 as an example. Each device constituting the communication system 1 will be described below.

[0013] <Recording device> The recording device 2 is a device that captures, for example, images of one or more participants participating in a predetermined event such as a conference, scenery, etc., and is equipped with a microphone that collects speech uttered by the participants. The recording device 2 is a device that has multiple built-in cameras (digital cameras) to simultaneously capture images of one or more participants. Furthermore, instead of having multiple built-in cameras, the recording device 2 may be an omnidirectional camera (also called an omnidirectional imaging device) that can capture spherical images (videos). The recording device 2 can communicate with the speech recognition device 3, the display device 5, and the speech content management server 6 via the communication network 100, but in the first embodiment, it is connected to the speech recognition device 3 by wire and transmits image information related to the captured images and audio information related to the collected audio to the speech recognition device 3.

[0014] <Speech recognition device> The speech recognition device 3 is realized by a computer system for communication equipped with a general OS or the like. The speech recognition device 3 is capable of communicating with the recording device 2, the display device 5, and the speech content management server 6 via a communication network 100. However, in the first embodiment, the speech recognition device 3 is connected to the recording device 2 by wire, and receives image information relating to images captured by the recording device 2 and audio information relating to collected audio. In addition, a browser application for communicating with the display device 5, for example, is installed in the speech recognition device 3.

[0015] The speech recognition device 3 may be a communication terminal having a communication function, such as a commonly used PC (Personal Computer), a portable notebook PC, a mobile phone, a smartphone, a tablet terminal, or a wearable terminal (sunglasses type, wristwatch type, etc.). The speech recognition device 3 may further be a communication device or a communication terminal capable of running software such as browser software.

[0016] <Display device> The display device 5 is a device that visualizes the content of an utterance by displaying, for example, image (video) information, text information, etc. transmitted by the speech recognition device 3, and is a general display terminal such as an electronic whiteboard.

[0017] <Speech content management server> The speech content management server 6 is realized by an information processing device (computer system) equipped with a general server OS or the like. The speech content management server 6 receives the speech content recognized and converted into text by the speech recognition device 3 via the communication network 100 instead of displaying it on the display device 5, and stores the received text information in a predetermined storage area. Furthermore, the speech content management server 6 may process the received text information and output it as conference minutes after the end of a predetermined event such as a conference. Note that even when the speech content recognized and converted into text by the speech recognition device 3 is displayed on the display device 5, the speech content management server 6 may receive the converted text in parallel and store the received text information in a predetermined storage area.

[0018] The speech content management server 6 may be a communication terminal with a communication function, such as a commonly used PC (Personal Computer), a portable notebook PC, or a tablet terminal. In this case, the speech content management server 6 may be constructed by a single computer, or may be constructed by multiple computers to which each section (function or means) such as storage is divided and arbitrarily assigned. Furthermore, all or part of the functions of the speech content management server 6 may be a server computer existing in a cloud environment, or may be a server computer existing in an on-premise environment.

[0019] [Hardware configuration] Next, the hardware configuration of the device or server constituting the communication system according to the first embodiment will be described with reference to Figures 3 and 4. Note that components (hardware resources) may be added or deleted from the hardware configuration of the device or server shown in Figures 3 and 4 as needed.

[0020] <Hardware configuration of the data acquisition device> First, the hardware configuration of the data collection device will be described with reference to Fig. 3. Fig. 3 is a diagram showing an example of the hardware configuration of the data collection device. As shown in Fig. 3, the data collection device 2 includes, for example, a CPU (Central Processing Unit) 201, a ROM (Read Only Memory) 202, a RAM (Random Access Memory) 203, an EEPROM (Electrically Erasable and Programmable ROM) 204, a short-range communication I / F 208, a CMOS sensor 209, an image sensor I / F 210, a network I / F 211, a touch panel 212, a media I / F 215, an external device connection I / F 216, an audio input / output I / F 217, a microphone 218, a speaker 219, and a bus line 220.

[0021] Of these, the CPU 201 controls the overall operation of the data collection device 2. The ROM 202 stores programs used in the processing of the CPU 201. The RAM 203 is used as a work area for the CPU 201. The EEPROM 204 reads and writes various data such as applications under the control of the CPU 201. The short-range communication I / F 208 is a communication circuit for performing short-range wireless communication with a communication device or communication terminal equipped with a wireless communication interface such as NFC (Near Field Communication), Bluetooth (registered trademark, omitted hereinafter), millimeter wave wireless communication, Wi-Fi (registered trademark, omitted hereinafter), QR Code (registered trademark, omitted hereinafter), visible light, ambient sound, or ultrasonic wave.

[0022] The CMOS sensor 209 is a type of built-in imaging means that captures an image of a subject and obtains image data or video data under the control of the CPU 201. The imaging means may be an imaging means configured with a CCD (Charge Coupled Device) sensor or the like instead of a CMOS sensor. The imaging element I / F 210 is a circuit that controls the driving of the CMOS sensor 209.

[0023] Furthermore, as described above, the recording device 2 may be a spherical (omnidirectional) camera instead of a device incorporating multiple cameras (digital cameras). In this case, the recording device 2 may have two CMOS sensors 209 and two imaging element I / Fs 210. When the recording device 2 is a spherical (omnidirectional) camera, the imaging unit includes two wide-angle lenses (so-called fisheye lenses) each with a field of view of 180° or more for forming a hemispherical image, and two imaging elements corresponding to each wide-angle lens. Each imaging element includes an image sensor such as a CMOS sensor or CCD (Charge Coupled Device) sensor that converts optical images captured by the two fisheye lenses into image data electrical signals and outputs the image data, a timing generation circuit that generates horizontal or vertical synchronization signals and pixel clocks for the image sensor, and a group of registers in which various commands, parameters, etc. required for the operation of the imaging element are set.

[0024] The network I / F 211 is a communication interface for communicating various data (information) with other devices via the communication network 100. The touch panel 212 is a type of display and operation means, such as a liquid crystal display (LCD) or organic electroluminescence (EL) display, that displays images, characters, various icons, etc. The media I / F 215 controls the reading and writing (storage) of data from and to a recording medium 214, such as a flash memory. The external device connection I / F 216 is an interface for connecting various external devices. In this case, the external device is, for example, a USB (Universal Serial Bus) memory. The sound input / output I / F 217 is a circuit that processes the input and output of sound signals between the microphone 218 and the speaker 219 under the control of the CPU 201. The microphone 218 is a built-in circuit that converts sound into an electrical signal and acquires information using the electrical signal by capturing voices or sound waves emitted from an external speaker or the like. The speaker 219 is a built-in circuit that converts electrical signals into physical vibrations to generate sounds such as music and voices. The bus line 220 is an address bus, a data bus, etc. for electrically connecting the components such as the CPU 201.

[0025] <Hardware configuration of speech recognition device, display device, and speech content management server> Next, the hardware configurations of the speech recognition device, the display device, and the speech content management server will be described with reference to Fig. 4. Fig. 4 is a diagram showing an example of the hardware configurations of the speech recognition device, the display device, and the speech content management server. As shown in Fig. 4, the speech recognition device 3 is constructed by, for example, a computer. The speech recognition device 3 includes, for example, a CPU 301, a ROM 302, a RAM 303, an EEPROM 304, a HD 305, an HDD controller 306, a display 307, a short-range communication I / F 308, a CMOS sensor 309, an image sensor I / F 310, a network I / F 311, a keyboard 312, a pointing device 313, a media I / F 315, an external device connection I / F 316, an audio input / output I / F 317, a microphone 318, a speaker 319, and a bus line 320.

[0026] Of these, CPU 301, ROM 302, RAM 303, EEPROM 304, short-range communication I / F 308, CMOS sensor 309, image sensor I / F 310, network I / F 311, media I / F 315, external device connection I / F 316, sound input / output I / F 317, microphone 318, and speaker 319 are similar to the respective components of CPU 201, ROM 202, RAM 203, EEPROM 204, short-range communication I / F 208, CMOS sensor 209, image sensor I / F 210, network I / F 211, media I / F 215, external device connection I / F 216, sound input / output I / F 217, microphone 218, and speaker 219 of the recording device 2 shown in Figure 2, so their explanation will be omitted.

[0027] The HD 305 stores various data such as programs. The HDD controller 306 controls the reading and writing of various data from and to the HD 305 under the control of the CPU 301. The display 307 is a type of display means such as a liquid crystal display or organic electroluminescence (EL) display that displays images, characters, various icons, etc. The keyboard 312 is a type of input means that has multiple keys for inputting characters, numbers, various instructions, etc. The pointing device 313 is a type of input means that selects and executes various instructions, selects processing targets, moves the cursor, etc.

[0028] 4, the display device 5 is constructed by, for example, a computer. The display device 5 includes, for example, a CPU 501, a ROM 502, a RAM 503, an EEPROM 504, a HD 505, a HDD controller 506, a display 507 as an example of a display means, a short-range communication I / F 508, a network I / F 511, a pointing device 513, a media I / F 515, an external device connection I / F 516, and a bus line 520. These hardware resources are similar to the respective components of the CPU 301, ROM 302, RAM 303, EEPROM 304, a HD 305, a HDD controller 306, a display 307, a short-range communication I / F 308, a network I / F 311, a pointing device 313, a media I / F 315, and an external device connection I / F 316 of the speech recognition device 3 shown in FIG. 4, and therefore, description thereof will be omitted.

[0029] 4, the utterance content management server 6 is constructed by, for example, a computer. The utterance content management server 6 includes, for example, a CPU 601, a ROM 602, a RAM 603, an EEPROM 604, a HD 605, an HDD controller 606, a short-range communication I / F 608, a network I / F 611, a keyboard 612, a pointing device 613, a media I / F 615, an external device connection I / F 616, and a bus line 620. These hardware resources are similar to the respective components of the CPU 301, ROM 302, RAM 303, EEPROM 304, HD 305, HDD controller 306, short-range communication I / F 308, network I / F 311, keyboard 312, pointing device 313, media I / F 315, and external device connection I / F 316 of the utterance recognition device 3 shown in FIG. 4, and therefore, description thereof will be omitted.

[0030] Furthermore, the above-mentioned program may be recorded on a computer-readable recording medium as an installable or executable file, or may be distributed by downloading via a network. Examples of recording media include CD-Rs (Compact Disc Recordables), DVDs (Digital Versatile Disks), Blu-ray Discs, SD cards, and USB memory. Furthermore, the recording media may be provided domestically or internationally as a program product. For example, the speech recognition device 3 realizes the speech recognition method according to the present invention by executing the program according to the present invention.

[0031] [Functional configuration of communication system] Next, the functional configuration of this embodiment will be described with reference to Fig. 5 to Fig. 8. Fig. 5 is a diagram showing an example of the functional configuration of a communication system.

[0032] <Functional configuration of the data acquisition device> As shown in Fig. 5, the recording device 2 includes a transmitting / receiving unit 21, an operation receiving unit 22, an imaging unit 23, a sound input / output unit 24, and a storage / readout unit 29. Each of these functional units is a function or means realized by operating any of the hardware resources shown in Fig. 3 in response to an instruction from the CPU 201 in accordance with a program for the recording device 2 that is loaded from at least one of the ROM 202 and the EEPROM 204 to the RAM 203. The recording device 2 also includes a storage unit 2000 constructed by at least one of the ROM 202 and the EEPROM 204 shown in Fig. 3. The storage unit 2000 stores a communication program (communication application) for communicating with the speech recognition device 3, a communication program (communication application) for communicating with the display device 5 and the utterance content management server 6 via the communication network 100, and the like.

[0033] <<Functional configuration of the data acquisition device>> Next, each functional configuration of the recording device 2 will be described in detail. The transmitting / receiving unit 21 of the recording device 2 shown in Fig. 5 is mainly realized by processing of the CPU 201 on the network I / F 211 and the short-range communication I / F 208. The transmitting / receiving unit 21 transmits and receives captured image (video) data and sound (audio) data to and from the utterance recognition device 3 via, for example, a wired cable. Furthermore, the transmitting / receiving unit 21 can also transmit and receive various data (or information) to and from the display device 5 and the utterance content management server 6 via the communication network 100. In this embodiment, the transmitting / receiving unit 21 functions as an example of at least one of a transmitting means and a receiving means.

[0034] The operation reception unit 22 is mainly realized by processing of the CPU 201 on the touch panel 212, and receives various inputs, settings, and other operations for the data collection device 2. The operation reception unit 22 also receives operations for capturing an image of a subject and collecting sound. In this embodiment, the operation reception unit 22 functions as an example of an operation reception means.

[0035] The imaging unit 23 is mainly realized by the processing of the CPU 201 on the CMOS sensor 209 and the image sensor I / F 210, and captures an image (video) by capturing the faces, etc. of one or more participants present in a space such as a conference room as the subjects. In this embodiment, the imaging unit 23 functions as an example of an imaging means.

[0036] The sound input / output unit 24 is mainly realized by processing by the CPU 201 on the microphone 218, speaker 219, and sound input / output I / F 217, and performs processing to collect speech sounds uttered by one or more participants present in a space such as a conference room, ambient sounds generated in the space, etc. using the microphone 218 and convert them into sound (audio) data. The sound input / output unit 24 further performs processing to convert predetermined sound (audio) data into a sound (audio) signal and output it from the speaker 219. In this embodiment, the sound input / output unit 24 functions as an example of a sound input / output means.

[0037] The memory / reading unit 29 is mainly realized by the processing of the CPU 201 on at least one of the ROM 202 and the EEPROM 204, and stores various data (or information) in the memory unit 2000 and reads various data (or information) from the memory unit 2000. In this embodiment, the memory / reading unit 29 functions as an example of a memory / reading means.

[0038] <Functional configuration of the speech recognition device> As shown in Fig. 5, the speech recognition device 3 includes a transmitting / receiving unit 51, an operation receiving unit 32, an acquiring unit 33, a display control unit 34, a processing unit 35, and a storage / reading unit 39. Each of these functional units is a function or means realized when any of the hardware resources shown in Fig. 4 operates in response to an instruction from the CPU 301 in accordance with a program for the utterance recognition device 3 that is loaded from at least one of the ROM 302, the EEPROM 304, and the HD 305 to the RAM 303. The speech recognition device 3 also includes a storage unit 3000 that is constructed using at least one of the ROM 302, the EEPROM 304, and the HD 305 shown in Fig. 4. The storage unit 3000 stores a communication program (communication application) for communicating with the recording device 2, and communication programs (communication applications) for communicating with the display device 5 and the utterance content management server 6 via the communication network 100. The storage unit 3000 further stores a lip feature calculation model, a multimodal recognition model, a mouth shape recognition model, a voice recognition model, and the like, which are used in the multimodal speech recognition process.

[0039] <<Functional configuration of the speech recognition device>> Next, each functional component of the speech recognition device 3 will be described in detail. The transmission / reception unit 31 of the speech recognition device 3 shown in Fig. 5 is mainly realized by processing of the CPU 301 on the network I / F 311 and the short-range communication I / F 308. The transmission / reception unit 31 transmits and receives captured image (video) data and sound (audio) data to and from the recording device 2 via, for example, a wired cable. Furthermore, the transmission / reception unit 31 transmits and receives various data (or information) to and from the display device 5 and the utterance content management server 6 via the communication network 100. In this embodiment, the transmission / reception unit 31 functions as an example of at least one of a transmitting means and a receiving means.

[0040] The operation reception unit 32 is mainly realized by processing of the CPU 301 on the keyboard 312 and the pointing device 313, and receives operations such as various inputs and settings in the speech recognition device 3. In this embodiment, the operation reception unit 22 functions as an example of an operation reception means.

[0041] The acquisition unit 33 is mainly realized by the processing of the CPU 301, and acquires, for example, the captured image (video) data and sound (audio) data transmitted by the recording device 2 via the transmission / reception unit 31. In this embodiment, the acquisition unit 33 functions as an example of an acquisition means.

[0042] The display control unit 34 is mainly realized by processing of the CPU 301 on the display 307, and controls the display of various screens and information (data) in the speech recognition device 3. In this embodiment, the display control unit 34 functions as an example of a display control means.

[0043] The processing unit 35 is mainly realized by the processing of the CPU 301, and controls the overall processing related to the multimodal speech recognition processing in the speech recognition device 3. In this embodiment, the processing unit 35 performs multimodal recognition processing in parallel using lip features obtained based on each lip image that changes when one or more speakers speak, and speech features obtained based on each speech uttered by one or more speakers. In this way, the processing unit 35 recognizes a specific speaker from one or more speakers and recognizes the specific speech content uttered by the specific speaker. In this embodiment, the processing unit 35 functions as an example of a processing means.

[0044] <Functional configuration of multimodal speech recognition (pre-combination)> FIG. 6 is a diagram showing an example of a functional configuration during early fusion in multimodal speech recognition processing. Early fusion is a method of combining lip features and speech features at the stage described below, and is also called early fusion. The processing unit 35 described above has a detailed functional configuration as shown in FIG. 6. Specifically, the processing unit 35 includes an image input unit 351, an image conversion unit 352, a face area recognition unit 353, a lip area extraction unit 354, a lip pixel count conversion unit 355, a lip feature calculation unit 357, and a mouth shape recognition unit 358. The lip feature calculation unit 357 and the mouth shape recognition unit 358 constitute a machine lip reading pre-training unit 359. The functional configuration described above is a functional configuration related to processing image (video) data collected by the collection device 2.

[0045] The processing unit 35 further includes a speech input unit 371, a speech feature calculation unit 373, a feature integration unit 374, a first multimodal recognition unit 375, and an utterance recognition result output unit 376. The first multimodal recognition unit 375 performs multimodal recognition using a previously obtained multimodal recognition model 372 and outputs the recognition result to the utterance recognition result output unit 376. Specifically, the first multimodal recognition unit 375 receives multimodal features combining lip features and speech features and outputs a sequence of hiragana characters. The above-described functional configuration is related to processing the sound (speech) data collected by the collection device 2.

[0046] <<Details of Multimodal Speech Recognition (Pre-Combination) Function>> Next, a detailed description will be given of each function constituting the processing unit 35. First, the image input unit 351 inputs the captured image (video) data acquired by the acquisition unit 33 described above.

[0047] For example, if the image (video) data input by the image input unit 351 has conditions (parameters and specifications) of 1920 x 1080 pixels and 30 fps, the image conversion unit 352 converts the image (video) data acquired under these conditions into a frame image sequence. The image conversion unit 352 also converts RGB images to grayscale and converts the number of pixels to speed up processing.

[0048] The face area recognition unit 353 recognizes the face areas of multiple participants in consecutive frame images acquired from the acquired video of participants who have taken part in a predetermined event such as a conference.

[0049] The lip area extraction unit 354 extracts detailed coordinates of facial features such as the mouth from the face area recognized by the face area recognition unit 353. The lip area extraction unit 354 may use a model that has been trained in advance using a neural network or the like with a large amount of data. The lip area extraction unit 354 may also use existing technology such as Dlib that has machine learning capabilities.

[0050] Lip pixel number conversion unit 355 converts the number of lip pixels to a predetermined image size. The size of the lip area extracted by lip area extraction unit 354 varies depending on the distance between recording device 2 and the participants in a conference or the like. Therefore, lip pixel number conversion unit 355 converts the image size to the size when the lip area was trained by mouth shape recognition model 361, and enlarges or reduces the image to a uniform size such as 150 x 150 pixels, for example. This is the image size when mouth shape recognition model 361 was trained.

[0051] Lip feature calculation unit 357 calculates lip features using lip feature calculation model 356 obtained in advance, and outputs the calculation results to mouth shape recognition unit 358 and feature integration unit 374, which will be described later. The calculated lip features represent features extracted from the lip feature calculation model that has been trained using the mouth shape pattern of a specific speaker out of one or more speakers as the correct answer.

[0052] The mouth shape recognition unit 358 recognizes the lip feature calculated by the lip feature calculation unit 357 as a sequence of mouth shape patterns. Note that the mouth shape recognition unit 358 is used only when the lip feature calculation unit 357 and the mouth shape recognition unit 358 are pre-trained, using the sequence of mouth shape patterns as correct labels. In this embodiment, Japanese is converted into hiragana and then converted into a corresponding mouth shape pattern. The correspondence between hiragana and mouth shape patterns will be explained in FIG. 7. Note that mouth shape refers to the shape of the participant's mouth (or the shape of the lips), and in this embodiment will be simply referred to as "mouth shape."

[0053] Based on the above, the machine lip reading pre-training unit 359 pre-trains the lip feature calculation unit 357 and the mouth shape recognition unit 358 using a sequence of mouth shape patterns as correct labels. However, the mouth shape recognition unit 358 is used during pre-training but is not used during speech recognition. Furthermore, by pre-training with mouth shape patterns, it is expected that the learning efficiency of lip features will improve.

[0054] The audio input unit 371 inputs sound (audio) data acquired by the above-described acquisition unit 33. In this embodiment, for example, the input conditions are monaural uncompressed data sampled at 16 kHz and 16 bits.

[0055] The speech feature calculation unit 373 extracts logarithmic mel filter bank features from the speech waveform input by the speech input unit 371 using, for example, a Hamming window with a width of 25 ms and a shift of 11 ms.

[0056] The feature integration unit 374 uses the speech features extracted by the speech feature calculation unit 373 and the lip feature calculation model trained as described above to combine the calculated lip features to obtain multimodal features. That is, the feature integration unit 374 recognizes a specific speaker and a specific utterance by combining the lip features and the speech features. The feature integration unit 374 further recognizes a specific utterance by combining the mouth shape pattern sequence recognition result, obtained by recognizing a mouth shape pattern sequence of a specific speaker from the lip features, with the speech recognition result. The feature integration unit 374 further obtains multimodal features for multimodal recognition processing by achieving temporal alignment according to the ratio of the lip features and the speech features per frame rate of a lip image sequence representing one utterance among the utterances spoken by one or more speakers. Using the pre-trained parameters, the lip features output from the lip feature calculation unit 357, which have already been trained, are used to generate multimodal features. Thereafter, the first multimodal recognition unit 375 learns the correct labels for hiragana.

[0057] The first multimodal recognition unit 375 performs fine-tuning including the lip feature amount calculation unit 357 to train the multimodal recognition model 372 .

[0058] The speech content recognition result output unit 376 outputs the speech content that has been multimodally recognized to the outside.

[0059] 5, the memory readout unit 39 is mainly realized by the processing of the CPU 301 on at least one of the ROM 302, the EEPROM 304, and the HD 305, and stores various data (or information) in the memory unit 3000 and reads various data (or information) from the memory unit 3000. In this embodiment, the memory readout unit 39 functions as an example of a memory readout means.

[0060] ●Mouth Shape Pattern Management Table● FIG. 7 is a conceptual diagram showing an example of a mouth shape pattern management table. A mouth shape pattern management DB 3001 configured with the mouth shape pattern management table shown in FIG. 7 is created in the storage unit 3000. The mouth shape pattern management table manages the hiragana characters corresponding to each mouth shape pattern. Among these, mouth shape patterns include "-A," "IA," "UA," "XA," and "-I." These mouth shape patterns represent the shape of a participant's mouth (mouth shape) in two states: an initial mouth shape at the beginning of movement and a final mouth shape at the end of movement. For example, "-" indicates that there is no initial mouth shape. Furthermore, "X" is an initial mouth shape representing a closed lip mouth shape. Since a geminate consonant and a nasal consonant cannot be defined as a mouth shape, they are represented as "-*" with neither an initial mouth shape nor a final mouth shape.

[0061] For example, if the correct label is "I twisted all reality to my own.", a sample sentence from the ATR Phoneme Balance 503 sentence published by the Speech Resource Consortium (http: / / research.nii.ac.jp / src / ATR503.html), "-A, IA, -U, -U, -E, -*, -I, -U, -O, -U, XE, IE, -I, -U, -*, UO, -O, -U, -E, IE, -I, XA, -E, IA, UO, IA" At this time, the particle "へ" is converted to "え" and the period is also removed. This is used as the correct label when training the mouth shape recognition model 361 of the mouth shape recognition unit 358, so the mouth shape recognition result is also a series of this mouth shape pattern, with reference to the following non-patent document: Non-patent literature: Miyazaki Tsuyoshi; Nakajima Toyoshiro. Coding characteristic mouth shapes during Japanese speech and proposing a method for displaying information on mouth shape changes. Transactions of the Institute of Electrical Engineers of Japan, C (Electronics, Information and Systems Division), 2009, 129.12: 2108-2114. The pre-learning using the mouth shape pattern management table described above is performed, for example, when the system is developed, but data may be collected in the user's environment and used for pre-learning. Specifically, data may be collected using the recording device 2 and speech recognition device 3 in the communication system 1, and learning (for Japanese) may be performed using that data. Furthermore, by appropriately changing the contents of the correspondence table in the mouth shape pattern management table, it can also be applied to foreign languages.

[0062] The mouth shape pattern management DB 3001 is used as a correspondence when calculating the commonly known Levenshtein distance. Furthermore, the mouth shape pattern management DB 3001 is also used in the case of subsequent connection in the speech recognition processing in the second embodiment, which will be described later.

[0063] ●Speech Recognition Result Management Table● Fig. 8 is a conceptual diagram showing an example of an utterance recognition result management table. In the storage unit 3000, an utterance recognition result management DB 3002 configured by an utterance recognition result management table such as that shown in Fig. 8 is constructed. In the utterance recognition result management table, a No. (1), (2), etc. is given to each sound (in Fig. 8, the part written with a circled number. Hereinafter, specific examples of No. will be written in the form of a number in parentheses). For each No., the correct answer, mouth shape, mouth shape recognition result, voice recognition result, operation, and correction result are associated with each other and stored and managed as the speech recognition result.

[0064] Among these, the correct answer represents the correct label for speech recognition, and is registered as the correct label after speech recognition processing is performed on the speech actually spoken by each participant. In the example shown in Figure 7, for example, the Speech Resource Consortium: The hiragana version of the above sentence is shown when the system recognizes the sample sentence "The fingers on both hands were deformed, and the joints were raised like bumps" from the ATR Phoneme Balance 503 sentence sample available at http: / / research.nii.ac.jp / src / ATR503.html.

[0065] The lip shape recognition result represents, for example, the lip shape recognition result recognized in a moving image sequence representing the above-mentioned utterance, "The fingers of both hands were deformed, and the joints were bulging like bumps."

[0066] The speech recognition result indicates the recognition result of the speech uttered by the participant.

[0067] The operation shows the operation when calculating the commonly used Levenshtein distance, and if the hiragana in the speech recognition result is mistakenly "inserted" into the mouth shape pattern in the mouth shape recognition result, "INS" is given; if it is deleted, "DEL" is given; if it is replaced, "SUB" is given; and if it is correct, "OK" is given.

[0068] The corrected result is given as the speech recognition result corrected by the speech content recognition result correction unit 382, ​​which operates in a later stage of the speech recognition process described later. In this embodiment, assuming recognition in a high-noise environment, where the reliability of speech recognition accuracy is reduced, the mouth shape recognition result is considered to be correct, and the speech recognition result is corrected using the mouth shape recognition result.

[0069] Before and after the speech recognition results for Nos. (1), (2), (39), (40), (41), (42), (43), (44), and (45), characters unrelated to the actual speaker's speech are output by the speech recognition due to the influence of surrounding participants' voices and noise other than the actual speaker (specific participant). In this case, the actual speaker's mouth shape is not present, so the speech recognition results are not output as they are unrelated to the actual speaker's speech. In other words, when there is no mouth shape recognition result (****) and the operation is "INS," the characters in the speech recognition results are deleted. Also, as in cases (16) and (21), when mouth shapes are output during the participant's speech but there is no speech recognition output, they are deleted. This correction allows for correction of insertion errors caused by surrounding noise, which adversely affects speech recognition, using the mouth shape recognition results. Furthermore, it also allows for detection of speech sections by the actual speaker. After this, the speech recognition device 3 can also generate sentences containing kanji using IME technology, etc. In this embodiment, the speech recognition result management table is created for each participant in an event such as a conference, and actual speakers and speech segments are identified.

[0070] The speech recognition result management DB 3002 is a DB used in the case of subsequent combination in the speech recognition process in the second embodiment, which will be described later.

[0071] <Functional configuration of the display device> As shown in Fig. 5, the display device 5 includes a transmitting / receiving unit 51, an operation receiving unit 52, a display control unit 54, a generating unit 57, and a storage / reading unit 59. Each of these functional units is a function or means realized when any of the hardware resources shown in Fig. 4 operates in response to an instruction from the CPU 501 in accordance with a program for the display device 5 that is loaded from at least one of the ROM 502, the EEPROM 504, and the HD 505 to the RAM 503. The display device 5 also includes a storage unit 5000 constructed by at least one of the ROM 502, the EEPROM 504, and the HD 505 shown in Fig. 4. The storage unit 5000 stores a communication program (communication application) and the like for communicating with the recording device 2, the speech recognition device 3, and the utterance content management server 6 via the communication network 100.

[0072] <<Functional configuration of the display device>> Next, each functional configuration of the display device 5 will be described in detail. The transmission / reception unit 51 of the display device 5 shown in Fig. 5 is mainly realized by processing of the CPU 501 on the network I / F 511 and the short-range communication I / F 508. The transmission / reception unit 51 can also transmit and receive various data (or information) between the recording device 2, the utterance recognition device 3, and the utterance content management server 6, for example, via the communication network 100. In this embodiment, the transmission / reception unit 51 functions as an example of at least one of a transmitting means and a receiving means.

[0073] The operation reception unit 52 is mainly realized by processing of the CPU 501 on the keyboard 312 and the pointing device 313, and receives operations such as various inputs and settings on the display device 5. In this embodiment, the operation reception unit 52 functions as an example of an operation reception means.

[0074] The display control unit 54 is mainly realized by processing of the CPU 501 on the display 507, and controls the display of various screens and information (data) on the display device 5. The display control unit 54 also uses, for example, a browser to cause the display device 5 to display a display screen created in HTML or the like. The display control unit 54 also displays at least one of a specific speech content and a combination content combining the specific speech content and a facial image of a specific speaker on the display 507. In this embodiment, the display control unit 54 functions as an example of a display control means.

[0075] The generation unit 57 is mainly realized by the processing of the CPU 501, and generates screen data for displaying the text as the speech recognition result transmitted by the speech recognition device 3 and the facial images of each participant on the display 507. In this case, the display device 5 may communicate with the recording device 2 to store facial images (video) of all participants attending a predetermined event such as a conference in a predetermined area of ​​the storage unit 5000, for example, in association with the participant identification information. Then, the generation unit 57 may read the facial images (video) of the participants managed in the storage unit 5000 based on the text information as the speech recognition result transmitted by the speech recognition device 3 and the participant identification information corresponding to the text information, and generate screen data for displaying the images on the display 507. In this embodiment, the generation unit 57 functions as an example of a generation means.

[0076] The memory readout unit 59 is mainly realized by the processing of the CPU 301 on at least one of the ROM 302, the EEPROM 304, and the HD 305, and stores various data (or information) in the memory unit 3000 and reads various data (or information) from the memory unit 3000. In this embodiment, the memory readout unit 39 functions as an example of a memory readout means.

[0077] <Functional configuration of the speech content management server> As shown in Fig. 5, the utterance content management server 6 has a transmitting / receiving unit 61, an acquiring unit 63, a generating unit 67, and a storing / reading unit 69. Each of these functional units is a function or means realized when any of the hardware resources shown in Fig. 4 operates in response to an instruction from the CPU 601 in accordance with a program for the utterance content management server 6 that is loaded from at least one of the ROM 602, the EEPROM 6504, and the HD 605 to the RAM 603. The utterance content management server 6 also has a storage unit 6000 constructed by at least one of the ROM 602, the EEPROM 6504, and the HD 605 shown in Fig. 4. The storage unit 6000 stores a communication program (communication application) and the like for communicating with the recording device 2, the speech recognition device 3, and the display device 5 via the communication network 100.

[0078] <<Functional configuration of the speech content management server>> Next, each functional configuration of the utterance content management server 6 will be described in detail. The transmitting / receiving unit 61 of the utterance content management server 6 shown in Fig. 5 is mainly realized by processing of the CPU 601 on the network I / F 611 and the short-range communication I / F 608. The transmitting / receiving unit 61 can also transmit and receive various data (or information) between the recording device 2, the utterance recognition device 3, and the display device 5, for example, via the communication network 100. In this embodiment, the transmitting / receiving unit 61 functions as an example of at least one of a transmitting means and a receiving means.

[0079] The acquisition unit 63 is mainly realized by processing of the CPU 601, and acquires, for example, captured image (video) data and sound (audio) data transmitted by the speech recognition device 3 via the transmission / reception unit 61. In this embodiment, the acquisition unit 63 functions as an example of an acquisition means.

[0080] The generation unit 57 is mainly realized by processing of the CPU 501. The generation unit 57 generates, for example, minutes of a predetermined event such as a conference based on text information as the speech recognition result transmitted by the speech recognition device 3. In this embodiment, the generation unit 57 functions as an example of a generating means. Note that in the communication system according to this embodiment, the generation unit 67 may generate the above-mentioned screen data instead of the function of the generation unit 57 possessed by the display device 5. Alternatively, another device capable of communicating with the display device 5 and the utterance content management server 6 via the communication network 100 may have a function equivalent to the generation unit 57 or the generation unit 67.

[0081] The memory / read unit 79 is mainly realized by processing of the CPU 701 on at least one of the ROM 702, the EEPROM 704, and the HD 705, and stores various data (or information) in the memory unit 7000 and reads various data (or information) from the memory unit 7000. In this embodiment, the memory / read unit 79 functions as an example of a memory / read means.

[0082] [Processing or Operation of the Embodiment] Next, the processing or operation of the first embodiment will be described with reference to Fig. 9 to Fig. 11. Fig. 9 is a sequence diagram showing an example of the overall processing according to the first embodiment. First, the machine lip reading pre-training unit 359 included in the processing unit 35 of the speech recognition device 3 pre-trains the lip feature calculation model 356 (step S1). More specifically, the lip feature calculation unit 357 pre-trains the lip feature calculation model 356. Furthermore, the processing unit 35 performs the machine lip reading pre-training using the mouth shape pattern management DB 3001 (see Fig. 7).

[0083] Next, the imaging unit 23 of the recording device 2 captures an image (video) of the face of each participant participating in a predetermined event such as a conference. Furthermore, the sound input / output unit 24 collects the voices of each participant and surrounding sounds (step S11). Note that the imaging unit 23 may capture the image (video) of each participant's face approximately simultaneously, as with a spherical camera, or may capture the image (video) of each participant's face using multiple cameras.

[0084] Next, the transmitter / receiver 21 transmits image information related to the image captured in step S11, audio information related to the collected audio, and participant identification information of each participant to the utterance recognition device 3 (step S12). As a result, the transmitter / receiver 31 of the utterance recognition device 3 receives the image information and audio information transmitted by the recording device 2, as well as the participant identification information of each participant.

[0085] Next, the processing unit 35 of the speech recognition device 3 executes the speech recognition process (step S13). When executing the speech recognition process, the processing unit 35 trains the multimodal recognition model 372. Specifically, the first multimodal recognition unit 375 trains the multimodal recognition model 372 by fine-tuning including the lip feature calculation unit 357. Then, the storage and reading unit 39 registers and manages the results of the execution of the speech recognition process in the speech recognition result management DB 3002 (see FIG. 8).

[0086] Next, the transmitting / receiving unit 31 transmits the speech recognition result to the display device 5 (step S14). As a result, the transmitting / receiving unit 51 of the display device 5 receives the speech recognition result transmitted by the speech recognition device 3. At this time, the speech recognition result includes the participant identification information and the text information that has been recognized by the speech.

[0087] 16 based on the participant identification information and text information received in step S14, and the display control unit 54 displays the generated display screen as the speech recognition result on the display 507 (step S15). The display screen generated at this time may be realized by, for example, the method described in the description of the generation unit 57 above.

[0088] In the communication system according to this embodiment, for example, when the process of step S14 described above is executed, other devices may exist between the utterance recognition device 3 and the display device 5. In other words, each piece of information (data) transmitted and received between the utterance recognition device 3 and the display device 5 may be transmitted and received once via another device. The above-described configuration is applicable even when other processing steps exist between the utterance recognition device 3 and the display device 5. Furthermore, the utterance recognition result transmitted from the utterance recognition device 3 in step S14 may be transmitted to the utterance content management server 6 instead of the display device 5. In that case, the utterance recognition result displayed in step S15 is based on the content of the utterance recognition result transmitted from the utterance content management server 6 to the display device 5.

[0089] <Processing Overview of Multimodal Speaker Recognition (Pre-Combination)> Next, an overview of the processing of multimodal speaker recognition (pre-combining) will be described. Fig. 10 is a schematic diagram showing an example of processing during pre-combining in the multimodal speaker recognition system according to the first embodiment. First, the lip feature calculation unit 357 inputs a lip image sequence (moving images) showing one utterance to acquire lip features. At this time, the lip feature calculation unit 357 pre-trains the lip feature calculation model 356 so that the series of mouth shape patterns can be recognized by the mouth shape recognition unit 358 as correct labels.

[0090] The speech feature calculation unit 373 also receives the speech waveform input by the speech input unit 371. Next, based on the input speech waveform, the speech feature calculation unit 373 obtains speech features as logarithmic mel filter bank features using a Hamming window with a width of 25 ms and a shift of 11 ms.

[0091] Next, the feature integration unit 374 combines the lip features extracted as a result of pre-training by the lip feature calculation unit 357 with the audio features obtained by the audio feature calculation unit 373 to obtain multimodal features. In this case, for example, if the frame rate of the lip image sequence is 30 fps (≒ 33.3 ms), the frame rate of the input audio waveform is 25 ms wide / 11 ms. Therefore, combining this with three frames of audio features enables temporal alignment. As a result, if the lip features per frame are 384-dimensional and the audio features are 40-dimensional, the multimodal features per frame will be 384 + (40 × 3) = 504-dimensional. In Figure 10, one "●●●···●●●" represents a 384-dimensional lip feature, and one "○○" represents a 40-dimensional audio feature. Therefore, a single "●●●···●●●", "○○", "○○", "○○" represents a multimodal feature with 504 dimensions per frame.

[0092] Next, the first multimodal recognition unit 375 inputs the multimodal features obtained by the feature integration unit 374, and fine-tunes the lip feature calculation unit 357 and the multimodal recognition model 372 using the hiragana sequence as the correct label.

[0093] In machine lip reading, if a sequence of hiragana characters is used as the correct label, different correct answers will be given for the same input due to the same mouth-shaped allophones. This makes it difficult to train parameters for lip feature extraction. However, this problem is resolved if a sequence of mouth shape patterns is used as the correct label, and it is expected that effective feature extraction will be performed by the lip feature calculation unit. Then, multimodal recognition is performed on Japanese hiragana characters, which are the final recognition result. This makes it possible to train a multimodal recognition model that can recognize with higher accuracy than training a sequence of hiragana characters using an end-to-end configuration that also includes lip feature extraction. After this, it is also possible to use IME technology or the like to create a sentence that includes kanji characters. Note that while the above explanation has been based on Japanese, this embodiment can also be applied to foreign languages ​​for which mouth shape patterns are defined.

[0094] ●Screen display example● Next, a screen displayed on the display device 5 will be described. FIG. 11 is an example of a screen display on the display device showing a speech recognition result. As shown in FIG. 11, a speech recognition result screen 5101 is displayed on the display 507 of the display device 5 by the display control unit 54. The speech recognition result screen 5101 displays the content of utterances made by participants A, B, C, and D participating in a specific conference in association with images (video) of the speakers' faces. In this case, the displayed content may be scrolled over time. Displaying such speech recognition result screen 5101 enables each participant to visually recognize who made which remark in real time during an event such as a conference. Furthermore, each participant can view the content of utterances displayed on the speech recognition result screen 5101 as a simple minutes of the meeting, thereby improving understanding of the meeting and streamlining its progress.

[0095] [Major Effects of the First Embodiment] As described above, according to this embodiment, the speech recognition device 3 performs multimodal recognition processing in parallel using lip features and speech features obtained based on the speech sounds of one or more speakers, thereby recognizing a specific speaker from one or more speakers and recognizing specific speech content uttered by the specific speaker. This has the effect of enabling the recognition of a specific speaker and specific speech content in a single processing step in speech recognition processing using multimodal recognition.

[0096] Furthermore, according to this embodiment, there is no need to provide a separate speech period detection function for detecting speech periods, so it is expected that the complexity of the system can be reduced.

[0097] Furthermore, according to this embodiment, multimodal recognition is performed using the results of prior learning regarding machine lip reading, so it is expected that the efficiency of the multimodal recognition process can be improved.

[0098] Second Embodiment Next, a second embodiment will be described with reference to Figs. 11 to 16. In the second embodiment, a case will be described in which post-stage combining is performed instead of pre-stage combining in the processing unit 35 in the multimodal utterance recognition processing according to the first embodiment. Note that the basic parts of the system configuration, hardware configuration, and functional configuration for realizing the second embodiment are the same as those of the first embodiment, and therefore their description will be omitted. Below, details of the post-stage combining in the multimodal utterance recognition processing of the processing unit 35 of the speech recognition device 3 will be described. Note that post-stage combining is a technique for combining the mouth shape recognition result and the voice recognition result described above, and is also called late fusion.

[0099] <Functional configuration of multimodal speech recognition (post-combination)> Fig. 12 is a diagram showing an example of the functional configuration at the time of subsequent stage combination in multimodal speech recognition processing. Note that, among the functional configurations shown in Fig. 12, descriptions of the same functions as those in the functional configuration shown in Fig. 6 will be omitted. The functional configurations and processing order that are the same as those in the first embodiment are from the image input unit 351 to the lip pixel number conversion unit 355, as well as the speech input unit 371 and the speech feature calculation unit 373, and descriptions of these units will be omitted.

[0100] When combined in the latter stage of the multimodal speech recognition processing, the processing unit 35 has functions that are different from or added to the functions described in the first embodiment, such as a mouth shape recognition unit 358, a mouth shape recognition model 361, a voice recognition model 377, a voice recognition unit 379, a speech section estimation unit 381, and an utterance content recognition result correction unit 382. The speech section estimation unit 381 and the utterance content recognition result correction unit 382 constitute a second multimodal recognition unit 383.

[0101] <<Details of Multimodal Speech Recognition (Post-Combination) Function>> Next, among the detailed functions constituting the processing unit 35, functions including those specific to the second embodiment will be described.

[0102] Lip feature calculation unit 357 calculates lip features for the sequence of continuous lip images whose image size has been changed so that the speech content can be easily recognized by mouth shape recognition unit 358, which will be described later. At this time, lip feature calculation model 356 obtained in advance is trained.

[0103] The mouth shape recognition unit 358 recognizes the lip feature calculated by the lip feature calculation unit 357 as a series of mouth shape patterns, and outputs the recognition result to the second multimodal recognition unit 383.

[0104] In this embodiment, the utterance recognition device 3 converts Japanese into hiragana and then converts it into a mouth shape pattern corresponding to the hiragana. The correspondence between hiragana and mouth shape patterns is determined using the mouth shape pattern management DB 3001 (see FIG. 7) described above. The utterance recognition device 3 trains the mouth shape recognition model 361 using the converted mouth shape pattern as a correct label, and recognizes the content of the utterance using the mouth shape recognition model 361. In this embodiment, the lip feature calculation unit 357 and the mouth shape recognition unit 358 are assumed to have an end-to-end configuration in which a single neural network implements processes from lip feature extraction (optimization of convolution parameters) to recognition, but they may also be configured separately.

[0105] The audio input unit 371 inputs sound (audio) data acquired by the above-described acquisition unit 33. In this embodiment, for example, the input conditions are monaural uncompressed data sampled at 16 kHz and 16 bits.

[0106] The speech feature calculation unit 373 extracts logarithmic mel filter bank features from the speech waveform input by the speech input unit 371, using, for example, a Hamming window with a width of 25 ms and a shift of 11 ms. A speech recognition model 377 to be used in the speech recognition unit 379 is trained on this feature sequence. The correct answer labels used in this case are those in hiragana, not those from the mouth shape recognition unit.

[0107] The mouth shape recognition unit 358 and the voice recognition unit 379 output the mouth shape recognition result and the voice recognition result, which are the respective recognition results, to the speech section estimation unit 381 .

[0108] The speech section estimation unit 381 recognizes the content of a specific utterance spoken by a specific speaker by extracting the speech section actually spoken by the specific speaker, based on the section of the lip image sequence input from the mouth shape recognition unit 358 and the speech input from the speech recognition unit 379.

[0109] The speech content recognition result correction unit 382 corrects the speech recognition result that has been erroneously output by the speech section estimation unit 381. The speech section estimation unit 381 and the speech content recognition result correction unit 382 are collectively referred to as a second multimodal recognition unit 383, which performs recognition taking into consideration both image features and speech features.

[0110] The speech content recognition result output unit 376 transmits the recognition result output from the second multimodal recognition unit 383 to the display device 5 as the final recognition result of the speech content, or transmits the recognition result to the speech content management server 6, thereby visually displaying the recognition result or saving it as a text file.

[0111] Next, the processing or operation of the second embodiment will be described. Fig. 13 is a sequence diagram showing an example of the overall processing according to the second embodiment. Here, the processing of steps S21 and S22 is the same as steps S11 and S12 shown in Fig. 9, and therefore description thereof will be omitted.

[0112] Next, the processing unit 35 of the speech recognition device 3 executes the speech recognition process using the mouth shape recognition model and the voice recognition model (step S23).

[0113] <Details of speech recognition processing> Next, the details of the speech recognition process will be described. Fig. 14 is an overall flowchart showing an example of multimodal processing according to the second embodiment. First, the speech input unit 371 acquires speech information (step S23-1).

[0114] Next, the speech feature calculation unit 373 calculates speech features from the speech information acquired in step S23-1 (step S23-2).

[0115] Next, the speech recognition unit 379 recognizes the speech from the speech feature amount calculated by the speech feature amount calculation unit 373 (step S23-3).

[0116] Next, the processing unit 35 repeats the following processes from step S23-4 to step S23-11 for all participants attending the event such as a conference. That is, the processing unit 35 executes the processes from step S23-4 to step S23-11 for each participant (step S23-4).

[0117] First, the face area recognition unit 353 recognizes the face area of ​​one participant displayed in the input image (video) acquired by the recording device 2 in the same section as the acquired voice (step S23-5).

[0118] Next, the lip area extraction unit 354 extracts the lip area from the face area recognized by the face area recognition unit 353 (step S23-6).

[0119] Next, the lip feature amount calculation unit 357 calculates the lip feature amount based on the number of lip pixels converted from the lip area extracted by the lip area extraction unit 354 (step S23-7).

[0120] Next, the mouth shape recognition unit 358 learns the mouth shape recognition model 361 and recognizes the mouth shape (step S23-8).

[0121] Next, the second multimodal recognition unit 383 associates the mouth shape recognition result obtained in step S23-8 with the speech recognition result obtained by the speech recognition unit 379 (step S23-9).

[0122] Next, the second multimodal recognition unit 383 selects one utterance as a multimodal recognition result from the mouth shape recognition results of the multiple participants. Specifically, the second multimodal recognition unit 383 selects one utterance as a multimodal recognition result using the speech recognition result and the mouth shape recognition result, using the mouth shape of the speaker that is most closely related to the speech recognition result, that is, the speaker that is most likely to actually be speaking the corresponding speech (step S23-10).

[0123] Next, in order to select the mouth shape with the highest degree of association, the loop processing from step S23-4 to step S23-11 described above is repeated a number of times corresponding to the number of participants (step S23-11). Note that the specific processing flow of step S23-9 and step S23-10 by the second multimodal recognition unit 383, surrounded by a dashed line, will be further explained using the flowchart in Fig. 15, which will be described later.

[0124] Next, the utterance content recognition result output unit 376 outputs the utterance content recognition result recognized by the second multimodal recognition unit 383 to the outside, and then the flow ends (step S23-12).

[0125] <<Output processing of multimodal recognition results>> Next, the output process of the multimodal recognition result will be described. Fig. 15 is a flowchart showing the output process of the multimodal recognition result according to the second embodiment. First, the second multimodal recognition unit 383 calculates the Levenshtein distance between one person's mouth shape recognition result and the speech recognition result (step S200-1). That is, the processing unit 35 including the second multimodal recognition unit 383 executes the process of step S200-1 for the number of participants when extracting an utterance section, and therefore calculates the Levenshtein distance between the mouth shape recognition result that recognizes each mouth shape of one or more speakers and the speech recognition result that recognizes each speech of one or more speakers. Here, a sequence of mouth shape patterns representing the mouth shape recognition results for one person is taken as the correct answer, and the Levenshtein distance between this and the sequence of hiragana characters resulting from speech recognition is calculated. In general speech recognition, there is a method for calculating the operation cost, i.e., how many characters (words) need to be "deleted," "inserted," or "replaced" from the recognized and output text to match the correct label. Using this method, the speech recognition device 3 expresses the difference from the correct label as a character error rate (CER), word error rate (WER), etc., and evaluates the accuracy. In this case, the smaller the character error rate and word error rate (the closer the distance), the smaller the difference from the correct label, i.e., the higher the recognition accuracy.

[0126] In this embodiment, a sequence of mouth shape patterns is considered correct, but the corresponding speech recognition result is a sequence of hiragana characters, so there is no matching portion. However, the speech recognition device 3 determines a match when a hiragana character corresponding to the mouth shape pattern that is the recognition result is recognized using the mouth shape pattern management DB 3001 (see FIG. 7) corresponding to Japanese hiragana. In other cases, the speech recognition device 3 calculates the Levenshtein distance by performing the operations of "delete," "insert," and "replace," similar to the general method of calculating the Levenshtein distance. For example, in the flowchart of FIG. 15, the second multimodal recognition unit 383 calculates the Levenshtein distance between the mouth shape recognition results for the number of participants and one speech recognition result, one by one.

[0127] Next, the second multimodal recognition unit 383 determines whether the Levenshtein distance is the minimum (step S200-2). If the Levenshtein distance is the minimum (step S200-2: YES), the second multimodal recognition unit 383 adopts the recognized mouth shape as the mouth shape of the speaker (step S200-3).

[0128] Next, the second multimodal recognition unit 383 estimates the actual speech section from each of the adopted sequences of the mouth shape recognition result and the speech recognition result (step S200-4).

[0129] Next, the second multimodal recognition unit 383 deletes the insertion error from the speech recognition result, outputs it as a final multimodal recognition result, and exits this flow (step S200-5).

[0130] On the other hand, if the Levenshtein distance is not the minimum (step S200-3: NO), the second multimodal recognition unit 383 determines that the recognized mouth shape and the speech are not synchronized, and discards (rejects) the recognition result as not being the mouth shape of the speaker who uttered the corresponding speech, and exits this flow (step S200-6). Thus, if the calculated Levenshtein distance is the minimum, the multimodal recognition unit 383 adopts the mouth shape used in the calculation as that of a specific speaker, and if the Levenshtein distance is not the minimum, discards the mouth shape used in the calculation as not being that of the specific speaker. Note that in this embodiment, an "INS" process (a process in which an unintended character is inserted) is determined to be an "insertion error," while a "DEL" process and a "SUB" process are excluded from the "insertion error" determination. However, the above-described targets are not limited to these targets, and a specification may also be adopted in which a "SUB" process is also determined to be an "insertion error."

[0131] Although the Levenshtein distance is used in this embodiment, the distance measure for calculating the edit cost by comparing character strings is not limited to this.

[0132] Returning to FIG. 13, the processes in steps S24 and S25 are similar to the processes in steps S14 and S15 described above, and therefore will not be described here.

[0133] <Overview of multimodal speaker recognition (post-combination) processing> Next, an overview of the processing of multimodal speaker recognition (post-combining) will be described. Fig. 16 is a schematic diagram showing an example of processing during post-combining in the multimodal speaker recognition system according to the second embodiment. First, the lip feature calculation unit 357 inputs a lip image sequence (moving image) showing one utterance and calculates lip features.

[0134] Next, the mouth shape recognition unit 358 recognizes the mouth shape recognition result from the mouth shape recognition model and lip feature amount.

[0135] On the other hand, the speech feature calculation unit 373 receives the speech waveform input by the speech input unit 371, and calculates speech features as logarithmic Mel filter bank features based on the received speech waveform using a Hamming window with a width of 25 ms and a shift of 11 ms.

[0136] Next, the speech recognition unit 379 recognizes the speech recognition result based on the speech recognition model and speech feature amount.

[0137] Next, the second multimodal recognition unit 383 receives the mouth shape recognition result and the speech recognition result in parallel, performs multimodal recognition, and outputs the obtained multimodal features to the speech content recognition result output unit 376.

[0138] The recognition process executed by the second multimodal recognition unit 383 may be the same as that executed by the first multimodal recognition unit 375. Therefore, a description of the process for acquiring multimodal features per frame and the concept of synchronization with the frame rate will be omitted.

[0139] In this way, lip feature learning can be performed efficiently by pre-learning using the mouth shape patterns managed in the mouth shape pattern management DB 3001 shown in Fig. 7. As a result, it is expected that the operation costs, editing costs, etc., incurred by system designers can be reduced.

[0140] [Major Effects of the Second Embodiment] As described above, according to this embodiment, the speech recognition device 3 estimates speech sections and corrects speech content based on lip features and speech sounds based on the content of an utterance spoken by a speaker. This provides the effect of achieving high recognition accuracy in multimodal recognition in addition to the effect of the first embodiment.

[0141] Furthermore, according to this embodiment, by pre-learning using mouth shape patterns, it is possible to efficiently learn lip features, and it is also expected to have the effect of reducing operation costs, editing costs, etc., incurred by system designers.

[0142] [Other application examples of the embodiment] Another application example of the speech recognition according to the above-described embodiment is, for example, a mobile control system having a speech recognition device included in a mobile object such as an automobile (hereinafter also referred to as a vehicle), in which speech recognition processing is performed using multimodal recognition on content spoken by one or more speakers (passengers) while driving the vehicle and operating various devices. For example, in a usage scenario in which one or more passengers are in a vehicle equipped with an autonomous driving system, when a destination is input into a car navigation system by spoken voice, speeches simultaneously spoken by one or more speakers can be input accurately by speech recognition processing using multimodal recognition according to the present embodiment. Note that other application examples of the embodiments can also be applied to vehicles not equipped with an autonomous driving system.

[0143] [Overall configuration of the mobile control system] <System configuration example> FIG. 17 is a diagram illustrating an example of the overall configuration of a mobile object control system. As illustrated in FIG. 17, the mobile object control system 11 includes a recording device 12, an utterance recognition device 13, a display device 15, and an utterance content management server 16, with each device and server connected to each other via a communication network 110. However, the utterance content management server 16 does not necessarily need to be included in the mobile object control system 11. The mobile object control system 11 also includes an utterance recognition system 14 configured with the recording device 12 and the utterance recognition device 13. Here, the utterance recognition device 13 is included, for example, in a general car navigation system installed in a vehicle. The display device 15 may be, for example, a PC owned by one or more speakers and installed in each speaker's home, and the mobile object control system 11 is provided with a mechanism for recording voice recordings on the display device 15 in synchronization with the movement of the mobile object. In this case, if the mobile object control system 11 includes the utterance content management server 16, the voice recordings may be stored and managed in the utterance content management server 16.

[0144] The recording device 12, speech recognition device 13, display device 15, and speech content management server 16 constituting the mobile object control system 11 have the same hardware configuration as the recording device 2, speech recognition device 3, display device 5, and speech content management server 6 described in the first embodiment. Furthermore, the contents of each functional configuration are also the same as those of each functional configuration described in the first embodiment, so detailed description will be omitted.

[0145] <Example of control using multimodal recognition> The speech recognition system 14 performs the following control when, for example, a usage scenario is assumed in which a destination or the like is input based on spoken voices in a mobile vehicle such as an automobile equipped with a car navigation system including the speech recognition device 13. That is, the car navigation system performs multimodal recognition processing in parallel using lip features and voice features obtained based on the respective speeches uttered by one or more speakers (passengers), recognizes a specific speaker from the one or more speakers, and transmits the results of recognizing the specific speech content uttered by the specific speaker to the display device 15. In this case, for example, by applying the multimodal speaker recognition (pre-combination) processing described in the first embodiment and pre-training a lip feature calculation model for the driver, the subsequent multimodal recognition processing can be performed accurately. Furthermore, by applying the above-described embodiment, the display device 15 can also display the speech recognition results (conversations, etc.) of other passengers, including the driver, in chronological order during a drive, trip, etc., in association with each person's facial photograph. At this time, a process may be performed in which a background image of the place where the vehicle was traveling at that time is displayed on the background of the display device 15.

[0146] By assuming such usage scenarios, for example, when driving or traveling, it becomes possible to recognize the speech of a specific speaker without being concerned about the speech of others emitted by passengers' conversations, music, video playback, etc. Therefore, even when inputting a destination, etc., it becomes possible to enjoy driving, traveling, etc. without disrupting the surrounding atmosphere.

[0147] [Supplementary explanation of the embodiment] Each function of the above-described embodiments can be realized by one or more processing circuits. Here, the term "processing circuit" includes a device programmed by software to perform each function, such as a processor implemented by electronic circuits. Examples of such devices include a processor, an application-specific integrated circuit (ASIC), a digital signal processor (DSP), a field programmable gate array (FPGA), a system on a chip (SOC), a graphics processing unit (GPU), and conventional circuit modules designed to perform each of the above-described functions.

[0148] The various pieces of information obtained by the above-described embodiments may be acquired through the learning effects of machine learning using artificial intelligence (AI). In this case, for example, the speech recognition device 3 may use machine learning to create minutes or the like based on text obtained by multimodal recognition processing. Furthermore, a device, database, or the like other than the speech recognition device 3 may acquire and process various pieces of information obtained through machine learning. Here, machine learning refers to a technology that allows a computer to acquire human-like learning capabilities. The computer autonomously creates algorithms necessary for judgments such as data classification from previously acquired training data, and applies these algorithms to new data to make predictions. The learning method for machine learning may be any of supervised learning, unsupervised learning, semi-supervised learning, reinforcement learning, and deep learning. Furthermore, the learning method for machine learning may be a combination of these learning methods, or any learning method for machine learning.

[0149] So far, we have described a speech recognition system, a communication system, a speech recognition device, a mobile object control system, a speech recognition method, and a program according to one embodiment of the present invention. However, the present invention is not limited to the above-described embodiment, and other modifications, such as additions, changes, or deletions, can be made within the scope that can be conceived by a person skilled in the art. Any modification that achieves the functions and effects of the present invention is included in the scope of the present invention. [Explanation of symbols]

[0150] 1. Communication Systems 2. Data Acquisition Device 3. Speech recognition device 4. Speech Recognition System 5 Display device 11. Mobile Control Systems 12 Data Acquisition Device 13 Speech recognition device 14 Speech Recognition System 15 Display device 31 Transmitting / receiving unit (an example of a transmitting means, an example of a receiving means) 35 Processing unit (an example of processing means) 54 Display control unit (an example of a display control means) 57 Generation unit (an example of generation means) 507 Display (an example of a display means) [Prior art documents] [Patent documents]

[0151] [Patent Document 1] Japanese Patent Application Laid-Open No. 2011-059186

Claims

1. A speech recognition system including: a recording device that records images and audio accompanying utterances by one or more speakers; and a speech recognition device that receives image information relating to the images and audio information relating to the audio transmitted by the recording device and recognizes the content of the utterances, The speech recognition device a processing means for performing multimodal recognition processing in parallel using lip features obtained based on lip images that change when the one or more speakers speak and speech features obtained based on speeches uttered by the one or more speakers, thereby recognizing a specific speaker from the one or more speakers and recognizing specific speech content uttered by the specific speaker; a transmitting means for transmitting the specific speech content identified by the processing means to a display device; and The processing means Recognizing a mouth shape based on the lip feature amount to obtain a mouth shape recognition result; Recognizing a speech based on the speech feature to obtain a speech recognition result; adopting a speech recognition result corresponding to the mouth shape recognition result of the specific speaker from among the speech recognition results of the one or more speakers; using the mouth shape recognition result of the specific speaker to delete insertion errors contained in the adopted speech recognition result, thereby recognizing the content of the specific utterance; A speech recognition system comprising:

2. The processing means Recognizing the specific utterance content uttered by the specific speaker by extracting a speech section in which the specific speaker actually spoke.

2. The speech recognition system of claim 1.

3. The processing means obtaining multimodal features for the multimodal recognition processing by performing temporal alignment in accordance with a ratio of the lip features and the speech features per frame rate of a lip image sequence representing one utterance among the contents of utterances made by the one or more speakers; 3. The speech recognition system according to claim 1 or 2.

4. The processing means When extracting the speech section, a Levenshtein distance is calculated between a mouth shape recognition result that recognizes each mouth shape of the one or more speakers and a speech recognition result that recognizes each voice of the one or more speakers.

3. The speech recognition system of claim 2.

5. The processing means If the calculated Levenshtein distance is the smallest, the mouth shape used in the calculation is adopted as that of the specific speaker, and if the Levenshtein distance is not the smallest, the mouth shape used in the calculation is discarded as not that of the specific speaker.

5. The speech recognition system of claim 4.

6. 6. A communication system including the speech recognition system according to claim 1 and a display device that displays a predetermined screen based on screen information transmitted by the speech recognition system, The display device includes: a display control means for displaying on a display means at least one of the specific speech content and a combination content of the specific speech content and a face image of the specific speaker; A communication system comprising:

7. A speech recognition device that recognizes predetermined speech content based on image information relating to the images and audio information relating to the audio transmitted from a recording device that records images and audio accompanying utterances by one or more speakers, a processing means for performing multimodal recognition processing in parallel using lip features obtained based on lip images that change when the one or more speakers speak and speech features obtained based on speeches uttered by the one or more speakers, thereby recognizing a specific speaker from the one or more speakers and recognizing specific speech content uttered by the specific speaker; a transmitting means for transmitting the specific speech content identified by the processing means to a display device; and The processing means Recognizing a mouth shape based on the lip feature amount to obtain a mouth shape recognition result; Recognizing a speech based on the speech feature to obtain a speech recognition result; adopting a speech recognition result corresponding to the mouth shape recognition result of the specific speaker from among the speech recognition results of the one or more speakers; using the mouth shape recognition result of the specific speaker to delete insertion errors contained in the adopted speech recognition result, thereby recognizing the content of the specific utterance; A speech recognition device characterized by:

8. 6. A mobile object control system for controlling a mobile object, comprising: the speech recognition system according to claim 1; and a display device that displays a predetermined screen based on screen information transmitted by the speech recognition system, The display device includes: a display control means for displaying, on a display means, at least one of a speech content for controlling the moving body as the specific speech content and a combination content combining the speech content for controlling the moving body and a face image of the specific speaker; A mobile object control system.

9. A speech recognition method executed by a speech recognition device that recognizes predetermined speech content based on image information related to the images and audio information related to the audio transmitted by a recording device that records images and audio accompanying utterances by one or more speakers, comprising: a processing step of performing multimodal recognition processing in parallel using lip features obtained based on lip images that change when the one or more speakers speak and speech features obtained based on speeches uttered by the one or more speakers, thereby recognizing a specific speaker from the one or more speakers and recognizing specific speech content uttered by the specific speaker; a transmitting step of transmitting the specific utterance content identified by the processing step to a display device; Perform a process including The processing steps include: Recognizing a mouth shape based on the lip feature amount to obtain a mouth shape recognition result; Recognizing a speech based on the speech feature to obtain a speech recognition result; adopting a speech recognition result corresponding to the mouth shape recognition result of the specific speaker from among the speech recognition results of the one or more speakers; using the mouth shape recognition result of the specific speaker to delete insertion errors contained in the adopted speech recognition result, thereby recognizing the content of the specific utterance; A speech recognition method comprising:

10. an utterance recognition device that recognizes predetermined utterance content based on image information related to the images and audio information related to the audio transmitted from a recording device that records images and audio accompanying utterances by one or more speakers; a processing step of performing multimodal recognition processing in parallel using lip features obtained based on lip images that change when the one or more speakers speak and speech features obtained based on speeches uttered by the one or more speakers, thereby recognizing a specific speaker from the one or more speakers and recognizing specific speech content uttered by the specific speaker; a transmitting step of transmitting the specific utterance content identified by the processing step to a display device; Execute a process including The processing steps include: Recognizing a mouth shape based on the lip feature amount to obtain a mouth shape recognition result; Recognizing a speech based on the speech feature to obtain a speech recognition result; adopting a speech recognition result corresponding to the mouth shape recognition result of the specific speaker from among the speech recognition results of the one or more speakers; using the mouth shape recognition result of the specific speaker to delete insertion errors contained in the adopted speech recognition result, thereby recognizing the content of the specific utterance; program.

Citation Information

Patent Citations

  • Speech section detecting device and speech recognition device, program and recording medium

    JP2011059186A

  • Identity authentication method and device

    JP2019522840A

  • Voice recognition system and voice recognition method

    JP2020086048A