Voice recognition device, voice recognition method, and recording medium

By providing a acquisition unit, a storage unit, a speech start detection unit and a speaker distinction unit in the voice recognition device, the problem that the first speaker cannot recognize the voice of the second speaker is solved, and accurate speech recognition of the second speaker is realized, recognition efficiency is improved and power and communication burden is reduced.

CN111755000BActive Publication Date: 2025-08-08PANASONIC HOLDINGS CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010230365.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-10-30
Filing Date
2020-03-23
Publication Date
2025-08-08
Estimated Expiration
2040-03-23

AI Technical Summary

Technical Problem

In the prior art, the first speaker cannot effectively recognize the voice of the second speaker because the second speaker does not understand the use method of the voice recognition device, resulting in the inability to accurately perform speech recognition.

Method used

By setting a acquisition unit, a storage unit, a speech start detection unit and a speaker distinction unit in the voice recognition device, the voices of the first speaker and the second speaker are respectively acquired, stored and distinguished by the speech start detection unit, and the speech start detection unit detects the speech start position, and distinguishes whether the operation inputer is the first speaker or the second speaker through the speaker distinction unit, thereby performing speech recognition.

Benefits of technology

Accurate recognition of the voice of the second speaker is achieved, interference with the first speaker is reduced, accuracy and efficiency of speech recognition is improved, and power consumption and communication volume are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111755000B_ABST
    Figure CN111755000B_ABST
Patent Text Reader

Abstract

A speech recognition device, speech recognition method, and recording medium. The speech recognition device comprises: an acquisition unit that acquires individual speech sounds of a conversation between a first speaker and one or more second speakers; a storage unit that stores the individual speech sounds of the conversation between the first speaker and one or more second speakers; an input unit that receives an operation input; an utterance start detection unit that detects the start position of an utterance for each speech sound in response to the operation input to the input unit; and a speaker distinguishing unit that distinguishes between the first speaker who has made an operation input and one or more second speakers who have not made an operation input, based on a first time point set for each speech sound at which the operation input to the input unit is received and a second time point indicating the start position of the utterance detected by the utterance start detection unit for each speech sound. The distinguished speech sounds of the first speaker and one or more second speakers are then submitted to the speech recognition unit for speech recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a speech recognition device, a speech recognition method, and a recording medium. Background Art

[0002] For example, Patent Document 1 discloses a speech recognition device comprising: a speech timing indication acquisition mechanism for acquiring a user's indication of speech timing; a speech signal holding mechanism for holding an input speech signal and, when the speech timing indication acquisition mechanism acquires an indication of speech start, outputting the held speech signal and a subsequently input speech signal; a speech interval detection mechanism for detecting a speech interval based on the speech signal output by the speech signal holding mechanism; and an erroneous operation detection mechanism for comparing the moment information of the speech interval with the presence or absence of an indication of speech timing and the moment information, and detecting an erroneous operation by the user.

[0003] This voice recognition device detects a user's erroneous operation and can notify the user of the detected erroneous operation.

[0004] Prior art literature

[0005] Patent Literature

[0006] Patent Document 1: Japanese Patent No. 5375423 Summary of the Invention

[0007] Problems to be solved by the invention

[0008] However, in the technology disclosed in Patent Document 1, for example, if the first speaker is the owner of a speech recognition device, the first speaker, knowing how to use the speech recognition device, can correctly operate the device to have it recognize their speech. Therefore, the first speaker can have the speech recognition device recognize their speech from the beginning to the end. However, the second speaker, the first speaker's conversation partner, does not know how to use the speech recognition device, and the first speaker cannot recognize the timing of the second speaker's speech. Therefore, it is difficult for the first speaker to have the speech recognition device recognize the second speaker's speech from the beginning to the end. As a result, the second speaker's speech cannot be fully recognized, and the first speaker needs to prompt the second speaker to speak again.

[0009] Therefore, the present disclosure has been made in view of the above circumstances, and an object thereof is to provide a speech recognition device, a speech recognition method, and a recording medium capable of reliably acquiring the speech of a conversation partner and thereby performing speech recognition on the speech of the conversation partner.

[0010] Means for solving problems

[0011] A speech recognition device according to one embodiment of the present disclosure is a speech recognition device for a first speaker to have a conversation with one or more second speakers who are conversation partners of the first speaker, and comprises: an acquisition unit for acquiring individual speech sounds of a conversation between the first speaker and the one or more second speakers; a storage unit for storing the individual speech sounds of the conversation between the first speaker and the one or more second speakers acquired by the acquisition unit; an input unit for receiving at least an operation input from the first speaker; and a speech start detection unit for detecting the start of speech for each speech sound stored in the storage unit in response to the operation input to the input unit. a speaker distinguishing unit that distinguishes, from among the first speaker and the one or more second speakers, whether it is the first speaker who has made an operation input to the input unit or the one or more second speakers who has not made an operation input to the input unit, based on a first time point set for each voice at which an operation input to the input unit is accepted and a second time point indicating the start position of the speech detected by the speech start detection unit based on the respective voices, and the voices of the first speaker and the one or more second speakers that have been distinguished by the speaker distinguishing unit and subsequent to the start position in the respective voices are provided to the voice recognition unit for voice recognition.

[0012] Furthermore, some of these specific aspects may be implemented using systems, methods, integrated circuits, computer programs, or computer-readable recording media such as CD-ROMs, or any combination of these.

[0013] Effects of the Invention

[0014] According to the speech recognition device and the like of the present disclosure, it is possible to perform speech recognition on the speech of the conversation partner by reliably acquiring the speech of the conversation partner. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1A This figure shows an example of the appearance of a speech translation device equipped with the speech recognition device according to the first embodiment, and usage scenes of the speech translation device by a first speaker and a second speaker.

[0016] Figure 1B This is a diagram showing an example of the appearance of another speech translation device in Embodiment 1.

[0017] Figure 2 This is a block diagram showing the speech translation device in the first embodiment.

[0018] Figure 3 This is a flowchart showing the operation of the speech translation device when the first speaker speaks.

[0019] Figure 4 This is a diagram illustrating a time sequence between a first time point and a second time point during a conversation between a first speaker and a second speaker.

[0020] Figure 5 This is a flowchart showing the operation of the speech translation device when the second speaker speaks.

[0021] Figure 6 This is a flowchart showing the operation of the speaker distinguishing unit of the speech translation device in Embodiment 1.

[0022] Figure 7 This is a block diagram showing a speech translation device in Embodiment 2.

[0023] Description of reference numerals:

[0024] 10.10a Voice recognition device

[0025] 21 Acquisition Department

[0026] 22 Storage

[0027] 23 Speech start detection unit

[0028] 24 Input section

[0029] 25 Speaker Distinction

[0030] 26, 51 Speech Recognition Department

[0031] 29 Ministry of Communications. DETAILED DESCRIPTION

[0032] A speech recognition device according to one embodiment of the present disclosure is a speech recognition device for a first speaker to have a conversation with one or more second speakers who are conversation partners of the first speaker, and comprises: an acquisition unit for acquiring individual speech sounds of a conversation between the first speaker and the one or more second speakers; a storage unit for storing the individual speech sounds of the conversation between the first speaker and the one or more second speakers acquired by the acquisition unit; an input unit for receiving at least an operation input from the first speaker; and a speech start detection unit for detecting the start of speech for each speech sound stored in the storage unit in response to the operation input to the input unit. a speaker distinguishing unit that distinguishes, from among the first speaker and the one or more second speakers, whether the first speaker has made an operation input to the input unit or the one or more second speakers has not made an operation input to the input unit, based on a first time point set for each voice at which an operation input to the input unit is accepted and a second time point indicating the start position of the speech detected by the speech start detection unit based on the respective voices; and the voices of the first speaker and the one or more second speakers that have undergone the distinguishing processing by the speaker distinguishing unit and that have been processed after the start position are provided to the voice recognition unit for voice recognition.

[0033] Therefore, according to the present disclosure, each voice of a conversation between a first speaker and one or more second speakers is stored in a storage unit, so that it is possible to distinguish whether it is the first speaker or the second speaker based on the stored voice. As a result, the voice recognition unit can read the distinguished voices of the first speaker and the second speaker from the storage unit and perform voice recognition. In other words, if the first speaker speaks after the first speaker has made an operation input to the input unit, the voice recognition unit can perform voice recognition on the voice spoken by the first speaker. In addition, the second speaker usually starts speaking after the first speaker finishes speaking, so by the speaker making an operation input to the input unit corresponding to the second speaker's speech, the voice recognition unit can perform voice recognition on the voice spoken by the second speaker.

[0034] Therefore, in this speech recognition device, it is possible to perform speech recognition on the speech of the conversation partner by reliably acquiring the speech of the conversation partner.

[0035] In addition, another embodiment of the present disclosure relates to a speech recognition method for a first speaker to have a conversation with one or more second speakers who are the conversation partners of the first speaker, comprising: obtaining each speech of the conversation between the first speaker and the one or more second speakers; storing the obtained each speech of the conversation between the first speaker and the one or more second speakers in a storage unit; receiving at least an operation input from the first speaker to an input unit; and detecting and starting the conversation according to each speech stored in the storage unit in accordance with the operation input to the input unit. The invention relates to a method for detecting a start position of an utterance of a first speaker; a method for detecting a start position of an utterance of the ...

[0036] This voice recognition method also has the same operational effects as those of the above-mentioned voice recognition device.

[0037] Furthermore, a recording medium according to another aspect of the present disclosure is a computer-readable nonvolatile recording medium having recorded thereon a program for causing a computer to execute a speech recognition method.

[0038] This recording medium also has the same operational effects as the above-mentioned voice recognition device.

[0039] In addition, in the speech recognition device involved in another embodiment of the present disclosure, the speaker distinguishing unit compares the first time and the second time set for each of the voices in the conversation between the first speaker and the one or more second speakers, and distinguishes the first speaker from the first speaker and the one or more second speakers when the first time is earlier than the second time, and distinguishes one or more second speakers from the first speaker and the one or more second speakers when the second time is earlier than the first time.

[0040] Thus, for example, if the first speaker is the owner of a speech recognition device, the first speaker understands how to use the device and therefore begins speaking after inputting an operation to the input unit. That is, the first moment at which the first speaker's operation input to the input unit is accepted is earlier than the second moment at which the first speaker begins speaking, so the speaker distinguishing unit can distinguish the first speaker from the first speaker and one or more second speakers. Furthermore, the first speaker cannot recognize the timing of the second speaker's speech and therefore inputs an operation to the input unit after the second speaker begins speaking. That is, the first moment at which the first speaker's operation input to the input unit is accepted is later than the second moment at which the second speaker begins speaking, so the speaker distinguishing unit can distinguish the second speaker from the first speaker and one or more second speakers.

[0041] In this way, the speaker distinguishing unit can accurately distinguish whether the speaker who spoke most recently at the first time is the first speaker or the second speaker. Therefore, in this speech recognition device, the speech of the second speaker can be more reliably acquired, and speech recognition can be performed on the speech of the second speaker.

[0042] In addition, in the speech recognition device involved in other aspects of the present disclosure, when the first speaker is distinguished from the first speaker and the one or more second speakers, the speech recognition unit performs speech recognition on the speech uttered by the first speaker, and when the second speaker is distinguished from the first speaker and the one or more second speakers, the speech recognition unit performs speech recognition on the speech uttered by the second speaker.

[0043] Thus, the speaker distinguishing unit distinguishes whether the speaker who spoke is the first speaker or the second speaker, and the speech recognition unit can more reliably perform speech recognition on each of the speech uttered by the first speaker and the second speaker.

[0044] In addition, in the speech recognition device involved in another embodiment of the present disclosure, the speaker distinguishing unit distinguishes whether the speech is from the first speaker or the one or more second speakers based on the respective speech of the conversation between the first speaker and the one or more second speakers during a specified period before and after the first moment when the input unit accepted the operation input.

[0045] Thus, in order to distinguish between the first and second speakers, a predetermined period based on the first moment can be set. Therefore, it is possible to distinguish whether the most recent speech uttered by the speaker is from the first moment when the first speaker inputs an operation to a moment earlier than the first moment by the predetermined period, or from the first moment to a moment after the predetermined period. This allows for separate recognition of the speech of the first and second speakers. Therefore, this speech recognition device can accurately distinguish between the first and second speakers.

[0046] In addition, in the speech recognition device involved in another embodiment of the present disclosure, after speech recognition is performed on the speech uttered by the first speaker who has performed an operation input to the input unit, the storage unit starts storing the respective speech acquired by the acquisition unit in order to store the speech of the one or more second speakers.

[0047] Typically, the second speaker begins speaking after the first speaker finishes speaking and understands the content of the first speaker's speech. After the first speaker's speech is recognized, recording begins before the second speaker begins speaking. This allows the storage unit to reliably store the second speaker's speech. Furthermore, the speech recognition device can suspend speech storage at least from the moment the first speaker finishes speaking until the storage unit begins recording. This reduces power consumption by the speech recognition device, which is used to store the speech in the storage unit.

[0048] In addition, the speech recognition device involved in other embodiments of the present disclosure has a communication unit that can communicate with a cloud server having the speech recognition unit, and the communication unit sends the respective speech of the first speaker and the one or more second speakers that have been distinguished by the speaker distinguishing unit to the cloud server, and receives the results obtained by the speech recognition unit of the cloud server for the speech after the starting position of each speech.

[0049] In this way, the cloud server performs speech recognition on each of the speech uttered by the first speaker and one or more second speakers, thereby reducing the processing load of the speech recognition device.

[0050] Furthermore, a speech recognition device according to another aspect of the present disclosure includes the speech recognition unit that performs speech recognition on the speech from the start position onward in the speech of the first speaker and the one or more second speakers that have been subjected to the discrimination process by the speaker discrimination unit.

[0051] As a result, the voice recognition device performs voice recognition, and therefore does not need to send the voice to an external cloud server, thereby suppressing an increase in the amount of communication between the voice recognition device and the cloud server.

[0052] In addition, in a voice recognition device according to another aspect of the present disclosure, the input unit is one operation button provided on the voice recognition device.

[0053] This allows the first speaker to easily operate the speech recognition device.

[0054] In the speech recognition device according to another aspect of the present disclosure, the input unit receives an operation input from the first speaker each time the first speaker and the one or more second speakers speak.

[0055] Thus, by avoiding requesting operation input from the second speaker as much as possible and allowing the first speaker to actively input operation to the speech recognition device, one speaker can be reliably distinguished from the first speaker and the second speaker.

[0056] Furthermore, some of these specific aspects may be implemented using systems, methods, integrated circuits, computer programs, or computer-readable recording media such as CD-ROMs, or any combination of these.

[0057] The embodiments described below each represent a specific example of the present disclosure. The numerical values, shapes, materials, components, configuration positions and connection methods of the components, steps, and the order of the steps shown in the following embodiments are examples and are not intended to limit the present disclosure. In addition, among the components in the following embodiments, components that are not described in the independent claims are described as arbitrary components. In addition, in all embodiments, the respective contents can also be combined.

[0058] Hereinafter, a speech recognition device, a speech recognition method, and a recording medium according to one embodiment of the present disclosure will be described in detail with reference to the accompanying drawings.

[0059] (Implementation 1)

[0060] <Composition: Voice Translation Device 1>

[0061] Figure 1A This figure shows an example of the appearance of the speech translation device 1 equipped with the speech recognition device 10 according to the first embodiment, and scenes of use of the speech translation device 1 by a first speaker and a second speaker.

[0062] like Figure 1AAs shown, a speech translation device 1 is a device that recognizes a conversation between a first speaker speaking a first language and one or more second speakers speaking a second language, and bidirectionally translates the recognized conversation. In other words, the speech translation device 1 recognizes each speech uttered by the first speaker and one or more second speakers, speaking two different languages, and translates the recognized speech into the other speaker's language. The first language is a language different from the second language. The first and second languages may be Japanese, English, French, German, Chinese, or the like. In this embodiment, a scenario is illustrated in which a first speaker and a second speaker have a face-to-face conversation.

[0063] In this embodiment, the first speaker is the owner of the speech translation device 1, and the first speaker mainly performs operation input to the speech translation device 1. In other words, the first speaker is a user of the speech translation device 1 who understands how to operate the speech translation device 1.

[0064] In this embodiment, before a first speaker speaks, the first speaker performs an operation input to the speech translation device 1, whereupon the speech translation device 1 recognizes the speech of the first speaker speaking in a first language. Once the speech translation device 1 recognizes the speech of the first speaker speaking in the first language, it displays the recognized speech as a first text (characters) in the first language, displays a second text (characters) in the second language obtained by translating the speech in the first language into a second language, and outputs the translated second text in the second language via speech. In this manner, the speech translation device 1 simultaneously outputs the first text recognized by speech, the translated second text, and the speech of the translated second text.

[0065] In this embodiment, after the second speaker speaks, the first speaker inputs an operation to the speech translation device 1, causing the speech translation device 1 to recognize the speech of the second speaker in the second language. Once the speech translation device 1 recognizes the speech of the second speaker in the second language, it displays the recognized speech as a second text in the second language, displays a first text translated from the second language into the first language, and outputs the translated first text as speech. In this manner, the speech translation device 1 simultaneously outputs the recognized second text, the translated first text, and the speech of the translated first text.

[0066] The first speaker and the second speaker have a conversation face to face or side by side using the speech translation device 1. Therefore, the speech translation device 1 can also change the display mode.

[0067] The speech translation apparatus 1 is a portable terminal such as a smartphone or a tablet terminal that can be carried by a first speaker.

[0068] Next, the specific configuration of the speech translation device 1 will be described.

[0069] Figure 2 This is a block diagram showing the speech translation device 1 in the first embodiment.

[0070] like Figure 2 As shown, the speech translation device 1 includes a speech recognition device 10 , a translation processing unit 32 , a display unit 33 , a speech output unit 34 , and a power supply unit 35 .

[0071] [Speech Recognition Device 10]

[0072] The speech recognition device 10 is a device for a first speaker to conduct a conversation with one or more second speakers serving as the first speaker's conversation partners, and is a device for performing speech recognition on speech, which is a conversation between the first speaker speaking in a first language and the second speaker speaking in a second language.

[0073] The speech recognition device 10 includes an input unit 24 , an acquisition unit 21 , a storage unit 22 , an utterance start detection unit 23 , a speaker identification unit 25 , and a speech recognition unit 26 .

[0074] The input unit 24 is an operation input unit that receives at least an operation input from the first speaker. Specifically, the input unit 24 receives an operation input from the first speaker immediately before the first speaker speaks, or receives an operation input from the second speaker immediately after the second speaker speaks. In other words, the input unit 24 receives an operation input from the first speaker each time the first speaker and one or more second speakers speak. The operation input to the input unit 24 triggers whether or not to perform speech recognition on each speech in the conversation between the first speaker and one or more second speakers.

[0075] Furthermore, the input unit 24 may be used as a trigger to start recording of speech into the storage unit 22 , or may be used as a trigger to suspend or stop recording of speech into the storage unit 22 , based on an operation input from the first speaker.

[0076] The input unit 24 generates an input signal corresponding to the operation input and outputs the generated input signal to the utterance start detection unit 23. Furthermore, the input unit 24 generates an input signal including the first time at which the operation input from the first speaker is received and outputs the generated input signal to the speaker distinguishing unit 25. The input signal includes information indicating the first time (time stamp).

[0077] For example, the input unit 24 is a single operation button provided on the speech recognition device 10. Two or more input units 24 may be provided on the speech recognition device 10. In addition, in this embodiment, the input unit 24 is a touch sensor provided integrally with the display unit 33 of the speech translation device 1. In this case, Figure 1B As shown, the display unit 33 of the speech translation apparatus 1 may display a plurality of input units 24 as operation buttons for receiving operation inputs from the first speaker. Figure 1B This is a diagram showing an example of the appearance of another speech translation device in Embodiment 1.

[0078] like Figure 1A As shown, the acquisition unit 21 acquires each speech of a conversation between a first speaker and one or more second speakers. Specifically, the acquisition unit 21 acquires the speech of each of the first speaker and one or more second speakers in the conversation, converts the sound including the acquired speech of the speaker into a speech signal, and outputs the converted speech signal to the storage unit 22.

[0079] The acquisition unit 21 is a microphone unit that acquires a speech signal by converting it into a speech signal containing speech. Alternatively, the acquisition unit 21 may be an input interface electrically connected to a microphone. In other words, the acquisition unit 21 may acquire speech signals from a microphone. Alternatively, the acquisition unit 21 may be a microphone array unit composed of multiple microphones. As long as the acquisition unit 21 is capable of collecting the speech of a speaker present in the vicinity of the speech recognition device 10, the configuration of the acquisition unit 21 in the speech translation device 1 is not particularly limited.

[0080] The storage unit 22 stores the individual speech sounds of the conversation between the first speaker and one or more second speakers acquired by the acquisition unit 21. Specifically, the storage unit 22 stores speech information of the speech sounds included in the speech signals acquired by the acquisition unit 21. In other words, the storage unit 22 automatically stores speech information including the speech sounds uttered by the first speaker and one or more second speakers in the conversation.

[0081] The storage unit 22 resumes recording when the speech recognition device 10 is activated, that is, when the speech translation device 1 is activated. Alternatively, the storage unit 22 may also resume recording from the moment the first speaker first inputs an operation to the input unit 24 after the speech translation device 1 is activated. In other words, the storage unit 22 may also start recording speech in response to an operation input to the input unit 24. Furthermore, the storage unit 22 may also pause or stop recording speech in response to an operation input to the input unit 24.

[0082] Furthermore, after, for example, performing speech recognition on the speech uttered by the first speaker who inputs an operation to the input unit 24, the storage unit 22 begins storing the speech acquired by the acquisition unit 21 in order to store the speech of the second speaker. That is, the storage unit 22 does not store the sound acquired by the acquisition unit 21 at least from the time the speech information of the speech uttered by the first speaker is stored until the speech recognition of the speech is completed.

[0083] Furthermore, since the storage capacity of the storage unit 22 is limited, the oldest voice data may be automatically deleted when the voice information stored in the storage unit 22 reaches a predetermined capacity. In other words, the speaker's voice and information indicating the date and time (time stamp) may be added to the voice information.

[0084] The storage unit 22 is composed of a HDD (Hard Disk Drive) or a semiconductor memory.

[0085] The utterance start detection unit 23 is a detection device that detects the start position of utterance for each speech sound stored in the storage unit 22, after the first speaker has made an operation input to the input unit 24, in accordance with the operation input to the input unit. Specifically, the utterance start detection unit 23 detects the start position of speech sound uttered by the first speaker between the first moment when the first speaker made an operation input to the input unit 24 and the moment when a predetermined period has passed, and represented by speech information stored based on the first speaker's speech. In other words, the utterance start detection unit 23 detects the start position of the second moment, which is the start position of the speech sound uttered by the first speaker, between the first moment when the operation input to the input unit 24 is completed and the moment when the predetermined period has passed.

[0086] Furthermore, the utterance start detection unit 23 detects the start position of a speech indicated by speech information stored based on the speech of the second speaker, which is started by the second speaker between the first time when the first speaker inputs an operation to the input unit 24 and the time that is a predetermined period before the first time, from the individual speech stored in the storage unit 22. In other words, the utterance start detection unit 23 detects the start position of the second time, which is the start of the speech of the speech uttered by the second speaker, between the first time when the operation input to the input unit 24 is completed and the time that is a predetermined period before the first time.

[0087] The utterance start detection unit 23 generates start position information indicating the start position of speech, and outputs the generated start position information to the speaker identification unit 25 and the speech recognition unit 26. The start position information is information (time stamp) indicating the start time of the utterance of the speech uttered by the first speaker, and information (time stamp) indicating the start time of the utterance of the speech uttered by the second speaker.

[0088] When the speaker distinguishing unit 25 receives an input signal from the input unit 24, it distinguishes whether the speaker is the first speaker who has made an operation input to the input unit 24 or the second speaker who has not made an operation input to the input unit 24 based on the first time point set for each voice at which the operation input to the input unit 24 by the first speaker is received and the second time point of the start position of the speech detected by the speech start detection unit 23 based on each voice.

[0089] Specifically, speaker distinguishing unit 25 compares a first time and a second time set for each speech in a conversation between a first speaker and one or more second speakers. More specifically, speaker distinguishing unit 25 compares the first time included in the input signal received from input unit 24 with the second time, which is the utterance start position of the speech within a predetermined period before and after the first time. In this way, speaker distinguishing unit 25 distinguishes whether the speaker is the first speaker or the second speaker.

[0090] For example, when the first time is earlier than the second time, the speaker distinguishing unit 25 determines that the speech uttered by the first speaker is input to the speech recognition device 10 (stored in the storage unit 22), and distinguishes the first speaker from the first speaker and the second speaker. Alternatively, when the second time is earlier than the first time, the speaker distinguishing unit 25 determines that the speech uttered by the second speaker is input to the speech recognition device 10 (stored in the storage unit 22), and distinguishes the second speaker from the first speaker and the second speaker.

[0091] Furthermore, speaker distinguishing unit 25 distinguishes whether a speaker is the first speaker or the second speaker based on each speech uttered by the first speaker and one or more second speakers during a predetermined period of time, which is the period before and after the first moment when input unit 24 receives an operation input from the first speaker. Specifically, during a conversation between one or more first speakers and one or more second speakers, speaker distinguishing unit 25 uses the first moment when input unit 24 receives the operation input as a reference point and selects the most recent speech uttered by the speaker from among the various speech utterances stored in storage unit 22 between the first moment and a moment that is the predetermined period before the first moment, or between the first moment and a moment that is the predetermined period after the first moment. Speaker distinguishing unit 25 uses the selected speech to distinguish whether the speaker is the first speaker or the second speaker. The predetermined period is, for example, a number of seconds, such as one or two seconds, or may be, for example, ten seconds. Thus, the speaker distinguishing unit 25 distinguishes whether the speaker is the first speaker or the second speaker based on the first and second times of the most recently spoken speech of the first speaker and one or more second speakers. This is to avoid the problem that even if the speaker distinguishing unit 25 distinguishes whether the speaker is the first speaker or the second speaker based on speech that is too old, it cannot accurately distinguish whether the speaker who spoke most recently is the first speaker or the second speaker.

[0092] The speaker identifying unit 25 outputs result information including the speaker identification result to the speech recognition unit 26. The result information includes information indicating that the speech information stored based on the speech of the first speaker is the identified first speaker, or information indicating that the speech information stored based on the speech of the second speaker is the identified second speaker.

[0093] Upon receiving result information from the speaker identification unit 25 and start position information from the utterance start detection unit 23, the speech recognition unit 26 performs speech recognition on the speech starting from the start position in each of the speech of the first speaker and one or more second speakers identified by the speaker identification unit 25, based on the result information and the start position information. More specifically, if the speech recognition unit 26 identifies the first speaker from among the first speaker and one or more second speakers, it performs speech recognition in the first language on the speech indicated by the speech information of the most recently uttered speech of the identified first speaker. Furthermore, if the speech recognition unit 26 identifies the second speaker from among the first speaker and one or more second speakers, it performs speech recognition in the second language on the speech indicated by the speech information of the most recently uttered speech of the identified second speaker. Speech recognition refers to the speech recognition unit 26 recognizing the content of the speech uttered by the speaker in the first and second languages. The speech recognition unit 26 generates a first text document and a second text document representing the content of the recognized speech. The speech recognition unit 26 outputs the generated first text and second text to the translation processing unit 32 .

[0094] [Translation Processing Unit 32]

[0095] The translation processing unit 32 translates the recognized language (recognition language) indicated by the text text acquired from the speech recognition unit 26 into another language, and generates a text text indicated by the translated language which is the other language.

[0096] Specifically, upon receiving the first text from the speech recognition unit 26, the translation processing unit 32 translates the first text from the first language indicated in the first text into the second language, generating a second text translated into the second language. The translation processing unit 32 recognizes the content of the second text and generates a translated speech in the second language representing the content of the recognized second text. The translation processing unit 32 outputs the generated first and second texts to the display unit 33 and outputs information representing the generated translated speech in the second language to the speech output unit 34.

[0097] Furthermore, upon receiving the second text from the speech recognition unit 26, the translation processing unit 32 translates the second text from the second language indicated in the second text into the first language, generating a first text translated into the first language. The translation processing unit 32 recognizes the content of the first text and generates a translated speech in the first language representing the content of the recognized first text. The translation processing unit 32 outputs the generated second text and the first text to the display unit 33 and outputs information representing the generated translated speech in the first language to the speech output unit 34.

[0098] Furthermore, the speech translation device 1 may not include the translation processing unit 32, or the cloud server may include the translation processing unit 32. In this case, the speech translation device 1 may be communicatively connected to the cloud server via a network and transmit the first or second text recognized by the speech recognition device 10 to the cloud server. Furthermore, the speech translation device 1 may receive the translated second or first text and the translated speech, output the received second or first text to the display unit 33, and output the received translated speech to the speech output unit 34.

[0099] [Display unit 33]

[0100] The display unit 33 is a monitor such as a liquid crystal panel or an organic EL panel, and displays the first text and the second text received from the translation processing unit 32 .

[0101] The display unit 33 changes the screen layout for displaying the first and second texts according to the positional relationship between the first and second speakers relative to the speech recognition device 10. For example, if the first speaker is speaking, the display unit 33 displays the first text recognized by speech in the area of the display unit 33 located on the first speaker's side, and displays the translated second text in the area of the display unit 33 located on the second speaker's side. Alternatively, if the second speaker is speaking, the display unit 33 displays the second text recognized by speech in the area of the display unit 33 located on the second speaker's side, and displays the translated first text in the area of the display unit 33 located on the first speaker's side. In these cases, the display unit 33 displays the characters of the first and second texts in reversed orientation. Furthermore, when the first and second speakers are conversing side by side, the display unit 33 displays the characters of the first and second texts in the same orientation.

[0102] [Voice output unit 34]

[0103] The speech output unit 34 is a speaker that outputs the translated speech indicated by the information indicating the translated speech obtained from the translation processing unit 32. Specifically, when the first speaker speaks, the speech output unit 34 reproduces and outputs the translated speech having the same content as the second text displayed on the display unit 33. Furthermore, when the second speaker speaks, the speech output unit 34 reproduces and outputs the translated speech having the same content as the first text displayed on the display unit 33.

[0104] [Power supply unit 35]

[0105] The power supply unit 35 is, for example, a primary battery or a secondary battery, and is electrically connected to the speech recognition device 10, the translation processing unit 32, the display unit 33, the speech output unit 34, and the like via wiring. The power supply unit 35 supplies power to the speech recognition device 10, the translation processing unit 32, the display unit 33, the speech output unit 34, and the like. In this embodiment, the power supply unit 35 is provided in the speech translation device 1, but it may also be provided in the speech recognition device 10.

[0106] Action

[0107] The operation of the speech translation device 1 configured as above will be described.

[0108] Figure 3 This is a flowchart showing the operation of the speech translation device 1 in the first embodiment. Figure 4 This is a diagram illustrating the time sequence between the first moment and the second moment when the first speaker and the second speaker are talking. Figure 3 and Figure 4 In the example, a situation in which a first speaker and a second speaker have a one-on-one conversation is assumed. Furthermore, a situation in which the owner of the speech translation device 1 is the first speaker and the first speaker primarily operates the speech translation device 1 is assumed. Furthermore, the speech translation device 1 is preset so that the first speaker speaks in the first language and the second speaker speaks in the second language.

[0109] like Figure 1A 、 Figure 3 and Figure 4 As shown, first, when a first speaker and a second speaker are conversing, the first speaker performs an operation input on input unit 24 before uttering a speech. Specifically, input unit 24 receives the operation input from the first speaker (S11). Specifically, input unit 24 generates an input signal corresponding to the received operation input and outputs the generated input signal to utterance start detection unit 23. Furthermore, input unit 24 generates an input signal including the first moment at which the operation input from the first speaker was received and outputs the generated input signal to speaker identification unit 25.

[0110] Next, the first speaker, who owns the speech recognition device 10 and understands the timing of his or her speech, begins speaking after inputting an operation to the input unit 24. When the first speaker and the second speaker are conversing, the speech recognition device 10 acquires the speech of the first speaker (S12). That is, when the first speaker speaks, the acquisition unit 21 acquires the speech of the first speaker. The acquisition unit 21 converts the acquired speech of the first speaker into a speech signal containing the speech of the acquired speaker and outputs the converted speech signal to the storage unit 22.

[0111] Next, the storage unit 22 stores the voice information of the voice included in the voice signal acquired by the acquisition unit 21 in step S12 (S13). In other words, the storage unit 22 automatically stores the voice information of the most recent voice uttered by one speaker.

[0112] Next, upon receiving an input signal from the input unit 24, the utterance start detection unit 23 detects the start position (second time point) of the utterance start in the speech stored in the storage unit 22 in step S13 (S14). Specifically, the utterance start detection unit 23 detects the start position of the speech indicated by the speech information stored by the first speaker and uttered by the second speaker immediately after the first speaker inputs an operation to the input unit 24.

[0113] The utterance start detection unit 23 generates start position information indicating the start position of speech, and outputs the generated start position information to the speaker identification unit 25 and the speech recognition unit 26 .

[0114] Next, upon receiving an input signal from the input unit 24, the speaker distinguishing unit 25 distinguishes whether the speaker is the first speaker who has made an operation input to the input unit 24 or the second speaker who has not made an operation input to the input unit 24, based on the first time and the second time set for each voice (S15a). Specifically, the speaker distinguishing unit 25 compares the first time and the second time. In other words, the speaker distinguishing unit 25 determines whether the first time is earlier than the second time.

[0115] For example, when the first time is earlier than the second time, the speaker distinguishing unit 25 determines that the speech uttered by the first speaker, who is one of the speakers, is input to the speech recognition device 10 (stored in the storage unit 22), and distinguishes the first speaker from the first speaker and the second speaker. Alternatively, when the second time is earlier than the first time, the speaker distinguishing unit 25 determines that the speech uttered by the second speaker, who is the other speaker, is input to the speech recognition device 10 (stored in the storage unit 22), and distinguishes the second speaker from the first speaker and the second speaker.

[0116] Here, since the first time is earlier than the second time, speaker distinguishing unit 25 determines that the speech information uttered by the first speaker is input into speech recognition device 10 (stored in storage unit 22), and distinguishes the first speaker from the first speaker and the second speaker. Speaker distinguishing unit 25 outputs result information including the speaker distinction result to speech recognition unit 26. The result information includes information indicating that the speech information in step S12 is the distinguished first speaker.

[0117] Next, the speech recognition unit 26 obtains the result information from the speaker distinguishing unit 25 and the start position information from the utterance start detection unit 23, and then performs speech recognition on the speech of the first speaker distinguished by the speaker distinguishing unit 25 based on the result information and the start position information (S16).

[0118] Specifically, the speech recognition unit 26 obtains speech information of the speech most recently uttered by the first speaker in step S12 from the storage unit 22 via the utterance start detection unit 23. The speech recognition unit 26 performs speech recognition on the speech uttered by the first speaker indicated by the speech information obtained from the storage unit 22 via the utterance start detection unit 23.

[0119] More specifically, the speech recognition unit 26 recognizes the content of the speech uttered by the first speaker in the first language and generates a first text document representing the content of the recognized speech. In other words, the content of the first text document matches the content of the speech uttered by the first speaker and is represented in the first language. The speech recognition unit 26 outputs the generated first text document to the translation processing unit 32.

[0120] After receiving the first text from the speech recognition unit 26, the translation processing unit 32 translates the first text from the first language into the second language, thereby generating a second text translated into the second language. In other words, the content of the second text expressed in the second language is consistent with the content of the first text expressed in the first language.

[0121] The translation processing unit 32 recognizes the content of the second text and generates a translated speech in the second language representing the content of the recognized second text.

[0122] The translation processing unit 32 outputs the generated first text and second text to the display unit 33 , and outputs information indicating the generated translated speech in the second language to the speech output unit 34 .

[0123] The display unit 33 displays the first and second text documents received from the translation processing unit 32 (S17). Specifically, the display unit 33 displays the first text document on the screen on the first speaker's side and the second text document on the screen on the second speaker's side. The display unit 33 displays the characters of the first text document in the normal orientation relative to the first speaker so that the first speaker can read the first text document, and displays the characters of the second text document in the normal orientation relative to the second speaker so that the second speaker can read the second text document. In other words, the orientation of the characters of the first text document is reversed relative to the orientation of the characters of the second text document.

[0124] Furthermore, the speech output unit 34 outputs the translated speech in the second language indicated by the information indicating the translated speech in the second language obtained from the translation processing unit 32 (S18). In other words, the speech output unit 34 outputs the translated speech obtained by translating the speech from the first language into the second language. This allows the second speaker, who hears the translated speech in the second language, to understand the speech spoken by the first speaker. Furthermore, since the second text is displayed on the display unit 33, the second speaker can also accurately understand the speech spoken by the first speaker through the characters.

[0125] Next, regarding the second speaker's speech, use Figure 5 Provide explanation. Figure 5 This is a flowchart showing the operation of the speech translation device when the second speaker speaks. Figure 3 The description of the same processing is omitted as appropriate.

[0126] like Figure 1A 、 Figure 4 and Figure 5 As shown, first, the first speaker cannot recognize the timing of the utterance of the second speaker, and therefore performs an operation input to the input unit 24 after the utterance of the second speaker.

[0127] First, when a first speaker and a second speaker are conversing, the speech recognition device 10 acquires the speech uttered by the other speaker (S21). Specifically, when the other speaker speaks, the acquisition unit 21 acquires the speech uttered by the other speaker. The acquisition unit 21 converts the acquired speech uttered by the other speaker into a speech signal containing the speech uttered by the other speaker and outputs the converted speech signal to the storage unit 22.

[0128] Next, the other speaker speaks using a voice in the second language. During a conversation between the first speaker and the second speaker, after the other speaker speaks, the first speaker performs an operation input on input unit 24. Specifically, input unit 24 accepts the operation input from the first speaker (S22). Specifically, input unit 24 outputs an input signal corresponding to the accepted operation input to utterance start detection unit 23, and outputs an input signal including the time (first time) at which the operation input was accepted to speaker identification unit 25.

[0129] Next, the storage unit 22 stores the voice information of the voice included in the voice signal acquired by the acquisition unit 21 in step S21 (S13). That is, the storage unit 22 automatically stores the voice information of the most recent voice uttered by the other speaker.

[0130] Next, the utterance start detection unit 23 detects the start position (second time) of the speech uttered by the other speaker immediately before the first speaker inputs an operation to the input unit 24 and indicated by the speech information stored by the other speaker ( S14 ).

[0131] The utterance start detection unit 23 generates start position information indicating the start position of speech, and outputs the generated start position information to the speaker identification unit 25 and the speech recognition unit 26 .

[0132] Next, the speaker distinguishing unit 25 compares the first time and the second time to determine whether the first time is earlier than the second time, thereby distinguishing whether the other speaker is the first speaker or the second speaker ( S15 b ).

[0133] Here, the second time is earlier than the first time. Therefore, the speech uttered by the second speaker, determined by the speaker distinguishing unit 25 to be the other speaker, is input into the speech recognition device 10 (stored in the storage unit 22), and the second speaker is distinguished from the first speaker and the second speaker. The speaker distinguishing unit 25 outputs result information including the speaker distinction result to the speech recognition unit 26. The result information includes information indicating that the speech information in step S21 is the distinguished second speaker.

[0134] Next, the speech recognition unit 26 obtains the result information from the speaker distinguishing unit 25 and the start position information from the utterance start detection unit 23 , and then performs speech recognition on the speech of the second speaker distinguished by the speaker distinguishing unit 25 based on the result information and the start position information ( S16 ).

[0135] Specifically, the speech recognition unit 26 obtains the speech information of the speech most recently uttered by the second speaker in step S21 from the storage unit 22 via the utterance start detection unit 23. The speech recognition unit 26 performs speech recognition on the speech uttered by the second speaker indicated by the speech information obtained from the storage unit 22 via the utterance start detection unit 23.

[0136] More specifically, the speech recognition unit 26 recognizes the content of the speech uttered by the second speaker in the second language and generates a second text document representing the content of the recognized speech. In other words, the content of the second text document matches the content of the speech uttered by the second speaker and is represented in the second language. The speech recognition unit 26 outputs the generated second text document to the translation processing unit 32.

[0137] Upon receiving the second text from the speech recognition unit 26, the translation processing unit 32 translates the second text from the second language into the first language, thereby generating a first text translated into the first language. In other words, the content of the first text expressed in the first language is consistent with the content of the second text expressed in the second language.

[0138] The translation processing unit 32 recognizes the content of the first text and generates a translated speech in the first language representing the content of the recognized first text.

[0139] The translation processing unit 32 outputs the generated second text and first text to the display unit 33 , and outputs information indicating the generated translated speech in the first language to the speech output unit 34 .

[0140] Display unit 33 displays the second text and the first text received from translation processing unit 32 (S17). Specifically, display unit 33 displays the first text on the screen on the first speaker's side and the second text on the screen on the second speaker's side. On display unit 33, the characters of the first text are displayed in the normal orientation relative to the first speaker so that the first speaker can read the first text, and the characters of the second text are displayed in the normal orientation relative to the second speaker so that the second speaker can read the second text. In other words, the orientation of the characters of the first text is reversed relative to the orientation of the characters of the second text.

[0141] Furthermore, the speech output unit 34 outputs the translated speech in the first language indicated by the information indicating the translated speech in the first language obtained from the translation processing unit 32 (S18). In other words, the speech output unit 34 outputs the translated speech obtained by translating the second language into the first language. This allows the first speaker, who hears the translated speech in the first language, to understand the speech spoken by the second speaker. Furthermore, since the first text is displayed on the display unit 33, the first speaker can also accurately understand the speech spoken by the second speaker through the characters.

[0142] Then, the speech translation apparatus 1 ends the processing.

[0143] Figure 6 This is a flowchart showing the operation of the speaker identifying unit 25 of the speech translation apparatus 1 in the first embodiment. Figure 6 It's about Figure 3 Step S15a and Figure 5 The flowchart specifically explains the processing of step S15b.

[0144] like Figure 3 、 Figure 5 and Figure 6As shown, first, the speaker distinguishing unit 25 takes the first moment when the input unit 24 receives the operation input from the first speaker as a base point, and selects the most recent voice uttered by the speaker from among the various voices stored in the storage unit 22 between the first moment and a moment that is a specified period earlier than the first moment, or between the first moment and a moment that has passed the specified period (S31).

[0145] Next, the speaker identifying unit 25 compares the first time and the second time set each time the first speaker and the second speaker speak, and determines whether the first time is earlier than the second time ( S32 ).

[0146] If the speaker distinguishing unit 25 determines that the first time is earlier than the second time (S32: YES), it distinguishes the first speaker from the first and second speakers (S33). In other words, if the first time is earlier than the second time, it is because the first speaker understands the timing of their own speech, and therefore the first time is earlier than the second time. Thus, the speaker distinguishing unit 25 can distinguish the first speaker from the first and second speakers based on the first and second times.

[0147] The speaker identifying unit 25 outputs result information including the result of identifying the first speaker from the first speaker and the second speaker to the speech recognition unit 26. The speaker identifying unit 25 then ends the processing.

[0148] Furthermore, if the speaker distinguishing unit 25 determines that the second time is earlier than the first time (S32: No), it distinguishes the second speaker from the first and second speakers (S34). In other words, if the second time is earlier than the first time, it is because the first speaker cannot understand the timing of the second speaker's speech and therefore inputs the operation to the input unit 24 after the second speaker's speech, resulting in the second time being earlier than the first time. Thus, the speaker distinguishing unit 25 can distinguish the second speaker from the first and second speakers based on the first and second times.

[0149] The speaker identifying unit 25 outputs result information including the result of identifying the second speaker from the first speaker and the second speaker to the speech recognition unit 26. The speaker identifying unit 25 then ends the processing.

[0150] Effects

[0151] Next, the effects of the speech recognition device 10 in this embodiment will be described.

[0152] As described above, the speech recognition device 10 in this embodiment is a speech recognition device 10 for a first speaker to conduct a conversation with one or more second speakers who are the first speaker's conversation partners. The speech recognition device 10 includes: an acquisition unit 21 for acquiring individual speech sounds of a conversation between the first speaker and the one or more second speakers; a storage unit 22 for storing the individual speech sounds of the conversation between the first speaker and the one or more second speakers acquired by the acquisition unit 21; an input unit 24 for receiving an operation input from at least the first speaker; an utterance start detection unit 23 for detecting the start position of an utterance for each speech sound stored in the storage unit 22 in response to the operation input to the input unit 24; and a speaker distinguishing unit 25 for distinguishing, from among the first speaker and the one or more second speakers, whether the first speaker has made an operation input to the input unit 24 or the one or more second speakers have not made an operation input to the input unit 24, based on a first time point set for each speech at which the operation input to the input unit 24 is received and a second time point indicating the start position of the utterance detected by the utterance start detection unit 23 for each speech sound. Then, the speech from the start position onwards of each speech of the first speaker and the one or more second speakers that have been distinguished by the speaker distinguishing unit 25 is provided to the speech recognition unit 26 for speech recognition.

[0153] Therefore, in this embodiment, the individual voices of the conversation between the first speaker and one or more second speakers are stored in the storage unit 22, so that it is possible to distinguish between the first speaker and the second speaker based on the stored voices. Thus, the voice recognition unit 26 can read the individual voices of the conversation between the first speaker and the second speaker, which have been distinguished, from the storage unit 22 and perform voice recognition. In other words, if the first speaker speaks after the first speaker has input an operation to the input unit 24, the voice recognition unit 26 can perform voice recognition on the voice spoken by the first speaker. Furthermore, the second speaker usually begins speaking after the first speaker has finished speaking, so by the first speaker inputting an operation to the input unit 24 in accordance with the second speaker's speech, the voice recognition unit 26 can perform voice recognition on the voice spoken by the second speaker.

[0154] Therefore, in the speech recognition device 10 , the speech of the second speaker (conversation partner) can be reliably acquired, and speech recognition can be performed on the speech of the second speaker (conversation partner).

[0155] The speech recognition method according to the present embodiment is a speech recognition method for a first speaker to have a conversation with one or more second speakers who are the first speaker's conversation partners, and includes: obtaining speech sounds of a conversation between the first speaker and the one or more second speakers; storing the obtained speech sounds of the conversation between the first speaker and the one or more second speakers in a storage unit 22; accepting an operation input from at least the first speaker to an input unit 24; detecting a start position of an utterance for each speech sound stored in the storage unit 22 in response to the operation input to the input unit 24; distinguishing, from among the first speaker and the one or more second speakers, whether the first speaker has made an operation input to the input unit 24 or one or more second speakers who have not made an operation input to the input unit 24, based on a first time point set for each speech and at which the operation input to the input unit 24 is accepted and a second time point indicating the start position of the utterance detected for each speech sound; and using the speech sounds of the first speaker and the one or more second speakers that have been distinguished, after the start position, for speech recognition.

[0156] This voice recognition method also has the same operational effects as those of the voice recognition device 10 described above.

[0157] Furthermore, the recording medium in this embodiment is a computer-readable nonvolatile recording medium on which a program for causing a computer to execute the speech recognition method is recorded.

[0158] This recording medium also has the same operational effects as those of the above-described voice recognition device 10 .

[0159] In addition, in the speech recognition device 10 of the present embodiment, the speaker distinguishing unit 25 compares the first time and the second time set for each speech in the conversation between the first speaker and one or more second speakers, and distinguishes the first speaker from the first speaker and the one or more second speakers when the first time is earlier than the second time, and distinguishes the second speaker from the first speaker and the one or more second speakers when the second time is earlier than the first time.

[0160] Thus, for example, if the first speaker is the owner of speech recognition device 10, the first speaker understands how to use the speech recognition device 10 and therefore begins speaking after inputting an operation to input unit 24. That is, the first moment at which the first speaker's operation input to input unit 24 is accepted is earlier than the second moment at which the first speaker begins speaking, allowing speaker distinguishing unit 25 to distinguish the first speaker from the first speaker and one or more second speakers. Furthermore, the first speaker cannot recognize the timing of the second speaker's speech and therefore inputs an operation to input unit 24 after the second speaker begins speaking. That is, the first moment at which the first speaker's operation input to input unit 24 is accepted is later than the second moment at which the second speaker begins speaking, allowing speaker distinguishing unit 25 to distinguish the second speaker from the first speaker and one or more second speakers.

[0161] In this way, the speaker distinguishing unit 25 can accurately distinguish whether the speaker who spoke most recently at the first time is the first speaker or the second speaker. Therefore, in the speech recognition device 10, the speech of the second speaker can be more reliably acquired, and speech recognition can be performed on the speech of the second speaker.

[0162] In addition, in the speech recognition device 10 of the present embodiment, when the first speaker is distinguished from the first speaker and one or more second speakers, the speech recognition unit 26 performs speech recognition on the speech uttered by the first speaker, and when the second speaker is distinguished from the first speaker and one or more second speakers, the speech recognition unit 26 performs speech recognition on the speech uttered by the second speaker.

[0163] Thus, by having the speaker distinguishing unit 25 distinguish whether the speaker who spoke is the first speaker or the second speaker, the speech recognition unit 26 can more reliably perform speech recognition on each of the speech uttered by the first speaker and the second speaker.

[0164] In addition, in the speech recognition device 10 of the present embodiment, the speaker distinguishing unit 25 distinguishes whether the speaker is the first speaker or the second speaker based on each speech in the conversation between the first speaker and one or more second speakers during a predetermined period before and after the first moment when the input unit 24 receives the operation input.

[0165] Thus, in order to distinguish between the first and second speakers, a predetermined period can be set based on the first moment. Therefore, it is possible to distinguish whether the most recent speech uttered by the speaker is from the first moment when the first speaker inputs an operation to a moment earlier than the first moment by the predetermined period, or from the first moment to a moment after the predetermined period. This allows for separate recognition of the speech of the first and second speakers. Therefore, in this speech recognition device 10, it is possible to accurately distinguish between the first and second speakers.

[0166] Furthermore, in the speech recognition device 10 of this embodiment, after the speech of the first speaker who inputs an operation to the input unit 24 is recognized, the storage unit 22 starts storing the speech acquired by the acquisition unit 21 to store the speech of the second speaker.

[0167] Typically, the second speaker begins speaking after the first speaker finishes speaking and understands the content of the first speaker's speech. After the first speaker's speech is recognized, recording begins before the second speaker begins speaking. This allows storage unit 22 to reliably store the second speaker's speech. Furthermore, speech recognition device 10 can suspend speech storage at least from the moment the first speaker finishes speaking until storage unit 22 begins recording. This reduces power consumption by speech recognition device 10 for storage in storage unit 22.

[0168] Furthermore, the speech recognition apparatus 10 in this embodiment includes a speech recognition unit 26 that performs speech recognition on the speech starting from the start position of each speech of the first speaker and one or more second speakers that have been distinguished by the speaker distinguishing unit 25 .

[0169] Thus, since the speech recognition is performed by the speech recognition device 10 , there is no need to transmit the speech to the external cloud server, and thus it is possible to suppress an increase in the amount of communication between the speech recognition device 10 and the cloud server.

[0170] In the speech recognition device 10 according to the present embodiment, the input unit 24 is a single operation button provided on the speech recognition device 10 .

[0171] This allows the first speaker to easily operate the speech recognition device 10 .

[0172] Furthermore, in the speech recognition device 10 according to the present embodiment, the input unit 24 receives an operation input from the first speaker each time the first speaker and one or more second speakers each speak.

[0173] Thus, by avoiding requesting the second speaker to perform an operation input as much as possible and allowing the first speaker to actively perform an operation input to the speech recognition device 10 , one speaker can be reliably distinguished from the first speaker and the second speaker.

[0174] (Implementation Method 2)

[0175] <Composition>

[0176] use Figure 7 The configuration of the speech translation device 1 according to this embodiment will be described.

[0177] Figure 7 This is a block diagram showing the speech translation device 1 in the second embodiment.

[0178] In the first embodiment, the speech recognition device 10 includes the speech recognition unit 26 . However, this embodiment is different from the first embodiment in that the speech recognition unit 51 is provided in the cloud server 50 .

[0179] Unless otherwise specified, other configurations in this embodiment are the same as those in the first embodiment, and the same reference numerals are given to the same configurations, and detailed descriptions of the configurations are omitted.

[0180] like Figure 7 As shown, the speech recognition device 10 a includes a communication unit 29 in addition to the input unit 24 , the acquisition unit 21 , the storage unit 22 , the utterance start detection unit 23 , and the speaker identification unit 25 .

[0181] When the speaker identifying unit 25 identifies one speaker from the first speaker and the second speaker, it outputs result information including the speaker identification result to the storage unit 22 .

[0182] Upon acquiring the result information, the storage unit 22 outputs the distinguished voice information of the most recently uttered voice of the speaker to the communication unit 29 .

[0183] The communication unit 29 is a communication module capable of wireless or wired communication with the cloud server 50 having the voice recognition unit 51 via a network.

[0184] The communication unit 29 transmits the speech of the first speaker and one or more second speakers, which have been distinguished by the speaker distinguishing unit 25, to the cloud server 50. Specifically, the communication unit 29 obtains speech information of the speech of the speaker distinguished by the speaker distinguishing unit 25 at the first moment from the storage unit 22 via the utterance start detection unit 23, and transmits the obtained speech information to the cloud server 50 via the network.

[0185] Furthermore, communication unit 29 receives the results of speech recognition performed by speech recognition unit 51 of cloud server 50 on the speech starting from the start position of each speech. Specifically, communication unit 29 receives first and second texts representing the contents of the speech, which are the results of speech recognition performed on the speech of the first speaker and one or more second speakers, from cloud server 50, and outputs the received first and second texts to translation processing unit 32.

[0186] Furthermore, the speech translation device 1 may not include the translation processing unit 32, or the cloud server 50 may also include the translation processing unit 32. In this case, the speech recognition device 10a of the speech translation device 1 may be communicatively connected to the cloud server 50 via a network, and the speech recognition device 10a may transmit the speech of the first speaker and one or more second speakers to the cloud server 50. Furthermore, the speech translation device 1 may receive a first text document, a second text document, and a translated speech document representing the content of the speech document, output the received first text document and second text document to the display unit 33, and output the received translated speech document to the speech output unit 34.

[0187] Effects

[0188] Next, the effects of the speech recognition device 10a in this embodiment will be described.

[0189] As described above, the speech recognition device 10a in this embodiment has a communication unit 29 that can communicate with a cloud server 50 having a speech recognition unit 51. The communication unit 29 sends the respective speech of the first speaker and one or more second speakers that have been distinguished by the speaker distinguishing unit 25 to the cloud server 50, and receives the results obtained by the speech recognition unit 51 of the cloud server 50 performing speech recognition on the speech after the starting position of each speech.

[0190] Thus, each of the speech uttered by the first speaker and one or more second speakers is speech recognized by the cloud server 50 , thereby reducing the processing load of the speech recognition device 10 a .

[0191] In addition, this embodiment has the same operational effects as those of the first embodiment.

[0192] (Other modifications, etc.)

[0193] As mentioned above, although this disclosure was demonstrated based on Embodiment 1 and 2, this disclosure is not limited to these Embodiment 1, 2, etc.

[0194] For example, in the speech recognition device, speech recognition method, and recording medium according to each of the above-mentioned embodiments 1 and 2, the speech recognition device may automatically perform speech recognition corresponding to the speech of the first speaker and the second speaker, and translation into the language of the speech recognition, by pressing the input unit once when translation starts.

[0195] In addition, in the speech recognition devices, speech recognition methods, and recording media involved in each of the above-mentioned embodiments 1 and 2, the direction of the first speaker and one or more second speakers relative to the speech translation device can also be estimated based on the speech acquired by the acquisition unit. In this case, the acquisition unit of the microphone array unit can also be used to estimate the direction of the sound source relative to the speech translation device based on the speech uttered by the first speaker and one or more second speakers. Specifically, the speech recognition device can also calculate the time difference (phase difference) between the speech reaching each microphone in the acquisition unit, for example, using a delay time estimation method to estimate the sound source direction.

[0196] In addition, in the speech recognition device, speech recognition method, and recording medium involved in each of the above-mentioned embodiments 1 and 2, the speech recognition device may not be installed in the speech translation device. For example, the speech recognition device and the speech translation device may be separate devices. In this case, the speech recognition device may also include a power supply unit, and the speech translation device may also include a translation processing unit, a display unit, a speech output unit, and a power supply unit.

[0197] In addition, in the speech recognition devices, speech recognition methods, and recording media according to the first and second embodiments, the speech of the first speaker and one or more second speakers stored in the storage unit may be transmitted to the cloud server via a network for storage in the cloud server, or only the first and second texts obtained by recognizing the speech may be transmitted to the cloud server via a network for storage in the cloud server. In this case, the speech, first and second texts, etc. may be deleted from the storage unit.

[0198] In addition, in the speech recognition device, speech recognition method and recording medium involved in the above-mentioned embodiments 1 and 2, the speech recognition device can also detect the interval of the speaker's speech obtained by the acquisition unit, so that if it is possible to detect a period in which the speaker's speech cannot be obtained for more than a specified period, the recording is automatically terminated or stopped.

[0199] Furthermore, the speech recognition methods according to the above-mentioned embodiments 1 and 2 may be implemented by a program using a computer, and such a program may be stored in a storage device.

[0200] The speech recognition apparatus, speech recognition method, and processing units included in the recording medium according to the above-mentioned embodiments 1 and 2 are typically implemented as LSIs, which are integrated circuits. These can be implemented individually as single chips or in combination with some or all of them.

[0201] Furthermore, integrated circuits are not limited to LSIs and can also be implemented using dedicated circuits or general-purpose processors. FPGAs (Field Programmable Gate Arrays) that can be programmed after LSI manufacturing, or reconfigurable processors that can reconfigure the connections and settings of circuit cells within the LSI, can also be used.

[0202] In addition, in each of the above-mentioned embodiments 1 and 2, each component may be formed by dedicated hardware, or may be implemented by executing a software program suitable for each component. Each component may also be implemented by a program execution unit such as a CPU or a processor reading and executing a software program stored on a recording medium such as a hard disk or semiconductor memory.

[0203] In addition, all the numbers used above are exemplified to specifically describe the present disclosure, and Embodiments 1 and 2 of the present disclosure are not limited to the exemplified numbers.

[0204] The division of functional modules in the block diagram is merely an example. It is also possible to implement multiple functional modules as a single functional module, to divide a single functional module into multiple modules, or to transfer some functions to other functional modules. Furthermore, it is also possible to process the functions of multiple functional modules with similar functions in parallel or in a time-sharing manner by a single hardware or software.

[0205] In addition, the order in which each step in the flowchart is executed is exemplified for the purpose of specifically explaining the present disclosure, and may be an order other than the above-described order. In addition, some of the above-described steps may be executed simultaneously (in parallel) with other steps.

[0206] Other embodiments obtained by various modifications conceived by those skilled in the art to Embodiments 1 and 2, and embodiments implemented by arbitrarily combining the components and functions in Embodiments 1 and 2 without departing from the spirit of the present disclosure, are also included in the present disclosure.

[0207] Industrial Applicability

[0208] The present disclosure is applicable to a speech recognition device, a speech recognition method, and a recording medium used by a plurality of speakers who speak different languages to communicate with each other through conversation.

Claims

1. A speech recognition device for use by a first speaker in a conversation with one or more second speakers who are conversation partners of the first speaker, comprising: an acquisition unit that acquires each speech of a conversation between the first speaker and the one or more second speakers; a storage unit configured to store the respective speech sounds of the conversation between the first speaker and the one or more second speakers acquired by the acquisition unit; an input unit for receiving at least an operation input from the first speaker; an utterance start detecting unit that detects a start position of utterance start for each of the voices stored in the storage unit in response to an operation input to the input unit; as well as a speaker distinguishing unit that distinguishes, from among the first speaker and the one or more second speakers, whether the first speaker has made an operation input to the input unit or the one or more second speakers has not made an operation input to the input unit, based on a first time point set for each voice at which an operation input to the input unit is received and a second time point indicating a start position of an utterance detected by the utterance start detection unit based on each voice; The speech of the first speaker and the one or more second speakers that have been subjected to the discrimination process by the speaker discrimination unit are subjected to speech recognition by the speech recognition unit after the start position of the speech. The speaker distinguishing part is: comparing the first time and the second time set for each of the voices in the conversation between the first speaker and the one or more second speakers, When the first time is earlier than the second time, distinguishing the first speaker from the first speaker and the one or more second speakers; When the second time is earlier than the first time, the one or more second speakers are distinguished from the first speaker and the one or more second speakers.

2. The speech recognition device according to claim 1, When the first speaker is distinguished from the first speaker and the one or more second speakers, the speech recognition unit performs speech recognition on the speech uttered by the first speaker. When the second speaker is distinguished from the first speaker and the one or more second speakers, the speech recognition unit performs speech recognition on a speech uttered by the second speaker.

3. The speech recognition device according to claim 1, The speaker distinguishing unit distinguishes whether a speaker is the first speaker or the one or more second speakers based on each speech in the conversation between the first speaker and the one or more second speakers during a predetermined period, the predetermined period being a period before and after the first moment when the input unit receives an operation input.

4. The speech recognition device according to claim 1, After performing speech recognition on the speech uttered by the first speaker who has performed an operation input to the input unit, the storage unit starts storing the respective speech acquired by the acquisition unit in order to store the speech of the one or more second speakers.

5. The speech recognition device according to claim 1, comprising: a communication unit capable of communicating with the cloud server having the speech recognition unit, The communication unit sends the respective voices of the first speaker and the one or more second speakers that have been distinguished by the speaker distinguishing unit to the cloud server, and receives the results of voice recognition performed by the voice recognition unit of the cloud server on the voices after the start position of the respective voices.

6. The speech recognition device according to claim 1, comprising: The speech recognition unit performs speech recognition on the speech from the start position onwards in the speech of the first speaker and the one or more second speakers that have been subjected to the discrimination process by the speaker discrimination unit.

7. The speech recognition device according to claim 1, The input unit is an operation button provided on the voice recognition device.

8. The speech recognition device according to claim 1, The input unit receives an operation input from the first speaker each time the first speaker and the one or more second speakers speak.

9. A speech recognition method for a first speaker to conduct a conversation with one or more second speakers as conversation partners of the first speaker, comprising: obtaining each speech of a conversation between the first speaker and the one or more second speakers; storing the acquired speech of the conversation between the first speaker and the one or more second speakers in a storage unit; receiving at least an operation input from the first speaker to the input unit; detecting a start position of utterance for each of the voices stored in the storage unit in response to an operation input to the input unit; Based on a first time point set for each voice at which an operation input to the input unit is received and a second time point indicating a start position of an utterance detected based on each voice, distinguishing between the first speaker and the one or more second speakers as the first speaker who has made an operation input to the input unit or the one or more second speakers who have not made an operation input to the input unit; as well as The speech of the first speaker and the one or more second speakers that have been subjected to the distinguishing process and that have been performed thereon, starting from the start position, is used for speech recognition. In the treatment of the distinction, comparing the first time and the second time set for each of the voices in the conversation between the first speaker and the one or more second speakers, When the first time is earlier than the second time, distinguishing the first speaker from the first speaker and the one or more second speakers; When the second time is earlier than the first time, the one or more second speakers are distinguished from the first speaker and the one or more second speakers. 10 . A computer-readable non-volatile recording medium recording a program for causing a computer to execute the speech recognition method according to claim 9 .

Citation Information

Patent Citations

  • Speech translation apparatus and speech translation method

    CN103246643A

  • Voice processor, voice processing method, and voice processing program

    JP2007264473A