Voice input device, voice input method, and recording medium

By detecting the time relationship between the trigger and the speech start position in the voice input device, the speaker identification process is simplified, the calculation amount is reduced and the device structure is simplified.

CN111754986BActive Publication Date: 2025-09-16PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202010206519.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-10-30
Filing Date
2020-03-23
Publication Date
2025-09-16
Estimated Expiration
2040-03-23

AI Technical Summary

Technical Problem

Existing speech recognition devices need to learn in advance the time from user operation to speech, which increases the amount of calculation.

Method used

By acquiring the speaker's voice, storing the voice and detecting the speech start position when trigger input is made, the speaker is identified by using the time relationship between the trigger input time and the speech start position.

Benefits of technology

The processing process is simplified, the amount of calculation is reduced, and speaker recognition is achieved while avoiding a complex device structure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111754986B_ABST
    Figure CN111754986B_ABST
Patent Text Reader

Abstract

A speech input device, a speech input method, and a recording medium. A speaker recognition device (1) comprises: an acquisition unit (21) for acquiring individual speech sounds of one or more speakers when speaking; a storage unit (22) for storing the individual speech sounds of the one or more speakers acquired by the acquisition unit (21); a trigger input unit (23) to which a trigger is input; a speech start detection unit (24) for detecting a speech start position based on the individual speech sounds stored in the storage unit (22) each time a trigger is input to the trigger input unit (23); and a speaker recognition unit (26) for recognizing a speaker from among the one or more speakers based on at least a first time when the trigger input unit (23) is input with a trigger and a second time when the speech start position is detected by the speech start detection unit (24) based on the individual speech sounds.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a voice input device, a voice input method, and a recording medium. Background Art

[0002] For example, Patent Document 1 discloses a speech recognition device comprising: a speech input start operation mechanism that enables speech input operation through user operation; a speech input mechanism that obtains the user's speech; a speech start time learning data holding mechanism that holds a speech start learning time obtained by learning the time from the moment the user operates the speech input start operation mechanism until the user actually starts speaking; and a speech recognition mechanism that compares the measured time with the speech start learning time from the speech start time learning data holding mechanism, determines whether the speech for which the time has been measured is the user's input speech, and performs speech recognition if it is the user's input speech.

[0003] According to this speech recognition device, by learning for each user and using the learned utterance start time, it is possible to recognize whether the speech is from the user.

[0004] Prior art literature

[0005] Patent Literature

[0006] Patent Document 1: Japanese Patent Application Laid-Open No. 2006-313261 Summary of the Invention

[0007] Problems to be solved by the invention

[0008] However, the technology disclosed in Patent Document 1 requires pre-learning the period from when the user operates the voice input device until the user actually starts speaking. Therefore, in conventional voice recognition devices, the amount of calculation required for learning may increase.

[0009] Therefore, an object of the present disclosure is to provide a speech input device, a speech input method, and a recording medium that can identify a speaker through a simple process and suppress an increase in the amount of calculation.

[0010] Means for solving problems

[0011] A speech input device according to one embodiment of the present invention comprises: an acquisition unit for acquiring individual speech sounds when one or more speakers speak; a storage unit for storing the individual speech sounds of the one or more speakers acquired by the acquisition unit; a trigger input unit to which a trigger is input; a speech start detection unit for detecting a start position of speech based on the individual speech sounds stored in the storage unit each time the trigger is input to the trigger input unit; and a speaker recognition unit for recognizing a certain speaker from among the one or more speakers based at least on a first moment when the trigger is input to the trigger input unit and a second moment when the speech start position is detected by the speech start detection unit based on the individual speech sounds.

[0012] Furthermore, some of these specific aspects may be implemented using systems, methods, integrated circuits, computer programs, or computer-readable recording media such as CD-ROMs, or any combination of these.

[0013] Effects of the Invention

[0014] According to the voice input device and the like of the present disclosure, it is possible to identify a speaker through simple processing, thereby suppressing an increase in the amount of calculation. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 This is a diagram showing an external appearance of a speaker identification device according to an embodiment and an example of a usage scenario of the speaker identification device based on an utterance of a speaker.

[0016] Figure 2A This is a block diagram showing an example of a speaker identification device in an embodiment.

[0017] Figure 2B This is a block diagram showing an example of another speaker recognition device in the embodiment.

[0018] Figure 3 This is a flowchart showing the operation of the speaker identification device when the first speaker speaks.

[0019] Figure 4 The diagram exemplifies the time series of the first time and the second time for each speech in the case where the first speaker speaks and the case where the second speaker speaks.

[0020] Figure 5 This is a flowchart showing the operation of the speaker identification device when the second speaker speaks.

[0021] Figure 6 This is a flowchart showing the operation of the speaker identification unit of the speaker identification device according to the embodiment.

[0022] Description of reference numerals:

[0023] 1 Speaker recognition device (voice input device)

[0024] 21 Acquisition Department

[0025] 22 Storage

[0026] 23 Trigger input section

[0027] 24 Speech start detection unit

[0028] 25 Speech Timing Registration Department

[0029] 26 Speaker Recognition Unit DETAILED DESCRIPTION

[0030] A speech input device according to one embodiment of the present invention comprises: an acquisition unit for acquiring individual speech sounds when one or more speakers speak; a storage unit for storing the individual speech sounds of the one or more speakers acquired by the acquisition unit; a trigger input unit to which a trigger is input; a speech start detection unit for detecting a start position of speech based on the individual speech sounds stored in the storage unit each time the trigger is input to the trigger input unit; and a speaker recognition unit for recognizing a certain speaker from among the one or more speakers based at least on a first moment when the trigger is input to the trigger input unit and a second moment when the speech start position is detected by the speech start detection unit based on the individual speech sounds.

[0031] Thus, for example, based on the temporal relationship between a first moment when a trigger by one of the one or more speakers is detected and a second moment when the speaker utters speech, it is possible to identify a particular speaker from among the one or more speakers. In other words, even without learning the period from the first moment to the second moment, it is possible to identify which of the one or more speakers the speech acquired by the acquisition unit is being spoken by.

[0032] Therefore, according to this voice input device, it is possible to identify the speaker through simple processing and suppress an increase in the amount of calculation.

[0033] In particular, the voice input device can identify the speaker of the voice based on the timing of the speech relative to the first moment. Therefore, according to the voice input device, the speaker of the voice can be identified through simple operation. In addition, since the operation of the voice input device is simplified, the complexity of the voice input device, such as configuring multiple buttons on the voice input device, can be suppressed. Therefore, according to this voice input device, for example, in the case where the trigger input part is a button, even a single button can identify which speaker is the one among more than one speakers, thereby making the configuration of the voice input device simpler.

[0034] The speech input method involved in other embodiments of the present disclosure includes: obtaining individual speech sounds when one or more speakers speak; storing the individual speech sounds of the one or more speakers in a storage unit; being triggered by an input; detecting the starting position of the speech based on the individual speech sounds stored in the storage unit each time the trigger is input; and identifying a certain speaker from the one or more speakers based on at least a first moment when the trigger is input and a second moment when the starting position of the speech is detected based on the individual speech sounds.

[0035] This voice input method also has the same effects as the above-mentioned voice input device.

[0036] Furthermore, a recording medium according to another aspect of the present disclosure is a computer-readable nonvolatile recording medium having recorded thereon a program for causing a computer to execute the voice input method.

[0037] This recording medium also has the same effects as the above-mentioned voice input device.

[0038] The voice input device involved in other embodiments of the present disclosure includes: a speech timing registration unit, which registers at least which of the first moment and the second moment is the earlier time; the speaker recognition unit identifies a speaker from the one or more speakers based on the first moment, the second moment, and multiple registration information indicating the timing of the second moment relative to the first moment by the speech timing registration unit.

[0039] Thus, the temporal relationship between the first and second moments can be pre-registered as a desired condition for one or more speakers. Therefore, the speaker identification unit can identify a particular speaker from among the one or more speakers simply by determining whether the temporal relationship between the first and second moments is indicated in the registration information. As a result, this voice input device can more reliably identify a speaker through simplified processing.

[0040] In the voice input device involved in other embodiments of the present disclosure, the speech timing registration unit registers first registration information when registering the timing of each speech of the one or more speakers, and the first registration information is registration information that associates first time information with a certain speaker among the one or more speakers, and the first time information indicates that the second time of the starting position of starting to speak is a later time compared to the first time when the trigger is input to the trigger input unit, and registers second registration information, which is registration information that associates second time information with another certain speaker among the one or more speakers, and the second time information indicates that the second time of the starting position of starting to speak is a more earlier time compared to the first time when the trigger is input to the trigger input unit.

[0041] Thus, the speaker can register a condition such as input trigger before starting to speak, or register a condition such as input trigger after starting to speak. In this way, if the speaker registers the conditions in advance, the voice input device can easily and reliably identify the speaker without learning.

[0042] In the voice input device involved in other embodiments of the present disclosure, the speaker recognition unit calculates the timing of the second moment relative to the first moment, compares the calculated result representing the timing with the multiple registration information, and when the second moment is later than the first moment, determines that the speaker who spoke is the first speaker; when the second moment is earlier than the first moment, determines that the speaker who spoke is the second speaker different from the first speaker.

[0043] Thus, the speaker identification unit can calculate the timing of the second time relative to the first time based on the first time input by the trigger input unit and the second time detected by the utterance start detection unit. The speaker identification unit can thus calculate a result indicating the timing, indicating whether the first time is earlier or later than the second time. As a result, the speaker identification unit can more reliably identify which speaker among one or more speakers is being spoken by comparing the calculated timing result with the plurality of registered information.

[0044] Furthermore, when there are multiple speakers, for example, by registering the period from the first time to the second time, it is possible to identify which speaker is speaking even if there are multiple speakers.

[0045] In the voice input device according to another aspect of the present disclosure, the trigger input unit is a voice input interface that receives input of a preset voice, and the preset voice is input to the trigger input unit as the trigger.

[0046] Thus, the voice input device can perform magic word recognition and speaker identification simply by the speaker uttering a preset voice such as a wake-up word. Therefore, the voice input device has excellent operability.

[0047] In the voice input device involved in other aspects of the present disclosure, the trigger input unit is an operation button provided on the voice input device, and the accepted operation input is input to the trigger input unit as the trigger.

[0048] Thus, the speaker can reliably input a trigger to the trigger input unit by operating the trigger input unit.

[0049] Furthermore, some of these specific aspects may be implemented using systems, methods, integrated circuits, computer programs, or computer-readable recording media such as CD-ROMs, or any combination of these.

[0050] The embodiments described below each represent a specific example of the present disclosure. The numerical values, shapes, materials, components, configuration positions and connection methods of the components, steps, and the order of the steps shown in the following embodiments are examples and are not intended to limit the present disclosure. In addition, among the components in the following embodiments, components that are not described in the independent claims are described as arbitrary components. In addition, in all embodiments, the respective contents can also be combined.

[0051] Hereinafter, a voice input device, a voice input method, and a recording medium according to one embodiment of the present disclosure will be described in detail with reference to the accompanying drawings.

[0052] (Implementation Method)

[0053] <Configuration: Speaker Identification Device 1>

[0054] Figure 1 1 is a diagram showing an external appearance of the speaker identification device 1 according to the embodiment and an example of a usage scene of the speaker identification device 1 based on the speech of the speaker. Figure 1 exemplifies a case where a plurality of speakers share the speaker identification apparatus 1 and use the speaker identification apparatus 1 when speaking.

[0055] like Figure 1 As shown, speaker identification device 1 is a device that obtains speech uttered by one or more speakers and, based on the obtained speech, identifies which of the one or more speakers is speaking. Specifically, speaker identification device 1 obtains individual speech uttered by one or more speakers and identifies the speaker for each of the obtained speech. Speaker identification device 1 is an example of a speech input device.

[0056] Alternatively, the speaker identification device 1 may acquire a conversation between a speaker and a conversation partner, and identify which of the speaker and the conversation partner is the speaker based on the acquired conversation.

[0057] In the present embodiment, the speaker identification apparatus 1 acquires individual voices uttered by one or more speakers, and identifies the speakers based on the acquired individual voices and respective timings (timings) of input triggers.

[0058] In this embodiment Figure 1 , illustrates a scenario where a first speaker and a second speaker, each of which is a multiple speaker, each use speaker identification device 1 and speak. For example, after the first speaker's speech recognition is completed, the second speaker may use speaker identification device 1 indicated by the two-dash line. In other words, speaker identification device 1 can be used by each speaker at their own scheduled time or event, or by the first and second speakers simultaneously during a conversation. The first and second speakers are examples of speakers. Furthermore, the second speaker may also be the first speaker's conversation partner.

[0059] Here, the first and second speakers may speak in the same language or in two different languages. In this case, the speaker identification device 1 identifies whether the speaker is the first or second speaker for each speech produced by the first and second speakers, based on whether the first and second languages ​​are the same or different. For example, the first and second languages ​​may be Japanese, English, French, German, or Chinese.

[0060] In this embodiment, the first speaker is the owner of speaker identification device 1. The first speaker primarily inputs triggers to speaker identification device 1 and registers the timing of the speaker's utterance relative to the input trigger. In other words, the first speaker is a user of speaker identification device 1 who understands how to operate it.

[0061] In this embodiment, a speaker speaks after a trigger is input to speaker identification device 1, whereby speaker identification device 1 recognizes, for example, that the first speaker has spoken. Alternatively, after another speaker speaks, a trigger is input to speaker identification device 1, whereby speaker identification device 1 recognizes, for example, that the second speaker has spoken.

[0062] The speaker recognition device 1 is a portable terminal such as a smartphone or a tablet terminal that can be carried by a first speaker.

[0063] Figure 2A 1 is a block diagram showing the speaker identification device 1 according to the embodiment.

[0064] like Figure 2A As shown, the speaker identification device 1 includes an utterance timing registration unit 25 , an acquisition unit 21 , a storage unit 22 , a trigger input unit 23 , an utterance start detection unit 24 , a speaker identification unit 26 , an output unit 31 , and a power supply unit 35 .

[0065] [Speech timing registration unit 25]

[0066] The utterance timing registering unit 25 registers at least which of the first and second times is earlier. Specifically, the utterance timing registering unit 25 is a registering device that registers the timing of each utterance of one or more speakers relative to the input of a trigger.

[0067] The speech timing registration unit 25 can set desired conditions through operations by one or more speakers and register the set conditions. Specifically, when registering the timing of each speech by one or more speakers, the speech timing registration unit 25 registers first registration information. This first registration information is registration information that associates first time information with one of the one or more speakers. This first time information indicates that the second time at which the speech begins is later than the first time when the trigger is input to the trigger input unit 23. As a specific example, if a condition is set such that the first speaker begins speaking after the trigger is input to the trigger input unit 23, the speech timing registration unit 25 registers first registration information that associates the first time information indicating the set condition with label A. The speech timing registration unit 25 includes a memory for storing the set first registration information. Alternatively, the first registration information set by the speech timing registration unit 25 may be stored in the storage unit 22.

[0068] Furthermore, when registering the timing of each speech, the speech timing registering unit 25 registers second registration information that associates second time information with one of the one or more speakers. This second time information indicates that the second time at which the speech began was earlier than the first time at which the trigger was input to the trigger input unit 23. Specifically, a condition is set such that the second speaker begins speaking before the trigger is input to the trigger input unit 23. The speech timing registering unit 25 registers second registration information that associates the second time information indicating the set condition with the tag B. The speech timing registering unit 25 includes a memory that stores the set second registration information. Furthermore, the second registration information set by the speech timing registering unit 25 may also be stored in the storage unit 22.

[0069] For example, if the first speaker speaks under the conditions of the first registration information set in tag A, and the first speaker prompts the second speaker to speak under the conditions of the second registration information set in tag B (the conditions are previously agreed upon between the first and second speakers), different speakers can speak under different conditions. Therefore, if the conditions for each speech are registered by the speech timing registration unit 25, they become information used by the speaker identification unit 26 to determine the speaker's identity.

[0070] The utterance timing registration unit 25 outputs a plurality of registered information such as the registered first registration information and the second registration information to the speaker identification unit 26 .

[0071] Furthermore, the speech timing registering unit 25 can set the period from the first moment when a trigger is input to the trigger input unit 23 to the second moment when the speaker speaks. In other words, the speech timing registering unit 25 can also register the following condition as registration information: the speaker begins speaking ○○ seconds or more after the first moment when the trigger is input to the trigger input unit 23. Alternatively, the speech timing registering unit 25 can also register the following condition as registration information: the trigger is input to the trigger input unit 23 ○○ seconds or more after the speaker begins speaking. In other words, the speech timing registering unit 25 can set the second moment to ○○ seconds or more after the first moment, and the first moment to ○○ seconds or more after the second moment, and register these settings as registration information. "○○" is an arbitrary number and does not necessarily represent the same time.

[0072] Furthermore, the speech timing registering unit 25 may also register the length of time the trigger is continuously input to the trigger input unit 23 as registration information. For example, if the trigger input unit 23 is an operation button, and the length of time the operation button is pressed and held (continuously input to the trigger input unit 23) in accordance with the timing of the speaker's speech is also registered in advance by the speech timing registering unit 25, the speaker identifying unit 26 can use the registered long-press duration as a criterion for identifying the speaker.

[0073] For example, the utterance timing registering unit 25 may register the following condition as registration information: a trigger is input to the trigger input unit 23 continuously for 00 seconds after 00 seconds or more from the first time a trigger is input to the trigger input unit 23. Alternatively, the utterance timing registering unit 25 may register the following condition as registration information: a trigger is input to the trigger input unit 23 continuously for 00 seconds after 00 seconds or more from the start of the speaker's utterance.

[0074] [Acquisition Section 21]

[0075] The acquisition unit 21 acquires speech from one or more speakers. Specifically, the acquisition unit 21 acquires speech from each of the one or more speakers, converts the acquired speech from the speakers into speech signals, and outputs the converted speech signals to the storage unit 22.

[0076] The acquisition unit 21 is a microphone unit that acquires speech signals by converting speech into speech signals. Alternatively, the acquisition unit 21 may be an input interface electrically connected to a microphone. In other words, the acquisition unit 21 may acquire speech signals from a microphone. Alternatively, the acquisition unit 21 may be a microphone array unit composed of multiple microphones. As long as the acquisition unit 21 is capable of collecting the speech of speakers present in the vicinity of the speaker recognition device 1, the configuration of the acquisition unit 21 in the speaker recognition device 1 is not particularly limited.

[0077] [Storage unit 22]

[0078] The storage unit 22 stores the speech information of each of the one or more speakers acquired by the acquisition unit 21. Specifically, the storage unit 22 stores the speech information of the speech represented by the speech signal acquired by the acquisition unit 21. In other words, the storage unit 22 automatically stores the speech information of each of the one or more speakers.

[0079] Furthermore, the storage unit 22 resumes recording when the speaker identification device 1 is activated. Alternatively, the storage unit 22 may start recording from the moment the speaker first inputs a trigger to the trigger input unit 23 after the speaker identification device 1 is activated. In other words, the storage unit 22 may start recording speech when the speaker first inputs a trigger to the trigger input unit 23. Furthermore, the storage unit 22 may pause or stop recording speech when a trigger is input to the trigger input unit 23.

[0080] Furthermore, the storage unit 22 has a storage capacity limit. Therefore, if the voice information stored in the storage unit 22 reaches a specified capacity, the oldest voice data may be automatically deleted. In other words, the voice information may include the speaker's voice and information indicating the date and time (time stamp). The storage unit 22 deletes the oldest voice information based on the date and time information.

[0081] The storage unit 22 is composed of a HDD (Hard Disk Drive) or a semiconductor memory.

[0082] [Trigger input unit 23]

[0083] A trigger is input into the trigger input unit 23 by a speaker. Specifically, the trigger input unit 23 receives input of a preset trigger from the first speaker before the first speaker speaks. Alternatively, the trigger input unit 23 receives input of a preset trigger from the second speaker after the second speaker speaks. Specifically, the trigger input unit 23 receives input of a trigger from the first speaker before the first speaker speaks, and receives input of a trigger from the second speaker after the second speaker speaks. The trigger input unit 23 receives input of a trigger from each speaker each time one or more speakers speak.

[0084] Furthermore, the trigger input unit 23 can start recording of speech into the storage unit 22 or suspend or stop recording of speech into the storage unit 22 in response to an operation input from the speaker.

[0085] When the trigger input unit 23 detects an input trigger, it generates an input signal and outputs the generated input signal to the utterance start detection unit 24 and the speaker identification unit 26. The input signal includes information indicating a first time (time stamp).

[0086] In this embodiment, the trigger input unit 23 is a single operation button provided on the speaker recognition apparatus 1. In this case, an operation input generated by a speaker pressing the operation button is input as a trigger to the trigger input unit 23. In other words, in this embodiment, the trigger is an input signal provided by the speaker to the trigger input unit 23. Alternatively, the speaker recognition apparatus 1 may include two or more trigger input units 23.

[0087] Alternatively, the trigger input unit 23 may be a touch sensor provided integrally with the display unit 33 of the speaker recognition apparatus 1. In this case, the trigger input unit 23 may be displayed on the display unit 33 of the speaker recognition apparatus 1 as an operation button for accepting an operation input from the speaker.

[0088] Figure 2B This is a block diagram showing an example of another speaker identification device 1 in the embodiment.

[0089] like Figure 2BAs shown, the trigger input unit 23a can also be a voice input interface that accepts input of a preset voice. In this case, the preset voice is input as a trigger to the trigger input unit 23a via the acquisition unit 21a. That is, in this case, the voice uttered by the speaker input to the trigger input unit 23a as a trigger becomes the input signal. Here, the preset voice is a wake-up word, etc. If the speaker recognition device 1 is pre-set to identify the first speaker when the wake-up word is, for example, "OK! ○○, ××" and the second speaker when the wake-up word is, for example, "○○, OK! ××", the speaker is identified as the first speaker when uttering "OK! ○○, ××" and as the second speaker when uttering "○○, OK! ××". In addition, if the trigger input unit 23a is a voice input interface, by setting the speaker for each content of the voice, it is possible to reliably identify each speaker from the first and second speakers.

[0090] [Speech start detection unit 24]

[0091] like Figure 1 and Figure 2A As shown, the utterance start detection unit 24 is a detection device that detects the start position of utterance based on each voice stored in the storage unit 22 every time a trigger is input to the trigger input unit 23 .

[0092] Specifically, the utterance start detection unit 24 detects the start position of speech represented by speech information stored by the first speaker, which was uttered between the first moment when the speaker input a trigger to the trigger input unit 23 and the time when a predetermined period has elapsed, from among the speech information stored in the storage unit 22. In other words, the utterance start detection unit 24 detects the start position of a second moment, which is the start of speech, of speech uttered by the first speaker, between the first moment when the trigger input unit 23 detects the input of a trigger and the time when the predetermined period has elapsed.

[0093] Furthermore, the utterance start detection unit 24 detects the start position of speech indicated by speech information stored by the second speaker, which was started by the second speaker between the first time when the speaker input a trigger to the trigger input unit 23 and the time that is a predetermined period before the first time, from among the speech information stored in the storage unit 22. In other words, the utterance start detection unit 24 detects the start position of the second time, which is the start of speech of the speech uttered by the second speaker, between the first time and the time that is the predetermined period before the first time.

[0094] The utterance start detection unit 24 generates start position information indicating the start position of each speech, and outputs the generated start position information to the speaker identification unit 26. The start position information is information (time stamp) indicating the start time of the speech uttered by the speaker.

[0095] [Speaker Identification Unit 26]

[0096] The speaker recognition unit 26 is a device that identifies a speaker from one or more speakers based on a first moment when a trigger is input to the trigger input unit 23, a second moment when the start position of the speech is detected based on each voice by the speech start detection unit 24, and multiple registration information indicating the timing of the second moment relative to the first moment by the speech timing registration unit 25.

[0097] Specifically, upon receiving the input signal indicated at the first moment from the trigger input unit 23 and the start position information from the utterance start detection unit 24, the speaker identification unit 26 calculates the timing of the second moment relative to the first moment. In other words, the speaker identification unit 26 compares and calculates the temporal relationship between the second moment indicated by the start position information and the first moment indicated by the input signal. The result calculated by the speaker identification unit 26 represents the timing of the second moment relative to the first moment.

[0098] Furthermore, upon receiving registration information from the utterance timing registration unit 25, the speaker identification unit 26 compares the calculated timing of the second time relative to the first time with the plurality of registration information. If the second time is later than the first time, the speaker identification unit 26 determines that the speaker who spoke is the first speaker, thereby identifying the speaker. Furthermore, upon comparing the calculated timing with the plurality of registration information, if the second time is earlier than the first time, the speaker identification unit 26 determines that the speaker who spoke is the second speaker, thereby identifying the speaker.

[0099] More specifically, speaker identification unit 26 determines the speaker based on individual speech sounds uttered by one or more speakers during a predetermined period before and after a first moment when trigger input is received from trigger input unit 23. Speaker identification unit 26 uses the first moment as a reference point and selects the most recent (latest) speech sound uttered by the speaker from among the individual speech sounds stored in storage unit 22 between the first moment and a moment earlier than the first moment by the predetermined period, or between the first moment and a moment after the predetermined period. Speaker identification unit 26 identifies a particular speaker using the selected speech sound.

[0100] Here, the predetermined period is, for example, a few seconds such as 1 or 2 seconds, or may be 10 seconds. Thus, speaker identification unit 26 identifies the speaker based on the first and second time points of each of the most recently uttered speech sounds of one or more speakers. This is to avoid the problem of speaker identification unit 26 being unable to accurately identify the most recently uttered speaker even if it identifies the speaker based on premature speech sounds.

[0101] Speaker identification unit 26 outputs result information including the speaker identification result to output unit 31. The result information includes information indicating which speaker was identified from one or more speakers. For example, the result information includes information indicating that the voice information stored based on the speaker's speech is from the first identified speaker, or information indicating that the voice information stored based on the speaker's speech is from the second identified speaker.

[0102] [Display unit 33]

[0103] The display unit 33 is, for example, a monitor such as a liquid crystal panel or an organic EL panel. The display unit 33 displays the speaker indicated by the result information obtained from the speaker identification unit 26 as a text. For example, if a speaker speaks, the display unit 33 displays that the speaker is the first speaker. Alternatively, if a speaker speaks, the display unit 33 displays that the speaker is the second speaker. The display unit 33 is an example of the output unit 31.

[0104] Furthermore, speaker recognition device 1 may also include a speech output unit. In this case, the speech output unit may be a speaker that outputs speech from the speaker indicated in the result information obtained from speaker recognition unit 26. Specifically, when a speaker speaks, the speech output unit outputs speech indicating that the speaker indicated in the result information is the first speaker. Alternatively, when a speaker speaks, the speech output unit outputs speech indicating that the speaker indicated in the result information is the second speaker. The speech output unit is an example of output unit 31.

[0105] [Power supply unit 35]

[0106] The power supply unit 35 is, for example, a primary battery or a secondary battery, and is electrically connected to the utterance timing registration unit 25, the acquisition unit 21, the storage unit 22, the trigger input unit 23, the utterance start detection unit 24, the speaker recognition unit 26, and the output unit 31 via wiring. The power supply unit 35 supplies power to the utterance timing registration unit 25, the acquisition unit 21, the storage unit 22, the trigger input unit 23, the utterance start detection unit 24, the speaker recognition unit 26, and the output unit 31.

[0107] Action

[0108] The operation of the speaker identification device 1 configured as described above will be described.

[0109] Figure 3 This is a flowchart showing the operation of the speaker identification apparatus 1 when the first speaker speaks. Figure 4 The diagram exemplifies the time series of the first time and the second time for each speech in the case where the first speaker speaks and the case where the second speaker speaks.

[0110] exist Figure 3 and Figure 4 In the embodiment, the utterance timing registering unit 25 is configured such that first registration information is associated with tag A, indicating a condition that the first speaker starts uttering after the speaker inputs a trigger to the trigger input unit 23. Furthermore, the utterance timing registering unit 25 is configured such that second registration information is associated with tag B, indicating a condition that the second speaker starts uttering before the speaker inputs a trigger to the trigger input unit 23.

[0111] like Figure 2A 、 Figure 3 and Figure 4 As shown, first, the trigger input unit 23 receives an input trigger for the acquisition unit 21 to begin acquiring each speech. Specifically, before a speaker utters a speech, the trigger input unit 23 receives input of a trigger preset by the speaker. Thus, the trigger input unit 23 detects the trigger input from the speaker (S11). Upon detecting the trigger input, the trigger input unit 23 generates an input signal and outputs the generated input signal to the utterance start detection unit 24 and the speaker identification unit 26.

[0112] Next, the acquisition unit 21 acquires the speech uttered by one speaker ( S12 ). The acquisition unit 21 converts the acquired speech uttered by one speaker into a speech signal, and outputs the converted speech signal to the storage unit 22 .

[0113] Next, the storage unit 22 stores the voice information of the voice represented by the voice signal acquired by the acquisition unit 21 (S13). In other words, the storage unit 22 automatically stores the voice information of the most recent voice uttered by one speaker.

[0114] Next, upon receiving an input signal from the trigger input unit 23, the utterance start detection unit 24 detects the start position (second time point) of the start of utterance in the speech information stored in the storage unit 22 (S14). Specifically, the utterance start detection unit 24 detects the start position of the speech indicated by the speech information stored by the speaker immediately after the speaker inputs a trigger to the trigger input unit 23. The utterance start detection unit 24 generates start position information indicating the start position of the speech and outputs the generated start position information to the speaker identification unit 26.

[0115] Next, the speaker recognition unit 26 recognizes one speaker from the first speaker and the second speaker based on the first moment when the trigger is input to the trigger input unit 23, the second moment of the speech start position detected by the speech start detection unit 24 based on each voice, and the plurality of registration information indicating the timing of the second moment relative to the first moment by the speech timing registration unit 25 (S15). Figure 3 In the example, the first time is earlier than the second time, so the speaker identification unit 26 identifies the voice (speech voice) of the start position information as the first speaker. In other words, the speaker identification unit 26 identifies one speaker as the first speaker.

[0116] Next, the speaker identification unit 26 outputs result information including the result of identifying the first speaker to the output unit 31 ( S16 ).

[0117] Then, the speaker identification apparatus 1 ends the processing.

[0118] Figure 5 This is a flowchart showing the operation of the speaker identification device 1 when the second speaker speaks. Figure 3 The description of the same processing is omitted as appropriate.

[0119] like Figure 2A 、 Figure 4 and Figure 5 As shown, first, the acquisition unit 21 acquires the speech uttered by the other speaker ( S21 ). The acquisition unit 21 converts the acquired speech uttered by the other speaker into a speech signal and outputs the converted speech signal to the storage unit 22 .

[0120] Next, the trigger input unit 23 receives an input trigger for the acquisition unit 21 to begin acquiring each speech. Specifically, after the other speaker speaks, the trigger input unit 23 receives input of a trigger preset by the speaker. Thus, the trigger input unit 23 detects the trigger input from the speaker (S22). Upon detecting the trigger input, the trigger input unit 23 generates an input signal and outputs the generated input signal to the utterance start detection unit 24 and the speaker identification unit 26.

[0121] Next, the storage unit 22 stores the voice information of the voice represented by the voice signal acquired by the acquisition unit 21 (S13). In other words, the storage unit 22 automatically stores the voice information of the most recent voice uttered by the other speaker.

[0122] Next, upon receiving an input signal from the trigger input unit 23, the utterance start detection unit 24 detects the start position (second time point) of the start of utterance within the speech information stored in the storage unit 22 (S14). Specifically, the utterance start detection unit 24 detects the start position of speech indicated by the speech information stored by the other speaker, which was uttered immediately before the speaker input a trigger to the trigger input unit 23. The utterance start detection unit 24 generates start position information indicating the start position of the speech and outputs the generated start position information to the speaker identification unit 26.

[0123] Next, the speaker recognition unit 26 recognizes one speaker from the first speaker and the second speaker based on the first moment when the trigger is input to the trigger input unit 23, the second moment of the speech start position detected by the speech start detection unit 24 based on each voice, and the plurality of registration information indicating the timing of the second moment relative to the first moment by the speech timing registration unit 25 (S15). Figure 5 In the example, the second time is earlier than the first time, so the speaker identification unit 26 identifies the speech of the start position information as the second speaker. In other words, the speaker identification unit 26 identifies the other speaker as the second speaker.

[0124] Next, the speaker identification unit 26 outputs result information including the result of identifying the second speaker to the output unit 31 ( S16 ).

[0125] Then, the speaker identification apparatus 1 ends the processing.

[0126] Figure 6 This is a flowchart showing the operation of the speaker identification unit 26 of the speaker identification device 1 according to the embodiment.

[0127] like Figure 3 、 Figure 5 and Figure 6 As shown, first, when the speaker recognition unit 26 receives the input signal indicated at the first time from the trigger input unit 23 and the start position information indicated at the second time from the utterance start detection unit 24, it calculates the timing of the second time relative to the first time (S31). In other words, the speaker recognition unit 26 compares and calculates the temporal relationship between the second time and the first time.

[0128] The speaker identification unit 26 compares the calculated result indicating the timing of the second time relative to the first time with the registration information, and determines whether the first time is earlier than the second time ( S32 ).

[0129] When the first time is before the second time, the speaker identification unit 26 determines that the content is the same as that indicated by the first registration information in the registration information ( S32 : Yes), and determines that the speaker who spoke is the first speaker ( S33 ).

[0130] The speaker identification unit 26 outputs result information including the result of identifying the first speaker from among the first speaker and the second speaker to the display unit, and then ends the processing.

[0131] When the first time is later than the second time, the speaker identification unit 26 determines that the content is the same as that indicated by the second registration information in the registration information ( S32 : No), and determines that the speaker who spoke is the second speaker ( S34 ).

[0132] The speaker identification unit 26 outputs result information including the result of identifying the second speaker from among the first speaker and the second speaker to the display unit, and then ends the processing.

[0133] Effects

[0134] Next, the effects of the speaker identification apparatus 1 in this embodiment will be described.

[0135] As described above, the speaker recognition device 1 in this embodiment includes: an acquisition unit 21, which acquires each voice when one or more speakers speak; a storage unit 22, which stores the each voice of the speech of one or more speakers acquired by the acquisition unit 21; a trigger input unit 23, to which a trigger is input; a speech start detection unit 24, which detects the start position of the speech based on each voice stored in the storage unit 22 each time a trigger is input to the trigger input unit 23; and a speaker recognition unit 26, which recognizes a certain speaker from one or more speakers based on at least a first moment when the trigger is input to the trigger input unit 23 and a second moment when the start position of the speech is detected by the speech start detection unit 24 based on each voice.

[0136] Thus, for example, based on the temporal relationship between a first moment when a trigger by one of the one or more speakers is detected and a second moment when the speaker utters the speech, it is possible to identify a particular speaker from among the one or more speakers. In other words, even without learning the period from the first moment to the second moment, it is possible to identify which of the one or more speakers the speech acquired by the acquisition unit 21 is being spoken.

[0137] Therefore, according to the speaker identification apparatus 1 , it is possible to identify the speaker through a simple process and suppress an increase in the amount of calculation.

[0138] In particular, the speaker recognition device 1 can identify the speaker of a speech based on the timing of the utterance relative to the first moment. Therefore, the speaker recognition device 1 can identify the speaker of a speech through simple operation. Furthermore, the operation of the speaker recognition device 1 is simplified, thereby reducing the complexity of the speaker recognition device 1, which would otherwise require multiple buttons. Therefore, according to the speech input device 1, for example, if the trigger input unit 23 is a button, even a single button can identify which speaker is being spoken among one or more speakers, thereby further simplifying the configuration of the speech input device 1.

[0139] In addition, the speech input method in this embodiment includes: obtaining each speech when one or more speakers speak; storing the obtained each speech of one or more speakers in the storage unit 22; being triggered by input; each time the trigger is input, detecting the starting position of the speech based on each speech stored in the storage unit 22; and identifying a certain speaker from one or more speakers based on at least the first moment when the trigger is input and the second moment when the starting position of the speech is detected based on each speech.

[0140] This voice input method also has the same operational effects as those of the above-mentioned speaker recognition device 1 .

[0141] Furthermore, the recording medium in this embodiment is a computer-readable nonvolatile recording medium on which a program for causing a computer to execute the voice input method is recorded.

[0142] This recording medium also has the same operational effects as those of the above-described speaker identification device 1 .

[0143] Furthermore, the speaker identification device 1 in this embodiment includes an utterance timing registering unit 25 that registers at least which of the first and second times is earlier. Furthermore, the speaker identification unit 26 identifies one speaker from among one or more speakers based on the first time, the second time, and a plurality of registration information indicating the timing of the second time relative to the first time, which is recorded by the utterance timing registering unit 25.

[0144] Thus, the temporal relationship between the first and second times can be pre-registered as a desired condition for one or more speakers. Therefore, the speaker identification unit 26 can identify a particular speaker from among the one or more speakers simply by determining whether the temporal relationship between the first and second times is indicated in the registration information. As a result, the speaker identification device 1 can more reliably identify a speaker through simplified processing.

[0145] Furthermore, in the speaker recognition device 1 of the present embodiment, when registering the timing of each utterance of one or more speakers, the utterance timing registering unit 25 registers first registration information that associates first time information with one of the one or more speakers, the first time information indicating that the second time at which the utterance was started is later than the first time at which the trigger was input to the trigger input unit 23. Furthermore, when registering the timing of each utterance, the utterance timing registering unit 25 registers second registration information that associates second time information with another of the one or more speakers, the second time information indicating that the second time at which the utterance was started is earlier than the first time at which the trigger was input to the trigger input unit 23.

[0146] Thus, the speaker can register a condition such as inputting a trigger before starting to speak, or register a condition such as inputting a trigger after starting to speak. In this way, if the speaker registers the conditions in advance, the speaker recognition device 1 can easily and reliably recognize the speaker without learning.

[0147] In addition, in the speaker recognition device 1 of the present embodiment, the speaker recognition unit 26 calculates the timing of the second moment relative to the first moment, compares the calculated result indicating the timing with a plurality of registration information, and determines that the speaker who spoke is the first speaker when the second moment is later than the first moment, and determines that the speaker who spoke is the second speaker who is different from the first speaker when the second moment is earlier than the first moment.

[0148] Thus, speaker identification unit 26 can calculate the timing of the second time relative to the first time based on the first time input to trigger input unit 23 and the second time detected by utterance start detection unit 24. Speaker identification unit 26 can thus calculate a result indicating the timing, indicating whether the first time is earlier or later than the second time. Consequently, speaker identification unit 26 can more reliably identify which speaker among one or more speakers is being spoken by comparing the calculated timing result with a plurality of registered information.

[0149] Furthermore, when there are multiple speakers, for example, by registering the period from the first time to the second time, it is possible to identify which speaker is speaking even if there are multiple speakers.

[0150] In the speaker recognition device 1 according to the present embodiment, the trigger input unit 23 is a voice input interface that receives input of a preset voice. The preset voice is input to the trigger input unit 23 as a trigger.

[0151] Thus, the speaker recognition device 1 can perform magic word recognition and speaker identification simply by the speaker uttering a preset voice such as a wake-up word.

[0152] In the speaker identification apparatus 1 of the present embodiment, the trigger input unit 23 is an operation button provided on the speaker identification apparatus 1. The received operation input is input to the trigger input unit 23 as a trigger.

[0153] Thus, the speaker can reliably input a trigger to the trigger input unit 23 by operating the trigger input unit 23 .

[0154] (Other modifications, etc.)

[0155] As mentioned above, although this disclosure was demonstrated based on embodiment, this disclosure is not limited to these embodiment etc.

[0156] For example, in the speech input device, speech input method, and recording medium involved in each of the above-mentioned embodiments, the direction of the speaker relative to the speech input device can also be estimated based on the speech acquired by the acquisition unit. In this case, the acquisition unit of the microphone array unit can also be used to estimate the sound source direction of each speaker's speech relative to the speech input device. Specifically, the speech input device can also calculate the time difference (phase difference) between the speech reaching each microphone in the acquisition unit, for example, by estimating the sound source direction using a delay time estimation method.

[0157] In addition, in the voice input device, voice input method and recording medium involved in the above-mentioned embodiments, the voice input device can also detect the interval of the speaker's voice obtained by the acquisition unit, so that if it is detected that the acquisition unit cannot obtain the speaker's voice for more than a specified period of time, the recording is automatically terminated or stopped.

[0158] Furthermore, the voice input method according to each of the above-described embodiments may be implemented by a program using a computer, and such a program may be stored in a storage device.

[0159] In addition, the processing units included in the speech input device, speech input method, and program thereof according to the above embodiments are typically implemented as LSIs, which are integrated circuits. These can be implemented individually as single chips or in combination with some or all of them.

[0160] Furthermore, integrated circuits are not limited to LSIs and can also be implemented using dedicated circuits or general-purpose processors. FPGAs (Field Programmable Gate Arrays) that can be programmed after LSI manufacturing, or reconfigurable processors that can reconfigure the connections and settings of circuit cells within the LSI, can also be used.

[0161] In addition, in each of the above embodiments, each component may be formed by dedicated hardware, or implemented by executing a software program suitable for each component. Each component may also be implemented by a program execution unit such as a CPU or a processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory.

[0162] In addition, all the numbers used above are exemplified to specifically describe the present disclosure, and the embodiments of the present disclosure are not limited to the exemplified numbers.

[0163] The division of functional modules in the block diagram is merely an example. It is also possible to implement multiple functional modules as a single functional module, to divide a single functional module into multiple modules, or to transfer some functions to other functional modules. Furthermore, it is also possible to process the functions of multiple functional modules with similar functions in parallel or in a time-sharing manner by a single hardware or software.

[0164] In addition, the order in which each step in the flowchart is executed is exemplified for the purpose of specifically explaining the present disclosure, and may be an order other than the above-described order. In addition, some of the above-described steps may be executed simultaneously (in parallel) with other steps.

[0165] Other modes obtained by various modifications conceived by those skilled in the art to implement the embodiments, and modes implemented by arbitrarily combining the components and functions in the embodiments without departing from the gist of the present disclosure are also included in the present disclosure.

[0166] Industrial Applicability

[0167] The present disclosure is applicable to a speech input device, a speech input method, and a storage medium used to identify a speaker who made each of a plurality of speakers' speeches.

Claims

1. A voice input device comprising: An acquisition unit, for acquiring each voice of one or more speakers when speaking; a storage unit configured to store the respective speech sounds of the one or more speakers acquired by the acquisition unit; A trigger input section, triggered by input; an utterance start detecting unit configured to detect a start position of utterance based on the respective voices stored in the storage unit each time the trigger is input to the trigger input unit; a speaker recognition unit that recognizes a speaker from among the one or more speakers based on at least a first time when the trigger is input to the trigger input unit and a second time when the start position of the speech is detected by the utterance start detection unit based on the respective voices; as well as The speech timing registering unit registers at least which of the first time and the second time is earlier. The speaker identification unit identifies one speaker from among the one or more speakers based on the first time, the second time, and a plurality of registration information indicating a timing of the second time relative to the first time by the utterance timing registration unit.

2. The voice input device according to claim 1, When the speech timing registration unit registers the timing of each speech of two speakers, registering first registration information, the first registration information being registration information associating first time information with one of the two speakers, the first time information indicating that the second time of the start position of starting speech is later than the first time when the trigger is input into the trigger input unit, Registering second registration information, which is registration information that establishes an association between second time information and the other one of the two speakers, the second time information indicating that the second time of the start position of starting to speak is a time that is earlier than the first time when the trigger is input to the trigger input unit.

3. The voice input device according to claim 2, The speaker recognition unit is: Calculating the timing of the second moment relative to the first moment, The calculated presentation timing result is compared with the plurality of registration information, and when the second moment is later than the first moment, it is determined that the speaker who spoke is the first speaker; and when the second moment is earlier than the first moment, it is determined that the speaker who spoke is a second speaker different from the first speaker.

4. The voice input device according to claim 1 or 2, The trigger input unit is a voice input interface that accepts a preset voice input. A preset voice is input as the trigger to the trigger input unit.

5. The voice input device according to claim 1 or 2, The trigger input unit is an operation button provided on the voice input device. The accepted operation input is input to the trigger input unit as the trigger.

6. A voice input method comprising: Get the individual voices of more than one speaker; storing the acquired speech sounds of the one or more speakers in a storage unit; Triggered by input; detecting a starting position of utterance based on each of the voices stored in the storage unit each time the trigger is input; as well as identifying a speaker from among the one or more speakers based on at least a first time when the trigger is input and a second time when the start position of the utterance is detected based on each of the voices; In the voice input method, At least registering which of the first time and the second time is earlier, In the recognition, one speaker is recognized from among the one or more speakers based on the first time, the second time, and a plurality of registration information indicating a timing of the second time relative to the first time.

7. A computer-readable non-volatile recording medium recording a program for causing a computer to execute the voice input method according to claim 6.

Citation Information

Patent Citations

  • Voice recognition device and voice recognition program and computer readable recording medium with the voice recognition program stored

    JP2006313261A