Speech recognition method and speech recognition device
Patent Information
- Application Number
- JP2025031892
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2026-09-09
AI Technical Summary
【0006】 本発明の音声認識方法によれば、音声認識状況に応じて閾値を手動又は自動で調整することができ、発話者が新たな話者であるか否かを判定する精度を向上させることができる。
Smart Images

Figure 2026144538000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a speech recognition method and a speech recognition apparatus. Background Art
[0002] A speech recognition method is known, which stores a speaker in association with recognition parameters when the speaker is a newly recognized speaker (see, for example, Patent Document 1). Prior Art Literature Patent Literature
[0003] Patent Document 1 Japanese Unexamined Patent Application Publication No. 2001-005482 Summary of the Invention Problem to be Solved by the Invention
[0004] However, with conventional speech recognition methods, it is difficult to determine whether a speaker is a new speaker or not. The present invention has been made in view of such circumstances, and provides a speech recognition method capable of improving the accuracy of determining whether a speaker is a new speaker or not. Means for Solving the Problem
[0005] The present invention provides a speech recognition method comprising: an input speech data acquisition step of acquiring input speech data; an adjustment step of adjusting a first threshold based on predetermined conditions; and an identification step of calculating similarity for each registered user by comparing the input speech data or processed input speech data generated by processing the input speech data with registered speech data of utterances of registered users that have been registered in advance, and identifying which registered user the input speech data or processed input speech data belongs to based on the similarity, wherein in the identification step, if it is determined that the maximum value of the similarity calculated for each registered user is lower than the first threshold, the speaker of the input speech data or processed input speech data is registered as a new speaker. [Effects of the Invention]
[0006] According to the speech recognition method of the present invention, the threshold can be adjusted manually or automatically according to the speech recognition status, thereby improving the accuracy of determining whether the speaker is a new speaker or not. [Brief explanation of the drawing]
[0007] [Figure 1] This is a flowchart of the speech recognition method in one embodiment of the present invention. [Figure 2] This is a block diagram showing the configuration of a speech recognition device in one embodiment of the present invention. [Figure 3] This is a flowchart of the speech recognition method in one embodiment of the present invention. [Figure 4] This is a flowchart of the speech recognition method in one embodiment of the present invention. [Figure 5] This is a flowchart of the speech recognition method in one embodiment of the present invention. [Modes for carrying out the invention]
[0008] The speech recognition method of the present invention includes an input speech data acquisition step of acquiring input speech data; an adjustment step of adjusting a first threshold based on predetermined conditions; and an identification step of calculating similarity for each registered user by comparing the input speech data or processed input speech data generated by processing the input speech data with registered speech data of utterances of registered users that have been registered in advance, and identifying which registered user the input speech data or processed input speech data belongs to based on the similarity, wherein in the identification step, if it is determined that the maximum value of the similarity calculated for each registered user is lower than the first threshold, the speaker of the input speech data or processed input speech data is registered as a new speaker. The speech recognition method of the present invention may also be a speaker identification method.
[0009] Preferably, the identification step involves comparing a numerical vector obtained by transforming the features of the input voice data or the processed input voice data with a numerical vector obtained by transforming the registered voice data to calculate the similarity for each registrant, and identifying which of the previously registered registrants the input voice data or the processed input voice data is spoken by based on the similarity, and if it is determined that the maximum value of the similarity calculated for each registrant is lower than the first threshold, the step involves registering the speaker of the input voice data or the processed input voice data as a new speaker. Preferably, the speech recognition method of the present invention includes an operation detection step for detecting an operation by a user to adjust the first threshold, wherein the adjustment step adjusts the first threshold based on the operation by the user to adjust the first threshold.
[0010] Preferably, the speech recognition method of the present invention further includes a reliability determination step of determining the reliability of speech recognition of the input speech data or the processed input speech data, wherein in the adjustment step, a first threshold is automatically adjusted according to the reliability of the speech recognition. Preferably, the speech recognition method of the present invention includes a time length determination step in which the time length of the input speech data or the time length of the processed input speech data is determined to be longer than a second threshold, and if in the determination step it is determined that the time length of the input speech data or the time length of the processed input speech data is longer than the second threshold, and in the identification step it is determined that the maximum value of the similarity calculated for each registered user is lower than the first threshold, the speaker of the input speech data or the processed input speech data is registered as a new speaker.
[0011] Preferably, in the identification step, if the maximum similarity value calculated for each registrant is higher than the first threshold but lower than the third threshold, the speaker of the input voice data or the processed input voice data is determined to be an unknown speaker; if the maximum similarity value calculated for each registrant is greater than the third threshold, the speaker of the input voice data or the processed input voice data is identified as the registrant whose similarity value is the maximum. Preferably, the speech recognition method of the present invention further includes a reliability determination step of determining the reliability of speech recognition of the input speech data or the processed input speech data, wherein in the adjustment step, the first and third thresholds are automatically adjusted according to the reliability of the speech recognition.
[0012] Preferably, the processed input audio data is data generated by removing silent portions of the audio data from the input audio data. Preferably, the speech recognition method of the present invention includes a speaker integration step of integrating the new speaker with a pre-registered registrant, and if the number of times the new speaker is integrated with a pre-registered registrant exceeds a predetermined number of times, the adjustment step automatically adjusts the first threshold so that the first threshold becomes smaller.
[0013] The present invention also provides a speech recognition device comprising: an input speech data acquisition unit that acquires input speech data; an adjustment unit that adjusts a first threshold based on predetermined conditions; and an identification unit that calculates a similarity for each registered user by comparing the input speech data or processed input speech data generated by processing the input speech data with registered speech data of a registered user's utterance that has been registered in advance, and identifies which registered user the input speech data or processed input speech data belongs to based on the similarity, wherein the identification unit determines that the maximum value of the similarity calculated for each registered user is lower than the first threshold, and registers the speaker of the input speech data or processed input speech data as a new speaker.
[0014] The present invention will be described in more detail below with reference to several embodiments. The configurations shown in the drawings and the following description are illustrative, and the scope of the present invention is not limited to those shown in the drawings and the following description.
[0015] First Embodiment Figure 1 is a flowchart of the speech recognition method according to the first embodiment, and Figure 2 is a block diagram showing the configuration of the speech recognition device. The speech recognition method of the first embodiment includes an input speech data acquisition step (e.g., step S3) for acquiring input speech data, an adjustment step (e.g., step S2) for adjusting a first threshold based on predetermined conditions, and an identification step (e.g., steps S7, S9-S11, S16, etc.) for calculating a similarity for each registered user by comparing the input speech data or processed input speech data generated by processing the input speech data with registered speech data of utterances of registered users that have been registered in advance, and identifying which registered user the input speech data or processed input speech data belongs to based on the similarity. In the identification step, if it is determined that the maximum value of the similarity calculated for each registered user is lower than the first threshold, the speaker of the input speech data or processed input speech data is registered as a new speaker (e.g., step S17). Furthermore, the speech recognition method according to the first embodiment further includes an operation detection step (for example, step S1) of detecting an operation by a user to adjust the first threshold value, and in the adjustment step (for example, step S2), the first threshold value is adjusted based on the user's operation to adjust the first threshold value. Step S15 is a time length determination step.
[0016] The speech recognition method according to the first embodiment can be implemented, for example, by a speech recognition apparatus 20 as shown in FIG. 2. The speech recognition apparatus 20 according to the first embodiment includes: an input audio data acquisition unit 3 that acquires input audio data; an adjustment unit 4 that adjusts a first threshold value based on a predetermined condition; and an identification unit 5 that calculates a similarity for each registrant by comparing the input audio data or processed input audio data generated by processing the input audio data with registered audio data of an utterance of a pre-registered registrant, and identifies which pre-registered registrant's utterance the input audio data or the processed input audio data is based on the similarity, wherein the identification unit 5 is characterized by registering a speaker of the input audio data or the processed input audio data as a new speaker when it is determined that the maximum value of the similarity calculated for each registrant is lower than the first threshold value.
[0017] The speech recognition apparatus 20 includes a control unit 2, and the input audio data acquisition unit 3, the adjustment unit 4, and the identification unit 5 may be programs included in the control unit 2. The control unit 2 may include an output unit. Furthermore, the speech recognition apparatus 20 may include a microphone 10, a display unit 11, an input unit, and the like. The speech recognition apparatus 20 may also be connected to a computer (user interface) via a wired or wireless connection. The speech recognition apparatus 20 may also be connected to a computer (user interface) via a network such as a LAN. In this case, a monitor of the computer can be used as the display unit 11. A user can also input information such as the speaker's name using the computer. The speech recognition device 20 may be included in an automatic speech recognition-based minute system, a speech recognition-based conversation recording system, or a speech-to-text conversion system. The control unit 2 may include a processor, a storage unit, a communication unit, and the like. The processor may include at least one of, for example, a CPU, an MPU, a GPU, an NPU, and the like. The storage unit is a RAM, a storage, or the like. The communication unit is a portion provided to connect to the Internet, a local area network, or the like. The control unit 2 can be connected to the microphone 10 so that the audio signal output from the microphone 10 can be input thereto. Furthermore, the control unit 2 (output unit) can be connected to a user interface such as the display unit 11 so as to output the recognition result of the speech recognition method of the present embodiment to the user interface.
[0018] A specific example of the speech recognition method according to the first embodiment will be described with reference to the flowchart shown in FIG. 1. In step S1, the control unit 2 determines whether or not a threshold value has been changed. This threshold value includes at least one of, for example, a threshold value A used in step S11, a threshold value B used in step S12, a threshold value C used in step S15, and a threshold value D used in step S16. The speech identification method of the present embodiment may include a selection step in which a user selects between a "mode with lenient speaker similarity determination" and a "mode with strict speaker similarity determination". The set threshold value differs depending on the selected mode. When the control unit 2 detects that the user has not changed the mode in step S1, the process proceeds to step S3. When the control unit 2 detects that the user has changed the mode in step S1, the process proceeds to step S2.
[0019] In step S2, the control unit 2 (adjustment unit 4) adjusts the threshold. For example, if the control unit 2 detects a change in mode in step S1, the control unit 2 adjusts at least one of the thresholds used in step S11 (A), step S12 (B), step S15 (C), and step S16 (D) to match the mode in step S2. This allows the user to change the similarity determination criteria according to the situation, thereby improving speaker identification accuracy. The process then proceeds to step S3.
[0020] In step S3, the control unit 2 (input audio data acquisition unit 3) acquires input audio data from the microphone 10 or the like. Alternatively, in step S3, the control unit 2 may use VAD (Voice Activity Detection) to detect speech segments and acquire the audio data of these speech segments as input audio data. Alternatively, in step S3, the control unit 2 may accumulate and concatenate the input audio data and acquire the audio data with silent portions removed as processed input audio data. For example, the control unit 2 can determine whether each audio data is a silent portion and exclude the silent audio data from the concatenated audio data to create processed input audio data.
[0021] In step S4, the control unit 2 performs speech recognition on the input voice data or processed input voice data and obtains text data corresponding to the input voice data or processed input voice data. For example, the control unit 2 can perform speech recognition processing using an AI model. If the control unit 2 has an AI model stored, the control unit 2 can perform speech recognition processing. Alternatively, the control unit 2 may transmit the input voice data to a server on the Internet or a local area network via the communication unit, have the server perform speech recognition processing, and receive the result via the communication unit. In speech recognition processing, text data is generated from input speech data or processed input speech data by performing filtering processes such as utterance detection (VAD detection), language detection, confidence level determination of the recognition result, and text formatting (removal of hallucination symbols, etc.).
[0022] In step S5, the control unit 2 performs speaker separation processing on the input audio data or processed input audio data. In the speaker separation processing, if the audio data contains only the utterances of one speaker, the control unit 2 detects the utterance segments of that one speaker included in the audio data. If the audio data contains utterances from multiple speakers, the control unit 2 detects the utterance segments of each speaker included in the audio data. For example, if the audio data contains utterances by speaker A, speaker B, speaker A, speaker C, and speaker B in this order, the control unit 2 detects the utterance segments of speaker A's first utterance, speaker B's first utterance, speaker A's second utterance, speaker C's utterance, and speaker B's second utterance. The control unit 2 can perform speaker separation processing, for example, using an AI model for speaker separation. If the control unit 2 has stored an AI model, it can perform speaker separation processing. Alternatively, the control unit 2 may transmit input voice data to a server on the Internet or a local area network via the communication unit, have the server perform speaker separation processing, and receive the results via the communication unit.
[0023] In step S6, the control unit 2 determines whether the number of speakers in the utterances included in the input voice data or the processed input voice data is one. If the control unit 2 determines in step S6 that there is one speaker, the process proceeds to step S7, where the control unit 2 (identification unit 5) performs a process to vectorize the input voice data or processed input voice data (for example, calculation of Embedding). Then, the process proceeds to step S10. If the control unit 2 determines in step S6 that there are multiple speakers, in step S8 the control unit 2 calculates the total utterance time of each speaker included in the input audio data or processed input audio data. Then, proceeding to step S9, the control unit 2 (identification unit 5) performs a process to vectorize the audio data of the speaker with the longest total utterance time included in the input audio data or processed input audio data (for example, calculation of Embedding). After that, proceeding to step S10.
[0024] In step S10, the control unit 2 (identification unit 5) compares the speaker information of multiple registered users stored in the memory unit with the vector of input speech data or the vector of processed input speech data generated in step S7 or S9 and calculates the similarity (for example, cosine similarity). The memory unit of the control unit 2 stores speaker information for multiple registered users. For example, the memory unit of the control unit 2 stores the registered user's voice data and its vector along with the registered user's name as speaker information. Speaker information may be stored in the memory unit before implementing the speech recognition method of this embodiment, or it may be stored in the memory unit while repeatedly implementing the speech recognition method of this embodiment.
[0025] Cosine similarity takes values in the range of -1 to 1 and is used to calculate the relationship or similarity between vectorized audio data. With cosine similarity, the smaller the angle between two vectors, that is, the closer the value is to 1, the more similar the two audio data vectors are considered to be. In step S10, for example, the control unit 2 (identification unit 5) compares the vector of input voice data or the vector of processed input voice data generated in step S7 or S9 with the vector of voice data for each registrant stored in the storage unit, and calculates the cosine similarity for each registrant.
[0026] In step S11, the control unit 2 (identification unit 5) determines whether the largest cosine similarity value among all cosine similarities calculated in step S10 is greater than threshold A. Threshold A can be the lowest value in the range of cosine similarity in which the speaker of the input voice data or processed input voice data and the registered person can be considered to be the same person. If threshold A is set high, the similarity judgment becomes stricter because the speaker will not be considered to be the same person as the registered person unless the similarity is high, thus suppressing the control unit 2 from judging different speakers as the same person. For example, if the user selects the "mode with strict speaker similarity judgment," in step S2 the control unit 2 can set threshold A higher than the "mode with lenient speaker similarity judgment." Lowering threshold A can suppress the control unit 2 from misinterpreting utterances from the same speaker as those from different speakers. For example, if the user selects the "mode with lenient speaker similarity determination," in step S2, the control unit 2 can set threshold A lower than in the "mode with strict speaker similarity determination." These modes can be selected by the user according to the situation.
[0027] If, in step S11, the control unit 2 (identification unit 5) determines that the largest value among all cosine similarities is greater than threshold A, then in step S12, the control unit 2 determines whether the speech duration of the input voice data or processed input voice data acquired in step S1 is longer than threshold B. Threshold B can be the minimum speech duration required to determine, based on cosine similarity, that the speaker is the same person as the registrant. This threshold B can be adjusted (changed) in step S2. If the control unit 2 determines in step S12 that the utterance time is longer than threshold B, the process proceeds to step S13. The control unit 2 outputs the text data obtained by speech recognition in step S4 to the user interface, such as the display unit 11, as the utterance of the registrant corresponding to the highest cosine similarity. The control unit 2 also outputs to the user interface that it has a high degree of confidence that the utterance is that of the registrant. Specifically, the display unit 11 can display both the name of the registrant corresponding to the highest cosine similarity and the text data obtained by speech recognition in step S4. In this case, the control unit 2 can display a high degree of confidence along with the registrant's name on the display unit 11. For example, the control unit 2 may choose not to display a question mark (?) along with the registrant's name. The process then returns to step S1.
[0028] If the control unit 2 determines in step S12 that the utterance time is shorter than threshold B, the process proceeds to step S14, where the control unit 2 outputs the text data obtained by speech recognition in step S4 to the user interface, such as the display unit 11, as the utterance of the registrant corresponding to the highest cosine similarity. The control unit 2 also outputs to the user interface that the confidence level of the utterance being that of the registrant is relatively low. Specifically, the display unit 11 can display both the registrant's name corresponding to the highest cosine similarity and the text data obtained by speech recognition in step S4. At this time, the control unit 2 can display the registrant's name along with the indication that the confidence level is relatively low. For example, the control unit 2 can display a question mark (?) along with the registrant's name. The process then returns to step S1.
[0029] If, in step S11, the control unit 2 (identification unit 5) determines that the largest value among all cosine similarities is less than threshold A, then in step S15, it determines whether the speech duration of the input voice data or processed input voice data is longer than threshold C. Threshold C can be the minimum speech duration required to determine, based on cosine similarity, whether the speaker is the same person as the registrant. This threshold C can be adjusted (changed) in step S2.
[0030] If, in step S15, the control unit 2 determines that the speech duration of the input audio data or the processed input audio data is shorter than threshold C, in step S19, the control unit 2 outputs the text data obtained by speech recognition in step S4 to a user interface such as the display unit 11 as "Speaker name unknown". In this case, the user can input the speaker's name to the control unit 2 via an input device such as a keyboard, or select the name of an already registered speaker. If a speaker's name is entered or a registered speaker's name is selected, the control unit 2 corrects the "Speaker name unknown" display on the user interface to the entered speaker's name or the selected registered speaker's name.
[0031] If the speaker's name entered by the user is not registered, the control unit 2 stores the input voice data or processed input voice data acquired in step S3, the voice data vector generated in step S7 or S9, and the entered name in the storage unit as speaker information. This speaker information can be used as the registered speaker information in the subsequent step S10. If the speaker name entered by the user is already registered, or if the user selects the name of a registered user, the control unit 2 can integrate the input voice data or processed input voice data acquired in step S3 and the voice data vector generated in step S7 or S9 with the registered user's speaker information and store it in the storage unit. Then, return to step S1.
[0032] If, in step S15, the control unit 2 determines that the speech duration of the input voice data or processed input voice data is longer than threshold C, then in step S16, the control unit 2 determines whether the largest cosine similarity value among all the cosine similarities calculated in step S10 is less than threshold D. Threshold D can be the maximum value within the range of cosine similarities in which the speaker of the input voice data or processed input voice data can be considered to be a different person from any of the registered users. This threshold D can be adjusted (changed) in step S2.
[0033] If, in step S16, the control unit 2 determines that the largest cosine similarity value is greater than the threshold D, then in step S19, the control unit 2 outputs the text data obtained by speech recognition in step S4 to a user interface such as the display unit 11, indicating "Speaker name unknown". After that, the process returns to step S1.
[0034] If, in step S16, the control unit 2 determines that the largest cosine similarity value among all cosine similarities is less than threshold D, then in step S17, the control unit 2 considers the speaker to be a "new speaker" and stores (registers) the input speech data or processed input speech data acquired in step S3 and the speech data vector generated in step S7 or S9 as speaker information for the "new speaker" in the memory unit. If multiple "new speakers" are registered as the speech recognition method flow is repeated, they can be distinguished, for example, as "new speaker A," "new speaker B," "new speaker C," etc. Also, if a "new speaker" is registered, in the subsequent step S10, the "new speaker" registered in step S17 is included in the list of registered speakers.
[0035] Subsequently, the process proceeds to step S18, where the control unit 2 outputs the text data obtained by speech recognition in step S4 to a user interface such as the display unit 11 as a statement from a "new speaker". In this case, the user can input the speaker's name to the control unit 2 via an input device such as a keyboard, or select the name of an already registered user. If a speaker's name is entered or a registered user's name is selected, the control unit 2 updates the "New Speaker" display on the user interface to the entered speaker's name or the selected registered user's name.
[0036] If the speaker name entered by the user is not registered, the control unit 2 corrects the speaker information for the "new speaker" to the speaker information of the speaker name entered by the user. If the speaker name entered by the user is already registered, or if the user selects the name of an already registered registrant, the control unit 2 integrates the speaker information of the "new speaker" with the speaker information of that registrant (speaker integration step). Then, return to step S1.
[0037] Second Embodiment Figure 3 is a flowchart of the speech recognition method according to the second embodiment. The speech recognition method of the second embodiment is the same as the speech recognition method of the first embodiment, except that steps S1 and S2 are omitted, steps S21, S22 and S23 are performed, and steps S31 and S32 are performed instead of steps S12 to S14. In the second embodiment, step S21 is a confidence determination step, steps S22 and S23 are adjustment steps, and steps S7, S9 to S11, S16, S31, etc. are included in the identification steps. Steps S3, S4, S5-S11, and S15-S19 were explained in the first embodiment and will be omitted here. The explanation will focus on steps S21-S23 and steps S31 and S32.
[0038] In step S4, as described in the first embodiment, after performing speech recognition, in step S21, the control unit 2 determines whether the reliability of the speech recognition performed in step S4 is high or low. In step S21, the control unit 2 can determine whether the reliability of the speech recognition is high or low based on the results of filtering processes performed in the speech recognition process of step S4, such as speech determination (VAD determination), language determination, determination of the confidence level of the recognition result, and text formatting (removal of hallucinations such as symbols). For example, if the sound is clearly picked up by the microphone 10, the reliability of the speech recognition is considered to be high. Conversely, if the sound is not clearly picked up by the microphone 10, the reliability of the speech recognition is considered to be low.
[0039] If the control unit 2 determines in step S21 that the reliability of speech recognition is high, in step S22 the control unit 2 (adjustment unit 4) increases the threshold A used in step S11, or if threshold A is already at a high value, it does not change the threshold. In this case, the control unit 2 (adjustment unit 4) may also increase the threshold D used in step S16 and / or the threshold E used in step S31, or if threshold D and / or threshold E are already at a high value, it does not change the threshold. This improves the accuracy of speaker identification.
[0040] If the control unit 2 determines in step S21 that the reliability of speech recognition is low, in step S23 the control unit 2 (adjustment unit 4) lowers the threshold A used in step S11, or if threshold A is already at a low value, it does not change the threshold. In this case, the control unit 2 (adjustment unit 4) may also lower the threshold D used in step S16 and / or the threshold E used in step S31, or if threshold D and / or threshold E are already at a low value, it does not change the threshold. If the reliability of speech recognition is low, the accuracy of similarity determination decreases, and even utterances from the same speaker may be judged as having low similarity. However, by using a low threshold, it is possible to suppress the registration of already registered speakers as "new speakers" in step S17. After the control unit 2 adjusts the threshold in steps S22 and S23, in step S5 the control unit 2 performs speaker separation of the input audio data or the processed input audio data. Details of step S5 were described in the first embodiment and are therefore omitted here.
[0041] In step S11 described in the first embodiment, if the control unit 2 determines that the largest value among all cosine similarities is greater than threshold A, then in step S31, the control unit 2 can determine whether the largest value among all cosine similarities is greater than threshold E. Threshold E can be set to a value higher than threshold A, and can be set to a value higher than the lowest value in the range of cosine similarity in which the speaker of the input voice data or processed input voice data and the registrant can be considered to be the same person. This can improve the accuracy of speaker identification.
[0042] If, in step S31, the control unit 2 determines that the largest cosine similarity value among all cosine similarities is greater than the threshold E, then in step S32, the control unit 2 outputs the text data obtained by speech recognition in step S4 to a user interface such as the display unit 11 as the registrant's statement corresponding to the largest cosine similarity. Specifically, both the registrant's name corresponding to the largest cosine similarity and the text data obtained by speech recognition in step S4 can be displayed on the display unit 11. After that, the process returns to step S1.
[0043] If, in step S31, the control unit 2 determines that the largest value among all cosine similarities is less than the threshold E, then in step S19, the control unit 2 outputs the text data obtained by speech recognition in step S4 as "Speaker name unknown" to a user interface such as the display unit 11. Details of step S19 were explained in the first embodiment and are therefore omitted here. Furthermore, the description of the first embodiment described above also applies to the second embodiment, unless otherwise contradictory.
[0044] Third Embodiment Figure 4 is a flowchart of the speech recognition method according to the third embodiment. The speech recognition method of the third embodiment is the same as the speech recognition method of the first embodiment, except that steps S1 and S2 are omitted and steps S41 to S43 are performed. In the third embodiment, step S43 is an adjustment step, and for example, steps S7, S9 to S11, S16, etc. are included in the identification steps. Steps S3 to S19 were explained in the first embodiment and will be omitted here. We will focus on explaining steps S41 to S43.
[0045] In step S41, the control unit 2 determines whether or not the user has input to integrate the speaker information of the "new speaker" registered in step S17 with the speaker information of the registrant. The integration of speaker information is as described in step S18 of the first embodiment. If the control unit 2 determines in step S41 that there was no user input to integrate speaker information, the process proceeds to step S3, and the control unit 2 acquires the voice data. If the control unit 2 determines in step S41 that there has been user input to integrate speaker information, the control unit 2 determines in step S42 whether the number of integrations has exceeded a predetermined number. Specifically, the control unit 2 determines whether a certain number of speaker integrations (e.g., 3 times) have occurred within a given time period (e.g., 10 minutes). If the number of integrations is high, it is likely that the system has already identified many instances of registered utterances as utterances from new speakers.
[0046] If the control unit 2 determines in step S42 that the number of integrations has not exceeded a predetermined number, the process proceeds to step S3, and the control unit 2 acquires the audio data. If the control unit 2 determines in step S42 that the number of integrations exceeds a predetermined number, in step S43 the control unit 2 (adjustment unit 4) adjusts threshold A used in step S11 and / or threshold D used in step S16. Specifically, the control unit 2 can lower threshold A and / or threshold D. This makes it less likely for already registered users to be registered as "new speakers," thereby improving registration accuracy. Then, the process proceeds to step S3, where the control unit 2 acquires audio data. Furthermore, the description of the first embodiment described above also applies to the third embodiment, unless otherwise contradictory.
[0047] Fourth Embodiment Figure 5 is a flowchart of the speech recognition method according to the fourth embodiment. The speech recognition method of the fourth embodiment is the same as the speech recognition method of the first embodiment, except that steps S1 and S2 are omitted and steps S51 and S52 are performed. In the fourth embodiment, step S52 is an adjustment step, and for example, steps S7, S9 to S11, S16, etc. are included in the identification steps. Steps S3 to S19 were explained in the first embodiment and will be omitted here. We will focus on explaining steps S51 and S52.
[0048] In step S51, the control unit 2 determines whether the number of speakers exceeds a predetermined number. Specifically, the control unit 2 determines whether the number of speakers (registered users and "new speakers") output along with the text data in steps S13, S14, and S18 exceeds a predetermined number (for example, 10 people). If a "new speaker" is merged with a registered user, that "new speaker" is not added to the count.
[0049] If the control unit 2 determines in step S51 that the number of speakers does not exceed a predetermined number, the process proceeds to step S3, and the control unit 2 acquires the audio data. If the control unit 2 determines in step S51 that the number of speakers exceeds a predetermined number, in step S52 the control unit 2 (adjustment unit 4) adjusts the threshold A used in step S11 and / or the threshold D used in step S16. Specifically, the control unit 2 can increase the threshold A and / or the threshold D. When there are many speakers, the control unit 2 is more likely to acquire audio data of speakers (registered speakers and "new speakers") whose voices are highly similar. Therefore, in step S52, by increasing threshold A and / or threshold D, the control unit 2 can suppress the misidentification of different speakers as the same speaker, thereby improving registration accuracy. Then, the process proceeds to step S3, where the control unit 2 acquires audio data. Furthermore, the description of the first embodiment described above also applies to the fourth embodiment, unless otherwise contradictory. [Explanation of Symbols]
[0050] 2: Control Unit 3: Input Voice Data Acquisition Unit 4: Adjustment Unit 5: Identification Unit 10: Microphone 11: Display Unit 20: Voice Recognition Device
Claims
1. An input audio data acquisition step to acquire input audio data, An adjustment step to adjust the first threshold based on predetermined conditions, The process includes an identification step of comparing the input audio data or processed input audio data generated by processing the input audio data with registered audio data of utterances of registered users that have been registered in advance, thereby calculating a similarity for each registered user, and identifying which registered user the input audio data or processed input audio data belongs to based on the similarity. A speech recognition method characterized in that, in the identification step, if it is determined that the maximum value of the similarity calculated for each registered user is lower than the first threshold, the speaker of the input speech data or the processed input speech data is registered as a new speaker.
2. The speech recognition method according to claim 1, wherein the identification step is a step of calculating the similarity for each registrant by comparing a numerical vector obtained by converting the features of the input speech data or the processed input speech data with a numerical vector obtained by converting the registered speech data, and identifying which of the previously registered registrants the input speech data or the processed input speech data is spoken by based on the similarity, and if it is determined that the maximum value of the similarity calculated for each registrant is lower than the first threshold, the step of registering the speaker of the input speech data or the processed input speech data as a new speaker.
3. The method further includes an operation detection step for detecting an operation by a user to adjust the first threshold, The speech recognition method according to claim 1, wherein the first threshold is adjusted based on an operation by the user to adjust the first threshold in the adjustment step.
4. The method further includes a reliability determination step for determining the reliability of speech recognition of the input audio data or the processed input audio data, The speech recognition method according to claim 1, wherein in the adjustment step, the first threshold is automatically adjusted according to the reliability of the speech recognition.
5. The method further includes a time length determination step of determining whether the time length of the input audio data or the time length of the processed input audio data is longer than a second threshold, The speech recognition method according to claim 1, wherein in the determination step, it is determined that the duration of the input speech data or the duration of the processed input speech data is longer than a second threshold, and in the identification step, it is determined that the maximum value of the similarity calculated for each registered user is lower than a first threshold, and the speaker of the input speech data or the processed input speech data is registered as a new speaker.
6. The speech recognition method according to claim 1, wherein in the identification step, if the maximum value of the similarity calculated for each registrant is higher than a first threshold and lower than a third threshold, the speaker of the input voice data or the processed input voice data is determined to be an unknown speaker, and if the maximum value of the similarity calculated for each registrant is greater than a third threshold, the speaker of the input voice data or the processed input voice data is identified as the registrant whose similarity is the maximum value.
7. The method further includes a reliability determination step for determining the reliability of speech recognition of the input audio data or the processed input audio data, The speech recognition method according to claim 6, wherein in the adjustment step, the first and third thresholds are automatically adjusted according to the reliability of the speech recognition.
8. The speech recognition method according to claim 1, wherein the processed input audio data is data generated by removing silent portions of audio data from the input audio data.
9. The system further includes a speaker integration step that integrates the aforementioned new speaker with a pre-registered registrant, The speech recognition method according to claim 1, wherein if the number of times the new speaker is integrated into the pre-registered registered speakers exceeds a predetermined number of times, the first threshold is automatically adjusted in the adjustment step so that the first threshold becomes smaller.
10. An input audio data acquisition unit that acquires input audio data, An adjustment unit that adjusts a first threshold value based on predetermined conditions, The system includes an identification unit that calculates a similarity for each registered user by comparing the input voice data or processed input voice data generated by processing the input voice data with registered voice data of utterances of registered users that have been registered in advance, and identifies which registered user the input voice data or processed input voice data belongs to based on the similarity. The speech recognition device is characterized in that, when the identification unit determines that the maximum value of the similarity calculated for each registered user is lower than the first threshold, it registers the speaker of the input speech data or the processed input speech data as a new speaker.
Citation Information
Patent Citations
Voice recognizing method and device
JP2001005482A