Speech recognition method and speech recognition device
Patent Information
- Application Number
- JP2025031899
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2026-09-09
AI Technical Summary
【0006】 本発明の音声認識方法によれば、発話時間が短い場合でも優れた精度で話者識別を行うことができる。
Smart Images

Figure 2026144542000001_ABST
Abstract
Description
[[Technical Field]]
[0001] The present invention relates to a speech recognition method and a speech recognition apparatus. [[Background Art]]
[0002] A speaker identification method is known which calculates a speaker feature vector from input speech and performs speaker recognition (see, for example, Patent Document 1). [[Prior Art Documents]] [[Patent Documents]]
[0003] [[Patent Document 1]] Japanese Unexamined Patent Application Publication No. 2017-187642 [[Summary of the Invention]] [[Problem to be Solved by the Invention]]
[0004] However, in conventional speaker identification methods, when the utterance duration is short, the speaker identification accuracy tends to deteriorate. This poses a problem when performing transcription and speaker identification in real time, and particularly in active meetings, utterance durations are often short, which makes speaker identification difficult. The present invention has been made in view of such circumstances, and provides a speech recognition method having excellent speaker identification accuracy even when the utterance duration is short. [[Means for Solving the Problem]]
[0005] The present invention provides a speech recognition method comprising: an input speech data acquisition step of acquiring input speech data; a determination step of determining whether the duration of the input speech data is longer than a threshold; and an identification step of identifying which registered utterance the input speech data belongs to by comparing the input speech data or processed input speech data generated by processing the input speech data with registered speech data of a registered utterance by threshold in the determination step of a determination step of whether the duration of the input speech data is longer than a threshold; and the present invention provides a speech recognition method comprising: an input speech data acquisition step of acquiring input speech data; a determination step of determining whether the duration [Effects of the Invention]
[0006] According to the speech recognition method of the present invention, speaker identification can be performed with excellent accuracy even when the speaking time is short. [Brief explanation of the drawing]
[0007] [Figure 1] This is a flowchart of the speech recognition method in one embodiment of the present invention. [Figure 2] This is a block diagram showing the configuration of a speech recognition device in one embodiment of the present invention. [Figure 3] This is a flowchart of the speech recognition method in one embodiment of the present invention. [Figure 4] This is a flowchart of the speech recognition method in one embodiment of the present invention. [Figure 5] This is a flowchart of the speech recognition method in one embodiment of the present invention. [Modes for carrying out the invention]
[0008] The speech recognition method of the present invention includes an input speech data acquisition step of acquiring input speech data; a determination step of determining whether the duration of the input speech data is longer than a threshold; and an identification step of identifying which registered speaker the input speech data belongs to by comparing the input speech data or processed input speech data generated by processing the input speech data with registered speech data of a registered speaker that has been registered in advance. The method is characterized in that, if the duration of the input speech data is determined to be shorter than the threshold in the determination step, the identification step performs a process to improve identifiability. The speech recognition method of the present invention may also be a speaker identification method.
[0009] Preferably, the speech recognition method of the present invention includes a first processed input speech data generation step in which, if the duration of the input speech data is determined to be shorter than the threshold in the determination step, a first processed input speech data is generated by concatenating multiple copies of the input speech data, and in the identification step, the input speech data is identified as the utterance of a registered user by comparing the first processed input speech data with the registered speech data. Preferably, in the speech recognition method of the present invention, if the duration of the input speech data is determined to be shorter than the threshold in the determination step, the identification step identifies which registered utterance the input speech data belongs to by comparing the input speech data or the processed input speech data with a first registered speech data of a registered utterance from
[0010] Preferably, if the determination step determines that the duration of the input voice data is shorter than the threshold, the system includes a second processed input voice data generation step which generates a second processed input voice data by concatenating the input voice data and the registered voice data, and in the identification step, the system identifies which registered user the input voice data is uttered by comparing the second processed input voice data with the registered voice data. Preferably, in the speech recognition method of the present invention, in the identification step, the second processed input speech data is compared with the registered speech data to identify which registered speaker the second processed input speech data belongs to for each of a plurality of speech segments with different speakers, and the registered speaker identified as the speaker of the speech segment corresponding to the input speech data is identified as the speaker of the input speech data.
[0011] The present invention also provides a speech recognition method that includes: a speech data acquisition step of acquiring input speech data; a determination step of determining whether the duration of the input speech data is longer than a threshold; an identification step of identifying which of the previously registered registrants the input speech data belongs to by comparing the input speech data or processed input speech data generated by processing the input speech data with registered speech data of a previously registered registrant; and an output step of outputting the result of identifying which of the previously registered registrants the input speech data belongs to in the identification step. In this speech recognition method, if the duration of the input speech data is determined to be shorter than the threshold in the determination step, the output step outputs that the confidence level of the speaker is low, and if the duration of the input speech data is determined to be longer than the threshold in the determination step, the output step outputs that the confidence level of the speaker is high. In the speech recognition method of the present invention, it is preferable to identify which registered user the input speech data belongs to by comparing a numerical vector obtained by transforming the features of the input speech data or the processed input speech data with a numerical vector obtained by transforming the registered speech data.
[0012] The present invention also provides a speech recognition device comprising: an input speech data acquisition unit that acquires input speech data; a determination unit that determines whether the duration of the input speech data is longer than a threshold; and an identification unit that identifies which registered speaker the input speech data belongs to by comparing the input speech data or processed input speech data generated by processing the input speech data with registered speech data of a registered speaker that has been registered in advance. The device is characterized in that, if the determination unit determines that the duration of the input speech data is shorter than the threshold, the identification unit performs a process to improve identifiability. The speech recognition device of the present invention may also be a speaker identification device.
[0013] The present invention also provides a speech recognition device comprising: an input speech data acquisition unit that acquires input speech data; a determination unit that determines whether the duration of the input speech data is longer than a threshold; an identification unit that identifies which registered utterance the input speech data belongs to by comparing the input speech data or processed input speech data generated by processing the input speech data with registered speech data of a registered utterance by
[0014] The present invention will be described in more detail below with reference to several embodiments. The configurations shown in the drawings and the following description are illustrative, and the scope of the present invention is not limited to those shown in the drawings and the following description.
[0015] First Embodiment Figure 1 is a flowchart of the speech recognition method according to the first embodiment, and Figure 2 is a block diagram showing the configuration of the speech recognition device. The speech recognition method according to the first embodiment comprises: an input speech data acquisition step of acquiring input speech data (for example, step S1); a determination step of determining whether or not a time length of the input speech data is longer than a threshold (for example, step S3); and an identification step of identifying which pre-registered registrant's utterance the input speech data corresponds to by comparing the input speech data or processed input speech data generated by processing the input speech data with registered speech data of utterances from pre-registered registrants (for example, steps S4, S7, S9, S10, S11, etc.).
[0016] When it is determined in the determination step (for example, step S3) that the time length of the input speech data is shorter than the threshold, processing for improving distinguishability is performed in the identification step. Specifically, when it is determined in the determination step (for example, step S3) that the time length of the input speech data is shorter than the threshold, a processed input speech data generation step of generating the processed input speech data by concatenating a plurality of copies of the input speech data (for example, step S4) is performed. Further, in the identification step, which pre-registered registrant's utterance the input speech data corresponds to is identified by comparing the processed input speech data with the registered speech data (for example, steps S7, S9, S10, S11, etc.).
[0017] The speech recognition method according to the first embodiment can be implemented, for example, by a speech recognition apparatus 20 as shown in FIG. 2. The speech recognition apparatus 20 according to the first embodiment comprises: an input speech data acquisition unit 3 that acquires input speech data; a determination unit 4 that determines whether or not a time length of the input speech data is longer than a threshold; and an identification unit 5 that identifies which pre-registered registrant's utterance the input speech data corresponds to by comparing the input speech data or the processed input speech data with registered speech data of utterances from pre-registered registrants, wherein when the determination unit 4 determines that the time length of the input speech data is shorter than the threshold, the identification unit 5 performs processing for improving distinguishability.
[0018] The speech recognition device 20 includes the control unit 2, and the input speech data acquisition unit 3, the determination unit 4, and the identification unit 5 may be programs included in the control unit 2. The control unit 2 may include the output unit 6. Further, the speech recognition device 20 may include a microphone 10, a display unit 11, an input unit, and the like. Further, the speech recognition device 20 may be connected to a computer via a wired or wireless connection. In this case, a monitor of the computer can be used as the display unit 11. A user can also input information such as the name of a speaker using the computer. The speech recognition device 20 may be included in an automatic speech recognition-based minute-taking system, a speech recognition-based conversation recording system, or a speech-to-text system. The control unit 2 can include a processor, a storage unit, a communication unit, and the like. The processor can include at least one selected from, for example, a CPU, an MPU, a GPU, an NPU, and the like. The storage unit is a RAM, a storage, or the like. The communication unit is a part provided to be connected to the Internet, a local area network, or the like. The control unit 2 can be connected to the microphone 10 so as to be capable of receiving an audio signal output from the microphone 10. Further, the control unit 2 (the output unit 6) can be connected to a user interface such as the display unit 11 so as to be capable of outputting a recognition result obtained by the speech recognition method of the present embodiment to the user interface.
[0019] A specific example of the speech recognition method according to the first embodiment will be described with reference to the flowchart shown in FIG. 1. In step S1, the control unit 2 (input audio data acquisition unit 3) acquires input audio data from the microphone 10 or the like. Alternatively, in step S1, the control unit 2 may use VAD (Voice Activity Detection) to detect a speech segment and acquire the audio data of this speech segment as input audio data. Alternatively, in step S1, the control unit 2 may accumulate and concatenate the audio data and remove silent portions to acquire the resulting audio data as input audio data. For example, the control unit 2 can determine whether each audio data is a silent portion and exclude the silent audio data from the concatenated audio data to create the input audio data.
[0020] In step S2, the control unit 2 performs speech recognition on the input voice data and obtains text data corresponding to the input voice data. For example, the control unit 2 can perform speech recognition processing using an AI model. If the control unit 2 has an AI model stored, the control unit 2 can perform speech recognition processing. Alternatively, the control unit 2 may transmit the input voice data to a server on the Internet or a local area network via the communication unit, have the server perform speech recognition processing, and receive the results via the communication unit.
[0021] In step S3, the control unit 2 (determination unit 4) determines whether the speech duration of the input voice data is longer than threshold A. Threshold A can be the shortest speech duration required for accurate speaker recognition. If the control unit 2 (determination unit 4) determines in step S3 that the speech duration of the input voice data is longer than threshold A, the process proceeds to step S5. In this case, the input voice data is used instead of the processed input voice data in the subsequent steps.
[0022] If, in step S3, the control unit 2 (determination unit 4) determines that the utterance time of the input voice data is shorter than threshold A, then in step S4, the control unit 2 (identification unit 5) generates processed input voice data by concatenating multiple copies of the input voice data. For example, if the utterance time of the input voice data is 2 seconds and four identical input voice data are concatenated, the processed input voice data will be voice data in which the same utterance is repeated four times, resulting in an utterance time of 8 seconds. By concatenating multiple copies of the input voice data in this way to lengthen the utterance time, the accuracy of speaker identification can be improved. The number of copies of the input audio data to be concatenated is not particularly limited, but for example, it can be set to a number such that the speech duration of the processed input audio data becomes longer than threshold A. After generating the processed input audio data in step S4, the process proceeds to step S5. In this case, the processed input audio data, rather than the original input audio data, is used in subsequent steps.
[0023] In step S5, the control unit 2 performs speaker separation processing on the input audio data or processed input audio data. In the speaker separation processing, if the audio data contains only the utterances of one speaker, the control unit 2 detects the utterance segments of that one speaker included in the audio data. If the audio data contains utterances from multiple speakers, the control unit 2 detects the utterance segments of each speaker included in the audio data. For example, if the audio data contains utterances by speaker A, speaker B, speaker A, speaker C, and speaker B in this order, the control unit 2 detects the utterance segments of speaker A's first utterance, speaker B's first utterance, speaker A's second utterance, speaker C's utterance, and speaker B's second utterance. The control unit 2 can perform speaker separation processing, for example, using an AI model for speaker separation. If the control unit 2 has stored an AI model, it can perform speaker separation processing. Alternatively, the control unit 2 may transmit input voice data to a server on the Internet or a local area network via the communication unit, have the server perform speaker separation processing, and receive the results via the communication unit.
[0024] In step S6, the control unit 2 determines whether the number of speakers in the utterances included in the input voice data or the processed input voice data is one. If the control unit 2 determines in step S6 that there is one speaker, the process proceeds to step S7, where the control unit 2 (identification unit 5) performs a process to vectorize the input voice data or processed input voice data (for example, calculation of Embedding). Then, the process proceeds to step S10. If the control unit 2 determines in step S6 that there are multiple speakers, in step S8 the control unit 2 calculates the total utterance time of each speaker included in the input audio data or processed input audio data. Then, in step S9, the control unit 2 (identification unit 5) performs a process to vectorize the audio data of the speaker with the longest total utterance time included in the input audio data or processed input audio data (for example, calculation of Embedding). After that, the process proceeds to step S10.
[0025] In step S10, the control unit 2 (identification unit 5) compares the speaker information of multiple registered users stored in the memory unit with the vector of input speech data or the vector of processed input speech data generated in step S7 or S9 and calculates the similarity (for example, cosine similarity). The memory unit of the control unit 2 stores speaker information for multiple registered users. For example, the memory unit of the control unit 2 stores the registered user's voice data and its vector along with the registered user's name as speaker information. Speaker information may be stored in the memory unit before implementing the speech recognition method of this embodiment, or it may be stored in the memory unit while repeatedly implementing the speech recognition method of this embodiment.
[0026] Cosine similarity takes values in the range of -1 to 1 and is used to calculate the relationship or similarity between vectorized audio data. With cosine similarity, the smaller the angle between two vectors, that is, the closer the value is to 1, the more similar the two audio data vectors are considered to be. In step S10, for example, the control unit 2 (identification unit 5) compares the input voice data vector or processed input voice data vector generated in step S7 or S9 with the voice data vector of each registered user stored in the storage unit, and calculates the cosine similarity for each registered user.
[0027] In step S11, the control unit 2 (identification unit 5) determines whether the largest cosine similarity value among all cosine similarities calculated in step S10 is greater than threshold B. Threshold B can be the lowest value in the range of cosine similarity in which the speaker of the input voice data or processed input voice data and the registrant can be considered to be the same person. If, in step S11, the control unit 2 (identification unit 5) determines that the largest cosine similarity value among all cosine similarities is greater than threshold B, then in step S12, the control unit 2 (output unit 6) outputs the text data obtained by speech recognition in step S2 to a user interface such as the display unit 11 as the registrant's statement corresponding to the largest cosine similarity. Specifically, both the registrant's name corresponding to the largest cosine similarity and the text data obtained by speech recognition in step S2 can be displayed on the display unit 11. After that, the process returns to step S1. Note that when repeating the flow shown in Figure 1, the control unit 2 can process steps S1 and S2 and steps S3 to S13 in parallel.
[0028] In step S11, if the control unit 2 (identification unit 5) determines that the largest value among all cosine similarities is less than threshold B, in step S13, the control unit 2 (output unit 6) outputs the text data obtained by speech recognition in step S2 as "Speaker name unknown" to the user interface such as the display unit 11. In this case, the user can input the speaker's name to the control unit 2 via an input unit such as a keyboard. If the speaker's name is entered, the control unit 2 corrects the display of "Speaker name unknown" on the user interface to the entered speaker's name. The control unit 2 also stores the input speech data acquired in step S1, the vector of input speech data generated in step S7 or the vector of processed input speech data generated in S9, and the entered name as speaker information in the storage unit. This speaker information can be used in the subsequent step S10.
[0029] Second Embodiment Figure 3 is a flowchart of the speech recognition method according to the second embodiment. The speech recognition method of the second embodiment is the same as the speech recognition method of the first embodiment, except that steps S3 and S4 are omitted, and steps S20, S21, and S22 are performed instead of step S10. In the second embodiment, the control unit 2 uses the input speech data acquired in step S1 to perform steps S5, S7, S9, etc., and uses the cosine similarity calculated in step S21 or S22 to perform steps S11 and S12. In the second embodiment, step S20 is a determination step, and steps S21 and S22 are included in the identification steps. Steps S1, S2, S5-S9, and S11-S13 were explained in the first embodiment and will be omitted here. The explanation will focus on steps S20-S22.
[0030] In the second embodiment, the speaker information for each registered user stored in the storage unit of the control unit 2 includes a vector of the registered user's long voice data, a vector of the registered user's short voice data, and the registered user's name. The long voice data is, for example, voice data whose utterance time is longer than threshold C, and the short voice data is, for example, voice data whose utterance time is shorter than threshold C. The speaker information may be stored in the storage unit before implementing the speech recognition method of this embodiment, or the speaker information may be stored in the storage unit while repeatedly implementing the speech recognition method of this embodiment.
[0031] When the process proceeds from step S7 or S9 to step S20, the control unit 2 (determination unit 4) determines whether the speech duration of the input voice data acquired in step S1 is longer than threshold C. Threshold C can be, for example, any speech duration that separates relatively short speech duration from relatively long speech duration. If the control unit 2 (determination unit 4) determines in step S20 that the speech time is longer than threshold C, the process proceeds to step S21, where the control unit 2 (identification unit 5) compares the input speech data vector generated in step S7 or S9 with the long speech data vectors of each registrant stored in the memory unit, and calculates the cosine similarity for each registrant. The process then proceeds to step S11, where the cosine similarity calculated in step S21 is used to make a decision.
[0032] If the control unit 2 (determination unit 4) determines in step S20 that the speech time is shorter than threshold C, the process proceeds to step S22. The control unit 2 (identification unit 5) compares the input speech data vector generated in step S7 or S9 with the short speech data vectors of each registrant stored in the memory unit, and calculates the cosine similarity for each registrant. The process then proceeds to step S11, where the cosine similarity calculated in step S22 is used for the decision. By comparing the input audio data with registered users' audio data of varying lengths based on the speech duration of the input audio data, the speaker identification accuracy can be improved. Furthermore, the description of the first embodiment described above also applies to the second embodiment, unless otherwise contradictory.
[0033] Third Embodiment Figure 4 is a flowchart of the speech recognition method according to the third embodiment. The speech recognition method of the third embodiment is the same as the speech recognition method of the first embodiment, except that steps S3 and S4 are omitted and steps S30 to S35 are performed. In addition, in the third embodiment, steps S5, S7, S9, etc. are performed using the input speech data acquired in step S1. In the third embodiment, step S30 is a determination step, and steps S31 to S33 are included in the identification step. Steps S1, S2, and S5-S13 were explained in the first embodiment and will be omitted here. The explanation will focus on steps S30-S35.
[0034] After performing speech recognition in step S2, the process proceeds to step S30, where the control unit 2 (determination unit 4) determines whether the speech duration of the input speech data acquired in step S1 is longer than threshold D. Threshold D can be set to the shortest speech duration required for accurate speaker recognition. If the control unit 2 determines in step S30 that the utterance time is longer than threshold D, the process proceeds to step S5, and the control unit 2 performs speaker separation. If the control unit 2 determines in step S30 that the speech time is shorter than threshold D, the process proceeds to step S31.
[0035] In step S31, the control unit 2 (identification unit 5) concatenates the input voice data acquired in step S1 with the voice data of multiple registered users included in the speaker information stored in the memory unit to create processed input voice data. For example, if the speech duration of the input voice data acquired in step S1 is 2 seconds, and the memory unit stores 10 seconds of voice data for registered user A, 10 seconds of voice data for registered user B, and 10 seconds of voice data for registered user C as speaker information, the control unit 2 concatenates the input voice data with the voice data of registered user A, the voice data of registered user B, and the voice data of registered user C to create 32 seconds of processed input voice data.
[0036] In step S32, the control unit 2 (identification unit 5) performs speaker separation processing on the processed input voice data created in step S31, and in step S33, the control unit 2 (identification unit 5) determines whether or not there is voice data separated as the same speaker as the input voice data. In the speaker separation process, the speech segments of each speaker included in the processed input audio data are detected. Since the processed input audio data contains audio data from multiple registered users, it is expected that the audio data of each registered user included in the processed input audio data will be detected as speech segments of different speakers. Furthermore, if the speaker of the input audio data is a registered user, it is expected that the input audio data included in the processed input audio data will be detected as speech segments of the same speaker as that registered user. For example, if the control unit 2 performs speaker separation processing on the 32 seconds of processed input audio data mentioned above, it is possible that the input audio data included in the processed input audio data will be detected as speech segments of the same speaker as one of the registered users A, B, or C. In this case, the process proceeds to step S34, where the control unit 2 (output unit 6) outputs the text data obtained by speech recognition in step S2 to a user interface such as the display unit 11, as the statement of the registered user whose speech segment was detected as being the same speaker as the speech segment of the input audio data. After that, the process returns to step S1.
[0037] Furthermore, if the speaker of the input audio data is not a registered user, in the speaker separation step S32, the input audio data included in the processed input audio data is likely to be detected as speech segments from different speakers among the multiple registered users. In this case, the speaker of the input audio data is likely not one of these registered users. For example, if the control unit 2 performs speaker separation processing on the 32 seconds of processed input audio data mentioned above, the input audio data included in the processed input audio data may be detected as speech segments from a speaker different from registered users A, B, and C. In this case, the process proceeds to step S35, and the control unit 2 (output unit 6) outputs the text data obtained by speech recognition in step S2 as "Speaker name unknown" to a user interface such as the display unit 11. In this case, the user can input the speaker's name to the control unit 2 via an input unit such as a keyboard. After that, the process returns to step S1. Furthermore, the description of the first embodiment described above also applies to the third embodiment, unless otherwise contradictory.
[0038] Fourth Embodiment Figure 5 is a flowchart of the speech recognition method according to the fourth embodiment. The speech recognition method of the fourth embodiment is the same as the speech recognition method of the first embodiment, except that steps S3 and S4 are omitted and steps S40 to S42 are performed instead of step S12. In addition, in the fourth embodiment, steps S5, S7, S9, etc. are performed using the input speech data acquired in step S1. In the fourth embodiment, step S40 is a determination step, and steps S41 and S42 are included in the output steps. Steps S1, S2, S5-S11, and S13 were explained in the first embodiment and will be omitted here. The explanation will focus on steps S40-S42.
[0039] If, in step S11, the control unit 2 (identification unit 5) determines that the largest value among all cosine similarities is greater than threshold B, then in step S40, the control unit 2 (determination unit 4) determines whether the speech duration of the input speech data acquired in step S1 is longer than threshold E. If the control unit 2 determines in step S40 that the utterance time is longer than threshold E, the process proceeds to step S41, where the control unit 2 (output unit 6) outputs the text data obtained by speech recognition in step S2 to the user interface, such as the display unit 11, as the utterance of the registrant corresponding to the highest cosine similarity. The control unit 2 also outputs to the user interface that it has a high degree of confidence that the utterance is that of the registrant. Specifically, the display unit 11 can display both the registrant's name corresponding to the highest cosine similarity and the text data obtained by speech recognition in step S2. In this case, the control unit 2 can display a high degree of confidence along with the registrant's name on the display unit 11. For example, the control unit 2 may choose not to display a question mark (?) along with the registrant's name. The process then returns to step S1.
[0040] If the control unit 2 determines in step S40 that the utterance time is shorter than the threshold E, the process proceeds to step S42. The control unit 2 (output unit) outputs the text data obtained by speech recognition in step S2 to the user interface, such as the display unit 11, as the utterance of the registrant corresponding to the highest cosine similarity. The control unit 2 also outputs to the user interface that the confidence level of the utterance being that of the registrant is low. Specifically, the display unit 11 can display both the registrant's name corresponding to the highest cosine similarity and the text data obtained by speech recognition in step S2. In this case, the control unit 2 can display the low confidence level along with the registrant's name on the display unit 11. For example, the control unit 2 can display a question mark (?) along with the registrant's name. The process then returns to step S1. Furthermore, the description of the first embodiment described above also applies to the fourth embodiment, unless otherwise contradictory. [Explanation of Symbols]
[0041] 2: Control Unit 3: Input Voice Data Acquisition Unit 4: Determination Unit 5: Identification Unit 6: Output Unit 10: Microphone 11: Display Unit 20: Voice Recognition Device
Claims
1. An input audio data acquisition step to acquire input audio data, A determination step of determining whether the duration of the input audio data is longer than a threshold, The process includes an identification step of identifying which registered registrant the input audio data belongs to by comparing the input audio data or the processed input audio data generated by processing the input audio data with the registered audio data of a registered registrant's speech that has been registered in advance. A speech recognition method characterized in that, if the duration of the input speech data is determined to be shorter than the threshold in the determination step, processing is performed in the identification step to improve identifiability.
2. If the determination step determines that the duration of the input audio data is shorter than the threshold, the system includes a first processed input audio data generation step which generates a first processed input audio data by concatenating multiple copies of the input audio data. The speech recognition method according to claim 1, wherein in the identification step, the first processed input speech data is compared with the registered speech data to identify which registered user the input speech data belongs to.
3. If, in the determination step, it is determined that the duration of the input audio data is shorter than the threshold, then in the identification step, the input audio data or the processed input audio data is compared with a first registered audio data of a registered user's utterance that has been registered in advance, to identify which registered user's utterance the input audio data belongs to. The speech recognition method according to claim 1, wherein, in the determination step, it is determined that the duration of the input speech data is longer than the threshold, and in the identification step, the input speech data or the processed input speech data is compared with a second registered speech data of a registered registrant's utterance that has been registered in advance, thereby identifying which registered registrant the input speech data belongs to.
4. If the determination step determines that the duration of the input audio data is shorter than the threshold, the system includes a second processed input audio data generation step which generates a second processed input audio data by concatenating the input audio data and the registered audio data. The speech recognition method according to claim 1, wherein in the identification step, the second processed input speech data is compared with the registered speech data to identify which registered user the input speech data is spoken by.
5. The speech recognition method according to claim 4, wherein in the identification step, the second processed input speech data is compared with the registered speech data to identify which registered speaker the second processed input speech data belongs to for each of a plurality of speech segments with different speakers, and the registered speaker identified as the speaker of the speech segment corresponding to the input speech data is identified as the speaker of the input speech data.
6. A voice data acquisition step to acquire input voice data, A determination step of determining whether the duration of the input audio data is longer than a threshold, An identification step to identify which registered user the input audio data belongs to by comparing the input audio data or the processed input audio data generated by processing the input audio data with the registered audio data of a registered user's utterance that has been registered in advance; The process includes an output step which outputs the result of identifying which registered user the input voice data belongs to, based on the identification step, If, in the determination step, it is determined that the duration of the input audio data is shorter than the threshold, then in the output step, it is output that the speaker's confidence level is low. The speech recognition method according to claim 1, wherein, in the determination step, it is determined that the duration of the input speech data is longer than the threshold, and in the output step, it outputs that the speaker has a high degree of confidence.
7. The speech recognition method according to any one of claims 1 to 6, wherein in the identification step, the input speech data is identified as the utterance of a pre-registered registrant by comparing a numerical vector obtained by converting the features of the input speech data or the processed input speech data with a numerical vector obtained by converting the registered speech data.
8. An input audio data acquisition unit that acquires input audio data, A determination unit that determines whether the duration of the input audio data is longer than a threshold, The system includes an identification unit that identifies which registered user the input audio data belongs to by comparing the input audio data or the processed input audio data generated by processing the input audio data with the registered audio data of a registered user's utterance that has been registered in advance. The speech recognition device is characterized in that, when the determination unit determines that the duration of the input speech data is shorter than the threshold, the identification unit performs a process to improve identifiability.
9. An input audio data acquisition unit that acquires input audio data, A determination unit that determines whether the duration of the input audio data is longer than a threshold, An identification unit identifies which registered user the input audio data belongs to by comparing the input audio data or the processed input audio data generated by processing the input audio data with the registered audio data of a registered user's utterance that has been registered in advance. The identification unit includes an output unit that outputs the result of identifying which registered user the input voice data belongs to, If the determination unit determines that the duration of the input audio data is shorter than the threshold, the output unit outputs that the speaker's confidence level is low. If the determination unit determines that the duration of the input voice data is longer than the threshold, the output unit outputs that the speaker has a high degree of confidence.
Citation Information
Patent Citations
Registered utterance division device, speaker likelihood evaluation device, speaker identification device, registered utterance division method, speaker likelihood evaluation method, and program
JP2017187642A