Device and method
The apparatus and method improve speaker estimation accuracy by using speech information alongside speaker feature vectors for cluster determination, addressing real-time analysis challenges and enhancing precision in speaker identification.
Patent Information
- Application Number
- PCT/JP2024/012463
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-27
- Publication Date
- 2025-10-02
AI Technical Summary
Existing speaker estimation technologies face challenges in achieving accurate real-time analysis and clustering due to limitations in the accuracy of speaker feature vector analysis and parameter registration, which affect the precision of speaker estimation.
An apparatus and method that enhance speaker estimation accuracy by extracting speaker feature vectors and audio information from divided speech segments, using a cluster estimation unit to determine clusters based on both speaker feature vectors and audio information, and updating the cluster space as necessary based on speech content and volume.
Improves the accuracy of speaker estimation by utilizing speech information in addition to speaker feature vectors for cluster estimation, resulting in enhanced precision and quality of speaker identification.
Smart Images

Figure JP2024012463_02102025_PF_FP_ABST
Abstract
Description
Apparatus and method
[0001] The present invention relates to an apparatus and a method.
[0002] There are known techniques for separating speakers or the number of speakers from audio recorded with multiple utterances. Patent Document 1 describes a technique for estimating the number of speakers by performing principal component analysis on speaker feature vectors extracted for each speech segment. Patent Document 2 describes a technique for clustering while estimating the number of speakers using parameters calculated from the number of utterances and utterance lengths.
[0003] JP 2018-063313 A JP 2022-133120 A
[0004] In the technology of Patent Document 1, the estimation accuracy can be improved as the number of speaker feature vectors subjected to principal component analysis increases, but it is not suitable for real-time analysis. In the technology of Patent Document 2, the accuracy of clustering can depend on the accuracy of the parameters registered in the table. Therefore, it is desired to improve the accuracy of speaker estimation.
[0005] An object of the present disclosure is to provide an apparatus and method that can improve the accuracy of speaker estimation.
[0006] An apparatus according to one aspect of the present disclosure includes a vector extraction unit that extracts a speaker feature vector based on divided audio, which is audio divided into speech segments from input audio recorded from multiple people; an audio information extraction unit that extracts audio information, which is a physical quantity related to audio, based on the divided audio; and a cluster estimation unit that estimates a cluster to which the divided audio belongs based on the speaker feature vector and the audio information, and outputs a speaker estimation result based on the cluster estimation result.
[0007] In an apparatus according to one aspect of the present disclosure, a cluster to which a segmented speech belongs is estimated based on a speaker feature vector and speech information extracted from the segmented speech. By using speech information in addition to the speaker feature vector to estimate the cluster, the accuracy of clustering of the segmented speech is improved. As a result, the accuracy of speaker estimation is improved.
[0008] According to the present disclosure, it is possible to improve the accuracy of speaker estimation.
[0009] Fig. 1 is a block diagram showing an example of the configuration of the device. Fig. 2 is a diagram showing an example of a cluster space. Fig. 3 is a flowchart showing an example of the operation of the device. Fig. 4 is a flowchart showing an example of a cluster estimation process. Fig. 5 is a flowchart showing an example of a cluster space update determination process. Fig. 6 is a flowchart showing an example of a cluster space update process. Fig. 7 is a diagram showing an example of the hardware configuration of the device.
[0010] The present disclosure will be described with reference to the accompanying drawings. Whenever possible, the same parts are designated by the same reference numerals and redundant description will be omitted.
[0011] 1 is a block diagram showing an example of the configuration of a device 10. The device 10 is, for example, a server device that functions as a speaker estimation device. The type of the device 10 is not limited. For example, the device 10 may be a terminal such as a personal computer, a high-function mobile phone (smartphone), a tablet terminal, or a wearable terminal.
[0012] The device 10 includes, as functional elements, an input unit 11, a speech section extraction unit 12, a vector extraction unit 13, a speech content extraction unit 14, a voice information extraction unit 15, a cluster estimation unit 17, an update determination unit 18, and an update unit 19. The device 10 also includes a storage unit 16.
[0013] The input unit 11 acquires input speech recorded from multiple people. The input speech is, for example, data recording speech from a meeting, sales, customer service, etc. The input unit 11 may acquire the input speech in real time, or may acquire the input speech stored in a predetermined storage device, etc.
[0014] The speech interval extraction unit 12 extracts speech intervals based on the input speech. An utterance interval is a section of the input speech where a person is speaking. The speech interval extraction unit 12 extracts divided speech, which is speech divided into each utterance interval. The speech interval extraction unit 12 may extract time information of the division points for the input speech as the divided speech.
[0015] The vector extraction unit 13 extracts a speaker feature vector based on the divided speech. The speaker feature vector is a vector that quantifies the features of a speaker's speech. Examples of speech features include, but are not limited to, Mel-frequency cepstrum coefficients. The vector extraction unit 13 extracts a speaker feature vector for each speech section.
[0016] The speech content extraction unit 14 extracts speech content based on the divided speech. The speech content is the content or meaning of the speech. The speech content extraction unit 14 extracts the speech content for each speech section. For example, the speech content extraction unit 14 performs speech recognition processing on the divided speech to extract the speech recognition result, which is text information, as the speech content.
[0017] The audio information extraction unit 15 extracts audio information based on the divided audio. Audio information is a physical quantity related to audio. The audio information extraction unit 15 extracts audio information for each speech section. Examples of audio information include, but are not limited to, speech duration, speech volume, noise level, and speaking speed (speech rate). The audio information may be, for example, at least one of speech duration, speech volume, noise level, and speaking rate. In one example, the audio information may be speech volume and noise level.
[0018] The storage unit 16 is a non-transitory storage medium or storage device that stores the cluster space. The cluster space is information about the cluster to which a voice belongs. Each cluster is a collection of similar voice features. For example, if clusters are created so that voices from the same speaker are in the same cluster, each cluster can be considered a collection of voices for each speaker. A cluster space may be created for each input voice, or may be created based on a combination of different input voices. The storage unit 16 may be constructed as a single database or may be a collection of multiple databases. The storage unit 16 may be a component of the device 10, or may be provided in a computer system separate from the device 10.
[0019] The cluster estimation unit 17 estimates a cluster to which the divided speech belongs based on the speaker feature vector and speech information. For example, the cluster estimation unit 17 estimates a cluster by referring to the cluster space in the storage unit 16. The cluster estimation unit 17 outputs a speaker estimation result based on the cluster estimation result.
[0020] The update determination unit 18 determines whether or not the cluster space needs to be updated based on at least one of the voice information and the speech content. The update determination unit 18 filters the voice used to update the cluster space.
[0021] The update unit 19 updates the cluster space based on the speaker feature vector and the speech information. If the update necessity of the cluster space is "necessary", the update unit 19 may update the cluster space.
[0022] FIG. 2 is a diagram showing an example of a cluster space. For example, the storage unit 16 stores cluster IDs, representative vectors, and representative speech information as the cluster space. The cluster ID is an identifier for identifying a cluster (speaker). The cluster ID may be associated with the name of the speaker. The representative vector is a representative value (e.g., an average value) of the speaker feature vector. The representative speech information is a representative value of the speech information.
[0023] The representative voice information may include a representative speech duration [s], a representative speech volume [dB], a representative noise level [dB], and a representative speech rate [mora / s]. The representative speech duration is a representative value (e.g., an average value) of the speech duration. The representative speech volume is a representative value (e.g., an average value) of the speech volume. The representative noise level is a representative value (e.g., an average value) of the noise level. The representative speech rate is a representative value (e.g., an average value) of the speech rate.
[0024] The operation of the device 10 and the method according to the present disclosure (speaker estimation method) will be described with reference to Figures 3 to 6. Figure 3 is a flowchart showing an example of the operation of the device 10.
[0025] In step S1, the input unit 11 of the device 10 acquires input audio recorded from speeches of multiple people. In one example, the input audio may be data obtained by recording audio from a conference using a single microphone. In another example, the input audio may be data obtained by recording audio from a web conference. The input unit 11 may acquire the input audio in real time, or may acquire the input audio stored in a predetermined storage device or the like.
[0026] In step S2, the speech interval extraction unit 12 of the device 10 extracts speech intervals based on the input speech. The speech interval extraction unit 12 extracts divided speech, which is speech divided into each speech interval. The speech interval extraction unit 12 may extract time information of the division points for the input speech as the divided speech. For example, the speech interval extraction unit 12 may extract time information regarding the start of speech (e.g., start time or start point) and time information regarding the end of speech (e.g., end time or end point) for each speech interval. The time information regarding the start of speech and the time information regarding the end of speech may be expressed as the elapsed time from the beginning of the input speech.
[0027] The method for extracting the speech interval is not limited. For example, the speech interval extraction unit 12 may determine whether an input speech is speech or non-speech for each interval obtained by dividing the input speech by a certain frame width or shift width. The speech interval extraction unit 12 may determine whether an input speech is speech or non-speech using the volume [dB] of the speech and a threshold value. The speech interval extraction unit 12 may determine whether an input speech is speech or non-speech using a Gaussian Mixture Model (GMM) or a deep neural network (DNN). The speech interval extraction unit 12 may determine as a speech interval frames that are consecutively determined to be speech.
[0028] In step S3, the vector extraction unit 13 of the device 10 extracts a speaker feature vector based on the segmented speech, and outputs the speaker feature vector of the segmented speech.
[0029] The method for extracting the speaker feature vector is not limited. For example, the vector extraction unit 13 may extract the speaker feature vector using d-vector. The vector extraction unit 13 may construct a DNN for speaker recognition in advance and extract the output of the intermediate layer of the DNN as the speaker feature vector.
[0030] In step S4, the speech content extraction unit 14 of the device 10 extracts the speech content based on the divided speech. For example, the speech content extraction unit 14 performs speech recognition processing on the divided speech to extract the speech recognition result, which is character information, as the speech content.
[0031] The speech recognition method is not limited. For example, the speech content extraction unit 14 may perform speech recognition processing using a hybrid speech recognition technology configured from an acoustic model and a language model, or an end-to-end speech recognition technology.
[0032] In step S5, the speech information extraction unit 15 of the device 10 extracts speech information based on the segmented speech. For example, the speech information extraction unit 15 may extract a speech duration based on the difference between time information regarding the start of speech and time information regarding the end of speech for each speech section. The speech information extraction unit 15 may extract a speech volume based on the sound pressure of the segmented speech. The speech information extraction unit 15 may extract a noise level based on the volume of speech relative to noise (SNR: Signal-to-Noise Ratio) of the segmented speech. The speech information extraction unit 15 may extract a speech rate based on the speech content and speech duration. For example, the speech information extraction unit 15 may extract a speech rate by dividing the number of characters or moras in the speech recognition result by the speech duration.
[0033] In step S6, the cluster estimation unit 17 of the device 10 performs a cluster estimation process. The cluster estimation unit 17 estimates the cluster to which the divided speech belongs based on the speaker feature vector and speech information. Fig. 4 is a flowchart showing an example of the cluster estimation process.
[0034] In step S61, the cluster estimation unit 17 estimates a cluster ID based on the speaker feature vector. For example, the cluster estimation unit 17 may estimate a cluster ID based on the similarity between the speaker feature vector of the segmented audio and a representative vector in the cluster space. In one example, the cluster estimation unit 17 may calculate the cosine similarity between the speaker feature vector of the segmented audio and the representative vector in the cluster space, and determine the cluster ID with the highest cosine similarity as the estimation result. In another example, the cluster estimation unit 17 may determine the similarity between the vectors using audio information as a feature in addition to the speaker feature vector.
[0035] In step S62, the cluster estimation unit 17 assigns processing based on whether the cluster space is empty. If the cluster space is not empty (YES in step S62), the processing proceeds to step S63. If the cluster space is empty (NO in step S62), the processing proceeds to step S67.
[0036] In step S63, the cluster estimation unit 17 allocates processing based on the similarity between the vectors. For example, the cluster estimation unit 17 determines whether the cosine similarity between the vectors satisfies a predetermined threshold criterion. If the similarity between the vectors satisfies the criterion (YES in step S63), the processing proceeds to step S64. If the similarity between the vectors does not satisfy the criterion (NO in step S63), the processing proceeds to step S66.
[0037] In step S64, the cluster estimation unit 17 allocates processing based on a comparison process between pieces of audio information. For example, the cluster estimation unit 17 compares the audio information of the divided audio with the representative audio information in the cluster space to determine whether a predetermined threshold criterion is satisfied.
[0038] The cluster estimation unit 17 may perform the comparison process using the speech duration. For example, the cluster estimation unit 17 may determine whether or not the difference between the speech duration of the divided voices and the representative speech duration in the cluster space satisfies a predetermined threshold criterion.
[0039] The cluster estimation unit 17 may perform the comparison process using the speech volume. For example, the cluster estimation unit 17 may determine whether or not the difference between the speech volume of the divided speech and the representative speech volume in the cluster space satisfies a predetermined threshold criterion.
[0040] The cluster estimation unit 17 may perform the comparison process using the noise level. For example, the cluster estimation unit 17 may determine whether or not the difference between the noise level of the divided sounds and the representative noise level of the cluster space satisfies a predetermined threshold criterion.
[0041] The cluster estimation unit 17 may perform the comparison process using the speaking rate. For example, the cluster estimation unit 17 may determine whether or not the difference between the speaking rate of the divided voices and the representative speaking rate of the cluster space satisfies a predetermined threshold criterion.
[0042] The cluster estimation unit 17 may determine whether the speech information satisfies the criteria comprehensively based on the determination results of multiple comparison processes. If the speech information satisfies the criteria (YES in step S64), the process proceeds to step S65. If the speech information does not satisfy the criteria (NO in step S64), the process proceeds to step S66.
[0043] In step S65, the cluster estimation unit 17 determines the cluster estimation result as a "cluster ID." The "cluster ID" is the cluster ID estimated in step S61.
[0044] In step S66, the cluster estimation unit 17 determines the cluster estimation result as "no cluster to belong to." "No cluster to belong to" indicates that the speaker feature vector of the segmented speech is dissimilar to the representative vector of the cluster space.
[0045] In step S67, the cluster estimation unit 17 determines the cluster estimation result as "no cluster space." "No cluster space" means that no cluster space has been created. For example, if this is the first utterance in the input speech, no cluster space has been created.
[0046] Returning to Fig. 3, in step S7, the update determination unit 18 of the device 10 performs a process of determining whether to update the cluster space. The update determination unit 18 determines whether to update the cluster space based on at least one of the audio information and the speech content. For example, the update determination unit 18 may determine whether to update the cluster space based on at least one of the speech duration, speech volume, noise level, speech rate, and speech content. Fig. 5 is a flowchart showing an example of the process of determining whether to update the cluster space.
[0047] In step S71, the update determination unit 18 allocates processing based on the audio information. For example, the update determination unit 18 determines whether the audio information of the divided audio satisfies a predetermined threshold criterion.
[0048] The update determination unit 18 may perform the determination process based on the speech duration. For example, the update determination unit 18 may determine whether the speech duration of the divided voices satisfies a predetermined threshold.
[0049] The update determination unit 18 may perform the determination process based on the speech volume. For example, the update determination unit 18 may determine whether the speech volume of the divided audio satisfies a predetermined threshold.
[0050] The update determination unit 18 may make a determination based on the noise level. For example, the update determination unit 18 may determine whether the noise level of the divided sounds satisfies a predetermined threshold.
[0051] The update determination unit 18 may perform the determination process based on the speech rate. For example, the update determination unit 18 may determine whether the speech rate of the divided voices satisfies a predetermined threshold value.
[0052] If the audio information satisfies the criteria (YES in step S71), the process proceeds to step S72. If the audio information does not satisfy the criteria (NO in step S71), the process proceeds to step S74.
[0053] In step S72, the update determination unit 18 allocates processing based on the speech content. For example, the update determination unit 18 determines whether the speech content of the divided voices satisfies a predetermined criterion.
[0054] The update determination unit 18 may analyze the speech recognition result to determine whether the speech content satisfies a predetermined criterion. For example, the update determination unit 18 may perform morphological analysis on the speech recognition result to determine whether a sentence contains a predetermined excluded word. Examples of the predetermined excluded word include, but are not limited to, function words, interjections, particles, and fillers.
[0055] If a sentence contains words other than the predetermined excluded words, the update determination unit 18 may determine that the speech content of the divided audio satisfies the predetermined criterion. If a sentence contains only the predetermined excluded words, the update determination unit 18 may determine that the speech content of the divided audio does not satisfy the predetermined criterion.
[0056] If the utterance content satisfies the criteria (YES in step S72), the process proceeds to step S73. If the utterance content does not satisfy the criteria (NO in step S72), the process proceeds to step S74.
[0057] In step S73, the update determination unit 18 determines whether or not the cluster space needs to be updated as "necessary."
[0058] In step S74, the update determination unit 18 determines whether or not the cluster space needs to be updated as "unnecessary."
[0059] In relation to steps S71 and S72, the update determination unit 18 may make a comprehensive determination based on the determination results of the speech duration, speech volume, noise level, speech rate, and speech content (hereinafter referred to as "each item"). For example, the update determination unit 18 may determine that updating the cluster space is "unnecessary" when one or more of the determination results of the items are "unnecessary." The update determination unit 18 may combine the determination results of the items to ultimately determine whether or not updating is necessary. The update determination unit 18 may also assign a weight to each item to ultimately determine whether or not updating is necessary.
[0060] 3, in step S8, the update unit 19 of the device 10 performs a cluster space update process. For example, the update unit 19 performs the cluster space update process based on the speaker feature vector, speech information, cluster estimation results, and whether or not the cluster space needs to be updated. Fig. 6 is a flowchart showing an example of the cluster space update process.
[0061] In step S81, the update unit 19 allocates processing based on whether the cluster space needs to be updated. If the cluster space needs to be updated (YES in step S81), the processing proceeds to step S82. If the cluster space needs to be updated (NO in step S81), the cluster space update processing ends.
[0062] In step S82, the update unit 19 allocates processing based on the cluster estimation result. If the cluster estimation result is a "cluster ID" (YES in step S82), the processing proceeds to step S83. If the cluster estimation result is not a "cluster ID" (NO in step S82), the processing proceeds to step S84. If the cluster estimation result is not a "cluster ID", this means that the cluster estimation result is "no associated cluster" or "no cluster space".
[0063] In step S83, the update unit 19 updates the cluster space. For example, the update unit 19 updates information associated with the "cluster ID" that is the cluster estimation result.
[0064] The update unit 19 may update the representative vector in the cluster space based on the speaker feature vector. For example, the update unit 19 may update the representative vector associated with the same cluster ID by using an arithmetic average of all speaker feature vectors belonging to the same cluster ID.
[0065] The update unit 19 may update the representative voice information in the cluster space based on the voice information. For example, the update unit 19 may update the representative voice information associated with the same cluster ID by using an arithmetic average of all voice information belonging to the same cluster ID.
[0066] The updating unit 19 may update a representative speech length associated with a cluster ID by using an arithmetic average of all speech durations belonging to the same cluster ID. The updating unit 19 may update a representative speech volume associated with a cluster ID by using an arithmetic average of all speech volumes belonging to the same cluster ID. The updating unit 19 may update a representative noise level associated with a cluster ID by using an arithmetic average of all noise levels belonging to the same cluster ID. The updating unit 19 may update a representative speech rate associated with a cluster ID by using an arithmetic average of all speech rates belonging to the same cluster ID.
[0067] In step S84, the update unit 19 creates a new cluster ID. If a cluster space has not been created, the update unit 19 creates a new cluster space. Hereinafter, the newly created cluster ID will be referred to as a "new cluster ID."
[0068] In step S85, the update unit 19 updates the cluster space. For example, the update unit 19 adds information to the cluster space in association with a new cluster ID.
[0069] The update unit 19 may update the representative vector in the cluster space based on the speaker feature vector. For example, the update unit 19 may add the speaker feature vector as a representative vector in association with a new cluster ID.
[0070] The update unit 19 may update the representative voice information in the cluster space based on the voice information. For example, the update unit 19 may add the voice information as representative voice information in association with a new cluster ID.
[0071] The update unit 19 may add the speech duration as a representative speech length in association with the new cluster ID. The update unit 19 may add the speech volume as a representative speech volume in association with the new cluster ID. The update unit 19 may add the noise level as a representative noise level in association with the new cluster ID. The update unit 19 may add the speech rate as a representative speech rate in association with the new cluster ID.
[0072] In step S86, the update unit 19 may update the cluster estimation result to a "new cluster ID" using the new cluster ID.
[0073] Returning to Fig. 3, in step S9, the cluster estimation unit 17 outputs a speaker estimation result based on the cluster estimation result. The cluster estimation unit 17 may display the speaker estimation result on a display device. The cluster estimation unit 17 may transmit the speaker estimation result to another device, etc. The speaker estimation result may change depending on the processing up to step S9.
[0074] If the update necessity of the cluster space is "necessary" and the cluster estimation result is "cluster ID", the cluster estimation unit 17 outputs "cluster ID" as the speaker estimation result.
[0075] If the necessity of updating the cluster space is "necessary" and the cluster estimation result is "no belonging cluster", the cluster estimation unit 17 outputs a "new cluster ID" as the speaker estimation result.
[0076] If the necessity of updating the cluster space is "necessary" and the cluster estimation result is "no cluster space", the cluster estimation unit 17 outputs a "new cluster ID" as the speaker estimation result.
[0077] If the update necessity of the cluster space is "unnecessary" and the cluster estimation result is "cluster ID", the cluster estimation unit 17 outputs "cluster ID" as the speaker estimation result.
[0078] If the necessity of updating the cluster space is "unnecessary" and the cluster estimation result is "no belonging cluster", the cluster estimation unit 17 outputs "speaker unknown" as the speaker estimation result. The cluster estimation unit 17 may output the cluster ID determined to have the highest similarity to the speaker feature vector as the speaker estimation result.
[0079] If the necessity of updating the cluster space is "unnecessary" and the cluster estimation result is "no cluster space", the cluster estimation unit 17 outputs "speaker unknown" as the speaker estimation result. The cluster estimation unit 17 may also create a new cluster space and a new cluster ID, and output the new cluster ID as the speaker estimation result.
[0080] The order of steps S3 to S5 is not limited. For example, steps S3 to S5 may be performed in parallel. In relation to steps S8 to S9, the update unit 19 may output the speaker estimation result. For example, the update unit 19 may output the speaker estimation result last as part of the cluster space update process of step S8. In this case, step S9 may be omitted.
[0081] As described above, the device 10 according to one aspect of the present disclosure includes a vector extraction unit 13 that extracts a speaker feature vector based on divided audio, which is audio divided into speech segments from input audio recorded from multiple people; an audio information extraction unit 15 that extracts audio information, which is a physical quantity related to audio, based on the divided audio; and a cluster estimation unit 17 that estimates the cluster to which the divided audio belongs based on the speaker feature vector and the audio information, and outputs a speaker estimation result based on the cluster estimation result.
[0082] A method according to one aspect of the present disclosure comprises the steps of: extracting a speaker feature vector based on divided audio, which is audio divided into speech sections from an input audio recording of speech from multiple people; extracting audio information, which is a physical quantity related to audio, based on the divided audio; and estimating a cluster to which the divided audio belongs based on the speaker feature vector and the audio information, and outputting a speaker estimation result based on the cluster estimation result.
[0083] In the device 10 and method according to one aspect of the present disclosure, a cluster to which a segmented speech belongs is estimated based on a speaker feature vector and speech information extracted from the segmented speech. By using speech information in addition to the speaker feature vector to estimate the cluster, the accuracy of clustering the segmented speech is improved. As a result, the accuracy of speaker estimation is improved.
[0084] The system further includes an update unit 19 that updates the cluster space based on the speaker feature vector and speech information. By updating the cluster space, the quality of the cluster space is improved, resulting in improved speaker estimation accuracy.
[0085] The device 10 further includes an update determination unit 18 that determines whether or not the cluster space needs to be updated based on the speech information. If the update is necessary, the update unit 19 updates the cluster space. If the update is not necessary, the update unit 19 does not update the cluster space. Whether or not the cluster space needs to be updated is determined based on the speech information. If the speech information is suitable for updating the cluster space, the cluster space is updated, and if the speech information is not suitable for updating the cluster space, the cluster space is not updated. By filtering the update of the cluster space with the speech information, the quality of the cluster space is improved. As a result, the accuracy of speaker estimation is improved.
[0086] The device 10 further includes an utterance content extraction unit 14 that extracts utterance content based on the divided audio, and an update determination unit 18 that determines whether or not the cluster space needs to be updated based on the utterance content. If the update is necessary, the update unit 19 updates the cluster space. If the update is not necessary, the update unit 19 does not update the cluster space. Whether or not the cluster space needs to be updated is determined based on the utterance content. If the utterance content is suitable for updating the cluster space, the cluster space is updated, and if the utterance content is not suitable for updating the cluster space, the cluster space is not updated. By filtering the update of the cluster space with the utterance content, the quality of the cluster space is improved. As a result, the accuracy of speaker estimation is improved.
[0087] The speech information extraction unit 15 extracts the speech volume as speech information based on the sound pressure of the divided speech. By using the speech volume in addition to the speaker feature vector for cluster estimation, the accuracy of speaker estimation is improved.
[0088] The speech information extraction unit 15 extracts the noise level as speech information based on the volume of the speech relative to the noise of the divided speech. By using the noise level in addition to the speaker feature vector for cluster estimation, the accuracy of speaker estimation is improved.
[0089] In multiple utterances by the same speaker, speaker feature vectors are often similar regardless of the content of the utterance. Furthermore, in multiple utterances by the same speaker, there are often no significant differences in the speech information. In the present disclosure, speech information is used in addition to speaker feature vectors to estimate clusters, thereby improving the accuracy of cluster estimation.
[0090] The input speech or the divided speech may include speech that is difficult to estimate as a speaker, such as speech that is difficult to hear. If the cluster space is updated using speech that is difficult to estimate as a speaker, deviation from the original cluster space may occur. As a result, this may cause a decrease in the accuracy of speaker estimation for other speech. In the present disclosure, the speech used to update the cluster space is filtered, thereby maintaining or improving the quality of the cluster space.
[0091] When the cluster ID is associated with the speaker's name, the speaker for each utterance section becomes clear. Furthermore, by combining the speaker estimation results of the present disclosure with natural language processing technology such as speech recognition results or machine translation results, or non-language recognition technology such as emotion recognition technology, it becomes possible to analyze the content of utterances or emotions of each speaker.
[0092] The present disclosure may have the following configuration. [1] A device comprising: a vector extraction unit that extracts a speaker feature vector based on divided speech, which is speech that has been recorded from a plurality of people and divided into speech segments; a speech information extraction unit that extracts speech information that is a physical quantity related to speech, based on the divided speech; and a cluster estimation unit that estimates a cluster to which the divided speech belongs based on the speaker feature vector and the speech information, and outputs a speaker estimation result based on the cluster estimation result. [2] The device described in [1], further comprising: an update unit that updates a cluster space based on the speaker feature vector and the speech information. [3] The device described in [2], further comprising: an update determination unit that determines whether the cluster space needs to be updated based on the speech information, wherein the update unit updates the cluster space if the update is necessary, and the update unit does not update the cluster space if the update is not necessary. [4] The device according to [2] or [3], further comprising: a speech content extraction unit that extracts speech content based on the divided speech; and an update determination unit that determines whether the cluster space needs to be updated based on the speech content, wherein if the update is necessary, the update unit updates the cluster space, and if the update is not necessary, the update unit does not update the cluster space. [5] The device according to any of [1] to [4], wherein the speech information extraction unit extracts speech volume as the speech information based on sound pressure of the divided speech. [6] The device according to any of [1] to [5], wherein the speech information extraction unit extracts noise level as the speech information based on the volume of speech relative to noise in the divided speech. [7] A method comprising the steps of: extracting a speaker feature vector based on divided speech, which is speech divided into speech sections, from an input speech recorded from a plurality of people; extracting speech information, which is a physical quantity related to speech, based on the divided speech; and estimating a cluster to which the divided speech belongs based on the speaker feature vector and the speech information, and outputting the cluster estimation result.
[0093] The block diagrams used to explain the above embodiments show functional blocks. These functional blocks (components) are realized by any combination of hardware and / or software. Furthermore, the method for realizing each functional block is not particularly limited. That is, each functional block may be realized using a single device that is physically or logically coupled, or may be realized using two or more physically or logically separated devices that are directly or indirectly connected (e.g., wired, wireless, etc.) and these multiple devices. The functional block may also be realized by combining software with the single device or multiple devices.
[0094] Functions include, but are not limited to, judgment, determination, assessment, calculation, computation, processing, derivation, investigation, search, confirmation, reception, transmission, output, access, resolution, selection, selection, establishment, comparison, assumption, expectation, consideration, broadcasting, notifying, communicating, forwarding, configuring, reconfiguring, allocating, mapping, and assignment. For example, a functional block (component) that performs transmission is called a transmitting unit or transmitter. As mentioned above, there are no particular limitations on how these functions are implemented.
[0095] For example, the device 10 according to an embodiment of the present disclosure may function as a computer that performs information processing according to the present disclosure. Fig. 7 is a diagram illustrating an example of a hardware configuration of the device 10 according to an embodiment of the present disclosure. The device 10 described above may be physically configured as a computer device including a processor 1001, a memory 1002, a storage 1003, a communication device 1004, an input device 1005, an output device 1006, a bus 1007, etc.
[0096] In the following description, the term "apparatus" can be interpreted as a circuit, a device, a unit, etc. The hardware configuration of the apparatus 10 may be configured to include one or more of the apparatuses shown in the drawings, or may be configured to exclude some of the apparatuses.
[0097] Each function of the device 10 is realized by loading specified software (programs) onto hardware such as the processor 1001 and memory 1002, causing the processor 1001 to perform calculations, control communication via the communication device 1004, and control at least one of reading and writing data in the memory 1002 and storage 1003.
[0098] The processor 1001 controls the entire computer by running, for example, an operating system. The processor 1001 may be configured as a central processing unit (CPU) including an interface with peripheral devices, a control device, an arithmetic unit, a register, etc. For example, each function of the device 10 described above may be realized by the processor 1001.
[0099] The processor 1001 also reads programs (program codes), software modules, data, etc. from at least one of the storage 1003 and the communication device 1004 into the memory 1002 and executes various processes in accordance with these. The programs used are those that cause a computer to execute at least some of the operations described in the above-described embodiments. For example, each function of the device 10 may be implemented by a control program stored in the memory 1002 and running on the processor 1001. While the above-described various processes have been described as being executed by one processor 1001, they may also be executed simultaneously or sequentially by two or more processors 1001. The processor 1001 may be implemented by one or more chips. The programs may also be transmitted from a network via a telecommunications line.
[0100] The memory 1002 is a computer-readable recording medium and may be configured, for example, by at least one of a read-only memory (ROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), a random access memory (RAM), etc. The memory 1002 may also be called a register, a cache, a main memory (primary storage device), etc. The memory 1002 can store executable programs (program codes), software modules, etc. for performing information processing according to an embodiment of the present disclosure.
[0101] Storage 1003 is a computer-readable recording medium and may be, for example, at least one of an optical disk such as a CD-ROM (Compact Disc ROM), a hard disk drive, a flexible disk, a magneto-optical disk (e.g., a compact disk, a digital versatile disk, a Blu-ray® disk), a smart card, a flash memory (e.g., a card, a stick, a key drive), a floppy disk, a magnetic strip, etc. Storage 1003 may also be referred to as an auxiliary storage device. The storage medium provided in device 10 may be, for example, a database, a server, or other appropriate medium including at least one of memory 1002 and storage 1003.
[0102] The communication device 1004 is hardware (transmission / reception device) for communicating between computers via at least one of a wired network and a wireless network, and is also called, for example, a network device, a network controller, a network card, or a communication module.
[0103] The input device 1005 is an input device (e.g., a keyboard, a mouse, a microphone, a switch, a button, a sensor, etc.) that receives input from the outside. The output device 1006 is an output device (e.g., a display, a speaker, an LED lamp, etc.) that outputs to the outside. The input device 1005 and the output device 1006 may be integrated into one device (e.g., a touch panel).
[0104] Furthermore, each device, such as the processor 1001 and the memory 1002, is connected by a bus 1007 for communicating information. The bus 1007 may be configured using a single bus, or may be configured using different buses between each device.
[0105] The device 10 may also be configured to include hardware such as a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a programmable logic device (PLD), or a field programmable gate array (FPGA), and some or all of the functional blocks may be realized by the hardware. For example, the processor 1001 may be implemented using at least one of these pieces of hardware.
[0106] The order of the procedures, sequences, flowcharts, etc. of each aspect / embodiment described in this disclosure may be changed unless it is consistent. For example, the methods described in this disclosure present elements of various steps using an example order, and are not limited to the particular order presented.
[0107] Input and output information may be stored in a specific location (for example, memory) or may be managed using a management table. Input and output information may be overwritten, updated, or added to. Output information may be deleted. Input information may be sent to another device.
[0108] The determination may be made based on a value represented by one bit (0 or 1), a Boolean value (true or false), or a numerical comparison (e.g., comparison with a predetermined value).
[0109] The aspects / embodiments described in this disclosure may be used alone, in combination, or switched depending on the implementation. Notification of predetermined information (e.g., notification that "X is true") is not limited to explicit notification, but may be implicit (e.g., not notifying the predetermined information).
[0110] Although the present disclosure has been described in detail above, it is clear to those skilled in the art that the present disclosure is not limited to the embodiments described herein. The present disclosure can be implemented in modified and altered forms without departing from the spirit and scope of the present disclosure as defined by the claims. Therefore, the description of the present disclosure is intended to be illustrative and does not have any limiting meaning on the present disclosure.
[0111] Software shall be construed broadly to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, threads of execution, procedures, functions, etc., whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise.
[0112] Software, instructions, information, etc. may also be transmitted or received over a transmission medium. For example, if software is transmitted from a website, server, or other remote source using wired technologies (such as coaxial cable, fiber optic cable, twisted pair, Digital Subscriber Line (DSL)), and / or wireless technologies (such as infrared, microwave), then these wired and / or wireless technologies are included within the definition of transmission media.
[0113] As used in this disclosure, the terms "system" and "network" are used interchangeably.
[0114] Furthermore, the information, parameters, etc. described in this disclosure may be expressed using absolute values, may be expressed using relative values from a predetermined value, or may be expressed using other corresponding information.
[0115] As used in this disclosure, the terms "determining" and "determining" may encompass a wide variety of actions. "Determining" and "determining" may include, for example, judging, calculating, computing, processing, deriving, investigating, looking up, searching, inquiring (e.g., searching in a table, database, or other data structure), ascertaining, and the like. "Determining" and "determining" may also include receiving (e.g., receiving information), transmitting (e.g., sending information), input, output, accessing (e.g., accessing data in memory), and the like. Furthermore, "judgment" and "decision" can include regarding resolving, selecting, choosing, establishing, comparing, etc. as having been "judged" or "decided." In other words, "judgment" and "decision" can include regarding some action as having been "judged" or "decided." Furthermore, "judgment (decision)" can be interpreted as "assuming," "expecting," "considering," etc.
[0116] The terms "connected," "coupled," or any variation thereof, refer to any direct or indirect connection or coupling between two or more elements, and may include the presence of one or more intermediate elements between two elements that are "connected" or "coupled" to each other. The coupling or connection between elements may be physical, logical, or a combination thereof. For example, "connected" may be read as "access." As used in this disclosure, two elements may be considered to be "connected" or "coupled" to each other using one or more wires, cables, and / or printed electrical connections, as well as electromagnetic energy having wavelengths in the radio frequency range, microwave range, and optical (both visible and invisible) range, as some non-limiting and non-exhaustive examples.
[0117] As used in this disclosure, the phrase "based on" does not mean "based only on," unless expressly stated otherwise. In other words, the phrase "based on" means both "based only on" and "based at least on."
[0118] As used in this disclosure, any reference to an element using a designation such as "first," "second," etc. does not generally limit the quantity or order of those elements. These designations may be used in this disclosure as a convenient method of distinguishing between two or more elements. Thus, a reference to a first and a second element does not imply that only two elements may be employed or that the first element must in some way precede the second element.
[0119] When the terms "include," "including," and variations thereof are used in this disclosure, these terms are intended to be inclusive, similar to the term "comprising." Furthermore, when the term "or" is used in this disclosure, it is not intended to be an exclusive or.
[0120] In this disclosure, where articles are added by translation, such as a, an, and the in English, the disclosure may include that the nouns following these articles are in the plural form.
[0121] In the present disclosure, the term "A and B are different" may mean "A and B are different from each other." The term may also mean "A and B are each different from C." Terms such as "separate" and "coupled" may also be interpreted in the same way as "different."
[0122] 10...device, 11...input unit, 12...utterance section extraction unit, 13...vector extraction unit, 14...utterance content extraction unit, 15...audio information extraction unit, 16...storage unit, 17...cluster estimation unit, 18...update determination unit, 19...update unit.
Claims
1. An apparatus comprising: a vector extraction unit that extracts a speaker feature vector based on divided audio, which is audio divided into speech segments from input audio recorded from multiple people; a speech information extraction unit that extracts audio information, which is a physical quantity related to audio, based on the divided audio; and a cluster estimation unit that estimates a cluster to which the divided audio belongs based on the speaker feature vector and the audio information, and outputs a speaker estimation result based on the cluster estimation result.
2. The device according to claim 1, further comprising an update unit that updates a cluster space based on the speaker feature vector and the speech information.
3. The device according to claim 2, further comprising an update determination unit that determines whether or not the cluster space needs to be updated based on the audio information, wherein if the update is necessary, the update unit updates the cluster space, and if the update is not necessary, the update unit does not update the cluster space.
4. The device according to claim 2, further comprising: an utterance content extraction unit that extracts utterance content based on the divided audio; and an update determination unit that determines whether or not the cluster space needs to be updated based on the utterance content, wherein if the update is necessary, the update unit updates the cluster space, and if the update is not necessary, the update unit does not update the cluster space.
5. The device according to claim 1, wherein the audio information extraction unit extracts speech volume as the audio information based on the sound pressure of the divided audio.
6. The device according to claim 1, wherein the audio information extraction unit extracts a noise level as the audio information based on the volume of speech relative to noise in the divided audio.
7. A method comprising the steps of: extracting a speaker feature vector based on segmented audio, which is audio divided into speech segments from an input audio recording of speech from multiple people; extracting audio information, which is a physical quantity related to audio, based on the segmented audio; and estimating a cluster to which the segmented audio belongs based on the speaker feature vector and the audio information, and outputting the cluster estimation result.
Citation Information
Patent Citations
Voice processing device, voice processing method, and program storage medium
WO2020049687A1
Signal processing device, signal processing method, and program
WO2021246304A1