Information processing device, information processing method, and recording medium
The information processing apparatus improves speaker verification accuracy by combining long and short input voices to generate a combined voice for re-determination, effectively addressing the challenge of short input voices in existing systems.
Patent Information
- Application Number
- PCT/JP2023/041602
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-20
- Publication Date
- 2025-05-30
AI Technical Summary
Existing speaker verification systems face challenges in accurately distinguishing between registered user voices and other voices, particularly when input voices are short, leading to increased error rates.
The proposed solution involves an information processing apparatus that acquires a first input voice longer than a predetermined value and a second input voice shorter than the predetermined value, combines them to generate a combined voice, and re-determines whether the second input voice belongs to the registered user by comparing the combined voice with the registered voice.
This approach enhances the accuracy of speaker verification even when input voices are short, reducing the equal error rate and correctly re-determining misclassified voices as belonging to the registered user.
Smart Images

Figure JP2023041602_30052025_PF_FP_ABST
Abstract
Description
Information processing device, information processing method, and recording medium
[0001] The present disclosure relates to the technical fields of an information processing device, an information processing method, and a recording medium.
[0002] There are known devices that perform speaker verification by comparing registered speech with input speech. For example, Patent Document 1 discloses a technique that determines that the speaker of the registered speech data matches the speaker of the input speech data when the distance between the registered speech data and the input speech data is smaller than a verification threshold.
[0003] Japanese Patent Application Laid-Open No. 2015-055835
[0004] An object of this disclosure is to provide an information processing device, an information processing method, and a recording medium that aim to improve upon the techniques disclosed in prior art documents.
[0005] One aspect of the information processing device disclosed herein includes a first acquisition means for acquiring a first input voice that is longer than a predetermined value and is determined to be the voice of the registered user in speaker verification that determines whether an input voice is the voice of the registered user corresponding to the registered voice, a second acquisition means for acquiring a second input voice that is shorter than the predetermined value and is determined not to be the voice of the registered user in the speaker verification, a combination means for combining the first input voice and the second input voice to generate a combined voice, and a re-determination means for re-determining whether the second input voice is the voice of the registered user based on the result of comparing the combined voice with the registered voice.
[0006] One aspect of the information processing method disclosed herein is a method in which at least one computer performs speaker verification to determine whether an input voice is the voice of a registered user corresponding to a registered voice, by acquiring a first input voice that is longer than a predetermined value and is determined to be the voice of the registered user, acquiring a second input voice that is shorter than the predetermined value and is determined not to be the voice of the registered user in the speaker verification, combining the first input voice and the second input voice to generate a combined voice, and re-determining whether the second input voice is the voice of the registered user based on the result of comparing the combined voice with the registered voice.
[0007] One aspect of the recording medium of this disclosure is a computer program recorded on at least one computer that causes the computer to execute an information processing method, which includes acquiring a first input voice that is longer than a predetermined value and is determined to be the voice of the registered user in speaker verification that determines whether an input voice is the voice of a registered user corresponding to a registered voice, acquiring a second input voice that is shorter than the predetermined value and is determined not to be the voice of the registered user in the speaker verification, combining the first input voice and the second input voice to generate a combined voice, and re-determining whether the second input voice is the voice of the registered user based on the result of comparing the combined voice with the registered voice.
[0008] 1 is a block diagram showing the hardware configuration of a first information processing device. FIG. 2 is a block diagram showing the functional configuration of the first information processing device. FIG. 3 is a flowchart showing the operation flow of the first information processing device. FIG. 4 is a map showing the relationship between the length of input speech and the similarity calculated from the input speech. FIG. 5 is a table showing the relationship between the length of input speech and the error rate. FIG. 6 is a block diagram showing the configuration of a second information processing device. FIG. 7 is a flowchart showing the operation flow of the second information processing device. FIG. 8 is a flowchart showing the operation flow of a third information processing device. FIG. 9 is a chart showing an example of operation of the third information processing device. FIG. 10 is a flowchart showing the operation flow of a fourth information processing device. FIG. 11 is a chart showing an example of operation of the fourth information processing device.
[0009] Hereinafter, embodiments of an information processing device, an information processing method, and a recording medium will be described with reference to the drawings.
[0010] First Embodiment A first information processing apparatus will be described with reference to FIGS. 1 to 5. FIG.
[0011] (Hardware Configuration) First, the hardware configuration of the first information processing apparatus will be described with reference to Fig. 1. Fig. 1 is a block diagram showing the hardware configuration of the first information processing apparatus.
[0012] 1, a first information processing device 10 includes a processor 11, a RAM (Random Access Memory) 12, a ROM (Read Only Memory) 13, and a storage device 14. The information processing device 10 may further include an input device 15 and an output device 16. The processor 11, RAM 12, ROM 13, storage device 14, input device 15, and output device 16 are connected to each other via a data bus 17. The data bus 17 may be an interface other than a data bus (for example, a LAN, a USB, etc.).
[0013] The processor 11 loads a computer program. For example, the processor 11 is configured to load a computer program stored in at least one of the RAM 12, the ROM 13, and the storage device 14. Alternatively, the processor 11 may load a computer program stored in a computer-readable storage medium using a storage medium reading device (not shown). The processor 11 may acquire (i.e., load) the computer program from a device (not shown) located outside the information processing device 10 via a network interface. The processor 11 controls the RAM 12, the storage device 14, the input device 15, and the output device 16 by executing the loaded computer program. In particular, in this embodiment, when the processor 11 executes the loaded computer program, a functional block for performing speaker verification is realized within the processor 11. In other words, the processor 11 may function as a controller that executes each control in the information processing device 10.
[0014] The processor 11 may be configured as, for example, a central processing unit (CPU), a graphics processing unit (GPU), a field-programmable gate array (FPGA), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), or a quantum processor. The processor 11 may be configured as one of these, or may be configured to use multiple processors in parallel.
[0015] The RAM 12 temporarily stores computer programs executed by the processor 11. The RAM 12 temporarily stores data that the processor 11 temporarily uses while it is executing the computer programs. The RAM 12 may be, for example, a dynamic random access memory (D-RAM) or a static random access memory (SRAM). Alternatively, other types of volatile memory may be used instead of the RAM 12.
[0016] The ROM 13 stores computer programs executed by the processor 11. The ROM 13 may also store fixed data. The ROM 13 may be, for example, a programmable read-only memory (PROM) or an erasable read-only memory (EPROM). Alternatively, other types of non-volatile memory may be used instead of the ROM 13.
[0017] The storage device 14 stores data that the information processing device 10 stores for a long period of time. The storage device 14 may operate as a temporary storage device for the processor 11. The storage device may store computer programs executed by the processor 11. The storage device 14 may include, for example, at least one of a hard disk device, a magneto-optical disk device, an SSD (Solid State Drive), and a disk array device.
[0018] The input device 15 is a device that receives input instructions from a user of the information processing device 10. The input device 15 may include, for example, at least one of a keyboard, a mouse, and a touch panel. The input device 15 may be configured as part of a smartphone, a tablet terminal, an earphone-type terminal, a watch-type terminal, an HMD (Head Mounted Display) terminal, etc. The input device 15 may be, for example, a device that includes a microphone and is capable of voice input.
[0019] The output device 16 is a device that outputs information related to the information processing device 10 to the outside. For example, the output device 16 may be a display device (e.g., a display or digital signage) that can display information related to the information processing device 10. The output device 16 may also be a speaker or the like that can output information related to the information processing device 10 as audio.
[0020] 1 may be configured to be included in a device external to the first information processing device 10. For example, the first information processing device 10 may be configured to include a processor 11, a RAM 12, and a ROM 13, and the other devices, such as a storage device 14, an input device 15, and an output device 16, may be configured as external devices. That is, the first information processing device 10 may be configured as an information processing system including a plurality of different devices. Furthermore, some of the calculation functions of the first information processing device 10 may be realized by an external server, a cloud, or the like.
[0021] (Functional Configuration) Next, the functional configuration of the first information processing device 10 will be described with reference to Fig. 2. Fig. 2 is a block diagram showing the functional configuration of the first information processing device.
[0022] 2, the first information processing device 10 is configured to include, as processing blocks for realizing its functions, a first acquisition unit 110, a second acquisition unit 120, a combination unit 130, and a redetermination unit 140. Note that each of the first acquisition unit 110, the second acquisition unit 120, the combination unit 130, and the redetermination unit 140 may be realized by, for example, the above-mentioned processor 11 (see FIG. 1).
[0023] The first acquisition unit 110 is configured to acquire a first input speech. The first input speech is speech determined to be the speech of a registered user in speaker verification, which determines whether the input speech is the speech of a registered user corresponding to the registered speech, and the length of the first input speech is longer than a predetermined value. The "predetermined value" here is a threshold for classifying the input speech into long speech and short speech. The predetermined value may be set in advance depending on the accuracy of speaker verification. For example, if the accuracy of speaker verification is relatively high, the predetermined value may be set to a short value (e.g., 5 seconds). On the other hand, if the accuracy of speaker verification is low, the predetermined value may be set to a long value (e.g., 10 seconds). The first input speech may be acquired as waveform data. Alternatively, the first input speech may be acquired as features extracted from the waveform data. The first acquisition unit 110 may have a function of extracting features from the first input speech acquired as waveform data.
[0024] The second acquisition unit 120 is configured to acquire a second input speech. The second input speech is speech determined not to be the speech of a registered user in speaker verification, and its length is shorter than a predetermined value. That is, the second input speech is speech determined not to be the speech of a registered user (hereinafter referred to as "other person's speech"), unlike speech determined to be the speech of a registered user like the first input speech (hereinafter referred to as "personal speech"), and is speech determined not to be the speech of a registered user. The second input speech is acquired as speech shorter than the first input speech. The second input speech may be acquired as waveform data. Alternatively, the second input speech may be acquired as features extracted from the waveform data. The second acquisition unit 120 may have a function of extracting features from the second input speech acquired as waveform data.
[0025] The combining unit 130 is configured to be able to generate a combined speech by combining the first input speech acquired by the first acquisition unit 110 and the second input speech acquired by the second acquisition unit 120. The combining unit 130 may generate the combined speech by splicing the first input speech and the second input speech together. For example, the combining unit 130 may generate the combined speech by splicing the second input speech after the first input speech. For example, the combining unit 130 may generate the combined speech by splicing the first input speech after the second input speech. Furthermore, the combining unit 130 may be configured to be able to generate the combined speech by combining a feature extracted from the first input speech and a feature extracted from the second input speech.
[0026] The re-determination unit 140 is configured to be able to re-determine whether the second input speech, which was determined to be another person's voice when acquired, is the voice of the registered user. Specifically, the re-determination unit 140 determines whether the second input speech is the voice of the registered user based on the result of comparing the combined speech with the registered speech. The re-determination unit 140 may compare the combined speech with the registered speech and calculate a score indicating the similarity between the combined speech and the registered speech. If a result is obtained that indicates that the second input speech is the voice of the registered user, the re-determination unit 140 may re-determine the second input speech as the user's own voice. In other words, the re-determination unit 140 may determine that the second input speech, which was determined to be another person's voice in the first determination, was actually the user's own voice (i.e., the first determination was incorrect). More specific examples of the determination made by the re-determination unit 140 will be described in detail in other embodiments described later.
[0027] (Flow of Operation) Next, the flow of operation in the first information processing device 10 will be described with reference to Fig. 3. Fig. 3 is a flowchart showing the flow of operation of the first information processing device.
[0028] 3 , when the operation of the first information processing device 10 starts, the first acquisition unit 110 and the second acquisition unit 120 first acquire a first input voice and a second input voice, respectively (step S101). Note that the first input voice and the second input voice may be acquired at the same time. Alternatively, the first input voice and the second input voice may be acquired at different times. In this case, the first acquisition unit 110 may acquire the first input voice, and then the second acquisition unit 120 may acquire the second input voice, or the first acquisition unit 110 may acquire the first input voice, and then the second acquisition unit 120 may acquire the second input voice.
[0029] Next, the combining unit 130 combines the first input speech acquired by the first acquisition unit 110 and the second input speech acquired by the second acquisition unit 120 to generate a combined speech (step S102). Then, the re-determination unit 140 compares the combined speech with a registered speech (step S103). The registered speech compared with the combined speech here is the registered speech used when it was determined that the second input speech was not the voice of the registered user.
[0030] Next, the re-determination unit 140 determines whether the second input speech is the speech of a registered user based on the comparison result between the combined speech and the registered speech (step S104). The re-determination unit 140 may output the determination result. For example, if the re-determination unit 140 determines that the second input speech is the speech of a registered user, it may output information indicating that the speaker who spoke the second input speech is the registered user. Alternatively, if the re-determination unit 140 determines that the second input speech is not the speech of a registered user, it may output information indicating that the speaker who spoke the second input speech is not the registered user.
[0031] The determination result by the re-determination unit 140 (i.e., the speaker verification result) can be used for various processes. For example, the speaker verification result can be used for detecting criminals, detecting the voice of a speaker in a meeting, and assisting in automatic transcription for each speaker.
[0032] (Technical Effects) Next, technical effects obtained by the first information processing device 10 will be described with reference to FIGS.
[0033] As shown in Figure 4, when comparing input speech with registered speech, the similarity clearly distinguishes between the actual person and others when the input speech is long, but becomes mixed with others when the input speech is short. In other words, the shorter the input speech, the more difficult it becomes to distinguish between the actual person and others.
[0034] As shown in Figure 5, the error rate (EER: Equal Error Rate) of speaker verification increases as the length of the input speech becomes shorter. This is because the false rejection rate (FRR: False Rejection Rate) increases significantly as the input speech becomes shorter. Therefore, if we can prevent false rejection even when the input speech is short, the error rate of speaker verification can be reduced.
[0035] In response to this, the first information processing device 10 performs a re-determination of the short speech determined to be another person's speech (i.e., the second input speech). Specifically, the second input speech is combined with the long speech determined to be the user's speech (i.e., the first input speech), and based on the comparison result of the combined speech, it is again determined whether or not the speech is that of the registered user. In this way, even if the input speech is mistakenly determined to be another person's speech due to its short length, it is possible to correctly determine it as the user's speech by a subsequent re-determination. Therefore, speaker verification can be performed with high accuracy even when the input speech is short.
[0036] Second Embodiment A second information processing device 10 will be described with reference to Figures 6 and 7. The second information processing device 10 differs in some configurations and operations from the first information processing device 10 described above, but other parts may be similar to the first information processing device 10. Therefore, the following will describe in detail the parts that differ from the first embodiment, and will omit explanations of other overlapping parts as appropriate.
[0037] (Device Configuration) First, the configuration of the second information processing device 10 will be described with reference to Fig. 6. Fig. 6 is a block diagram showing the configuration of the second information processing device. Note that in Fig. 6, the same elements as those described in Fig. 2 are denoted by the same reference numerals.
[0038] 6, the second information processing device 10 is configured to include, as processing blocks for realizing its functions, a first acquisition unit 110, a second acquisition unit 120, a combining unit 130, a re-determination unit 140, a voice acquisition unit 210, a determination unit 220, and an output unit 230. That is, the second information processing device 10 further includes, in addition to the configuration of the first information processing device 10 already described (see FIG. 2), a voice acquisition unit 210, a determination unit 220, and an output unit 230. Note that each of the voice acquisition unit 210, the determination unit 220, and the output unit 230 may be realized by, for example, the above-described processor 11 (see FIG. 1).
[0039] The voice acquisition unit 210 is configured to be able to acquire input voice. The voice acquisition unit 210 may acquire, for example, spoken voice as input voice. In this case, the voice acquisition unit 210 may segment the spoken voice and acquire it as multiple input voices. Note that existing technology can be appropriately adopted as a method for segmenting and acquiring the spoken voice. For example, the voice acquisition unit 210 may detect silent intervals or noise to segment the voice.
[0040] The determination unit 220 is configured to be able to perform speaker verification of the input speech acquired by the speech acquisition unit 210. Specifically, the determination unit 220 determines whether the input speech is the speech of a registered user corresponding to the registered speech by comparing the input speech with the registered speech. The determination unit 220 may, for example, calculate a matching score indicating the similarity between the input speech and the registered speech to determine whether the input speech is the speech of the registered user. In this case, the determination unit 220 may determine that the input speech is the speech of the registered user if the matching score indicating the similarity exceeds a predetermined threshold, and may determine that the input speech is not the speech of the registered user if the matching score indicating the similarity is below the predetermined threshold.
[0041] The output unit 230 is configured to be able to output the second input voice in accordance with the determination result of the determination unit 220. Specifically, when voice determined by the determination unit 220 not to be the voice of a registered user includes voice shorter than a predetermined value, the output unit 230 outputs the voice as the second input voice to the second acquisition unit 120. Note that when there is no voice determined by the determination unit 220 not to be the voice of a registered user, or when there is no voice shorter than the predetermined value included in the voice determined not to be the voice of a registered user, the output unit 230 does not need to output the second input voice.
[0042] The output unit 230 may be configured to output the first input speech in accordance with the determination result of the determination unit 220. Specifically, the output unit 230 may output, as the first input speech, a speech that is longer than a predetermined value among the speeches determined by the determination unit 220 to be the speech of a registered user. In this way, the first acquisition unit 110 can appropriately acquire the first input speech to be combined with the second input speech.
[0043] (Operation Flow) Next, the operation flow in the second information processing device 10 will be described with reference to Fig. 7. Fig. 7 is a flowchart showing the operation flow of the second information processing device. Note that in Fig. 7, the same processes as those described in Fig. 3 are denoted by the same reference numerals.
[0044] 7, when the operation of the second information processing device 10 is started, the speech acquisition unit 210 first acquires an input speech (step S201). Then, the determination unit 220 compares the input speech acquired by the speech acquisition unit 210 with a registered speech (step S202). The determination unit 220 performs speaker determination based on the comparison result between the input speech and the registered speech (step S203).
[0045] Next, the output unit 230 determines whether or not the speech determined by the determination unit 220 to be other people's includes speech shorter than a predetermined value (step S204). If the speech determined to be other people's includes no speech shorter than the predetermined value (step S204: NO), the subsequent processing is omitted and the series of operations ends. In this case, the determination result by the determination unit 220 may be output as the speaker verification result as is.
[0046] On the other hand, if the voice determined to be a voice of a registered user includes a voice shorter than the predetermined value (step S204: YES), the output unit 230 outputs the voice to the second acquisition unit 120 as a second input voice (step S205). In this case, the output unit 120 may output a voice longer than the predetermined value among the voices determined to be a voice of a registered user by the determination unit 220 as a first input voice to the first acquisition unit 110.
[0047] Next, the first acquisition unit 110 and the second acquisition unit 120 acquire the speech output by the output unit 230. That is, the first acquisition unit 110 acquires the first input speech, and the second acquisition unit 120 acquires the second input speech (step S101).
[0048] Next, the combining unit 130 combines the first input speech acquired by the first acquisition unit 110 and the second input speech acquired by the second acquisition unit 120 to generate combined speech (step S102). The redetermination unit 140 compares the combined speech with registered speech (step S103). Then, based on the comparison result between the combined speech and the registered speech, the redetermination unit 140 determines whether the second input speech is the speech of the registered user (step S104).
[0049] (Technical Effects) Next, technical effects obtained by the second information processing device 10 will be described.
[0050] 6 and 7, in the second information processing device 10, the speaker is determined by comparing the input voice with the registered voice, and the second input voice is output based on the result. In this way, the re-determination unit 140 re-determines the voice (i.e., the second input voice) that the determination unit 220 erroneously determined to be a different person. As a result, speaker verification can be performed with higher accuracy.
[0051] <Third Embodiment> A third information processing device 10 will be described with reference to Figures 8 and 9. Note that the third information processing device 10 differs in some of its operations from the first and second information processing devices 10 described above, but other parts may be similar to the first and second information processing devices 10. Therefore, the following will describe in detail the parts that differ from the embodiments already described, and will omit explanations of other overlapping parts as appropriate.
[0052] (Operational Flow) First, the operational flow of the third information processing device 10 will be described with reference to Fig. 8. Fig. 8 is a flowchart showing the operational flow of the third information processing device. Note that in Fig. 8, the same processes as those described in Fig. 3 are denoted by the same reference numerals.
[0053] 8 , when the operation of the third information processing device 10 starts, the first acquisition unit 110 and the second acquisition unit 120 first acquire a first input speech and a second input speech, respectively (step S101). Then, the combining unit 130 combines the first input speech acquired by the first acquisition unit 110 and the second input speech acquired by the second acquisition unit 120 to generate a combined speech (step S102).
[0054] Next, the re-determination unit 140 compares the combined speech with the registered speech to calculate a matching score (step S301). The re-determination unit 140 then compares the calculated matching score between the combined speech and the registered speech with the matching score between the first input speech and the registered speech (step S302). The matching score between the first input speech and the registered speech may be a score calculated in a determination performed before the first input speech is acquired. For example, the matching score between the first input speech and the registered speech may be a matching score calculated by the determination unit 220 (see FIG. 6 ) in the second information processing device 10.
[0055] Next, the re-determination unit 140 re-determines whether the second input voice is the voice of the registered user based on the comparison result between the matching score between the first input voice and the registered voice (hereinafter referred to as the "first matching score") and the matching score between the combined voice and the registered voice (hereinafter referred to as the "second matching score") (step S303).
[0056] For example, the re-determination unit 140 may determine that the second input speech is the voice of the registered user if the second matching score is higher than the first matching score. In this case, the re-determination unit 140 may determine that the second input speech is not the voice of the registered user if the second matching score is lower than the first matching score. The combined speech is a speech obtained by combining the first input speech, which is the user's own voice, with the second input speech. Therefore, if the second input speech is a voice of another person (i.e., if the first determination result is correct), the second matching score is considered to be lower than the first matching score. On the other hand, if the second input speech is the user's own voice (i.e., if the first determination result is incorrect), the second matching score is considered to be higher than the first matching score. In this way, by determining the magnitude relationship between the first matching score and the second matching score, it is possible to appropriately determine whether the second input speech is the voice of the registered user.
[0057] Furthermore, the re-determination unit 140 may determine that the second input speech is the voice of the registered user if the value obtained by subtracting the first matching score from the second matching score is greater than a first predetermined threshold. In this case, the re-determination unit 140 may determine that the second input speech is not the voice of the registered user if the value obtained by subtracting the first matching score from the second matching score is less than the first predetermined threshold. Here, the "first predetermined threshold" is a threshold for determining the extent to which the matching score of the first input speech has changed due to being combined with the second input speech. For example, if the second input speech is a voice of another person (i.e., if the first determination result is correct), the second matching score is expected to decrease significantly. In this case, the value obtained by subtracting the first matching score from the second matching score is a negative value with a large absolute value. On the other hand, if the second input speech is the person's own voice (i.e., if the first determination result is incorrect), the second matching score is expected to increase or only decrease slightly. In this case, the value obtained by subtracting the first matching score from the second matching score is a positive value or a value slightly smaller than 0. Therefore, by setting a first predetermined threshold value that can distinguish between these values, it is possible to appropriately determine whether the second input voice is the voice of a registered user.
[0058] (Operation Example) Next, a specific operation example of the third information processing device 10 will be described with reference to Fig. 9. Fig. 9 is a chart showing an operation example of the third information processing device.
[0059] 9, it is assumed that input speech A and input speech B are input to the third information processing apparatus 10. Note that input speech A is longer than a predetermined value, and input speech B is shorter than the predetermined value.
[0060] Input voice A is compared with registered voices, and a matching score A is calculated. Then, as a result of the determination using score A, input voice A is determined to be the person's own voice. On the other hand, input voice B is also compared with registered voices, and a matching score B is calculated. Then, as a result of the determination using score B, input voice B is determined to be another person's voice.
[0061] Input speech A is determined to be the user's own speech and is longer than a predetermined value, and is therefore acquired as a first input speech by the first acquisition unit 110. On the other hand, input speech B is determined to be another person's speech and is shorter than a predetermined value, and is therefore acquired as a second input speech by the second acquisition unit 120. Then, input speech A and input speech B are combined by the combination unit 130 to form a combined speech.
[0062] The combined speech obtained by combining input speech A and input speech B is compared with the registered speech, and a matching score AB is calculated. The re-determination unit 140 determines whether this score AB is greater than the score A, which is the matching score for input speech A. If the score AB is greater than the score A, input speech B is determined to be the person's own voice.
[0063] (Technical Effects) Next, technical effects obtained by the third information processing apparatus 10 will be described.
[0064] 8 and 9, the third information processing device 10 re-determines whether the second input speech is the voice of a registered user based on the comparison result between the first matching score, which is the matching score between the first input speech and the registered speech, and the second matching score, which is the matching score between the combined speech and the registered speech. In this way, it is possible to easily and accurately re-determine the second input speech by comparing the matching scores.
[0065] <Fourth embodiment> A fourth information processing device 10 will be described with reference to Figures 10 and 11. The fourth information processing device 10 differs in some of its operations from the first to third information processing devices 10 described above, but other parts may be similar to the first to third information processing devices 10. Therefore, the following will describe in detail the parts that differ from the embodiments already described, and will omit explanations of other overlapping parts as appropriate.
[0066] (Flow of Operation) First, the flow of operation in the fourth information processing device 10 will be described with reference to Fig. 10. Fig. 10 is a flowchart showing the flow of operation of the fourth information processing device.
[0067] 10 , when the operation of the fourth information processing device 10 starts, the first acquisition unit 110 acquires a plurality of first input voices, and the second acquisition unit 120 acquires a second input voice (step S401). That is, for one second input voice, a larger number of first input voices are acquired.
[0068] Next, the combining unit 130 combines each of the first input voices acquired by the first acquisition unit 110 with the second input voice acquired by the second acquisition unit 120 to generate multiple combined voices (step S402). As a result, the number of combined voices generated is equal to the number of first input voices acquired by the first acquisition unit 110.
[0069] Next, the re-determination unit 140 compares each of the multiple combined speeches with the registered speech and calculates multiple matching scores (step S403). That is, multiple matching scores corresponding to each of the multiple combined speeches are calculated.
[0070] Next, the re-determination unit 140 calculates a first average score, which is the average of the matching scores corresponding to each of the multiple first input speeches acquired by the first acquisition unit 110 (step S404).The re-determination unit 140 also calculates a second average score, which is the average of the matching scores corresponding to each of the multiple combined speeches (step S405).
[0071] Next, the re-determination unit 140 compares a first average score, which is the average of the matching scores corresponding to each of the first input speeches, with a second average score, which is the average of the matching scores corresponding to each of the combined speeches (step S406).The re-determination unit 140 then re-determines whether the second input speech is the voice of the registered user based on the comparison result between the first average score and the second average score (step S407).
[0072] For example, the re-determination unit 140 may determine that the second input speech is the voice of a registered user if the second average score is higher than the first average score. In this case, the re-determination unit 140 may determine that the second input speech is not the voice of a registered user if the second average score is lower than the first average score. The combined speech is a speech obtained by combining the first input speech, which is the user's own voice, with the second input speech. Therefore, if the second input speech is a voice of another person (i.e., if the first determination result is correct), the second average score is considered to be lower than the first average score. On the other hand, if the second input speech is the user's own voice (i.e., if the first determination result is incorrect), the second average score is considered to be higher than the first average score. In this way, by determining the magnitude relationship between the first average score and the second average score, it is possible to appropriately determine whether the second input speech is the voice of a registered user.
[0073] Furthermore, the re-determination unit 140 may determine that the second input speech is the voice of a registered user if the value obtained by subtracting the first average score from the second average score is greater than a second predetermined threshold. In this case, the re-determination unit 140 may determine that the second input speech is not the voice of a registered user if the value obtained by subtracting the first average score from the second average score is less than the second predetermined threshold. The "second predetermined threshold" here is a threshold for determining how much the first average value, which is the average of the matching scores of multiple first input speeches, has changed due to the combination of the second input speech with the second input speech. For example, if the second input speech is a voice of another person (i.e., if the first determination result is correct), the second average score is expected to decrease significantly. In this case, the value obtained by subtracting the first average score from the second average score is a negative value with a large absolute value. On the other hand, if the second input speech is the user's own voice (i.e., if the first determination result is incorrect), the second average score is expected to increase or only decrease slightly. In this case, the value obtained by subtracting the first average score from the second average score is a positive value or a value slightly smaller than 0. Therefore, by setting a second predetermined threshold value that can distinguish between these values, it is possible to appropriately determine whether the second input voice is the voice of a registered user.
[0074] (Operation Example) Next, a specific operation example of the fourth information processing device 10 will be described with reference to Fig. 11. Fig. 11 is a chart showing an operation example of the fourth information processing device.
[0075] 11, it is assumed that input speech A, input speech B, and input speech C are input to the fourth information processing device 10. Note that input speech A is a speech longer than a predetermined value, input speech B is a speech shorter than a predetermined value, and input speech C is a speech longer than a predetermined value.
[0076] Input voice A is matched with registered voices, and a matching score A is calculated. Then, as a result of the determination using score A, input voice A is determined to be the person's own voice. Input voice B is also matched with registered voices, and a matching score B is calculated. Then, as a result of the determination using score B, input voice B is determined to be another person's voice. Input voice C is also matched with registered voices, and a matching score C is calculated. Then, as a result of the determination using score C, input voice A is determined to be the person's own voice.
[0077] Input speech A is determined to be the person's own speech and is longer than a predetermined value, and is therefore acquired by the first acquisition unit 110 as the first input speech. Input speech B is determined to be another person's speech and is shorter than a predetermined value, and is therefore acquired by the second acquisition unit 120 as the second input speech. Input speech C is determined to be the person's own speech and is longer than a predetermined value, and is therefore acquired by the first acquisition unit 110 as the first input speech.
[0078] Input voice A, which is the first input voice, and input voice B, which is the second input voice, are combined into a combined voice by the combining unit 130. Similarly, input voice C, which is the first input voice, and input voice B, which is the second input voice, are combined into a combined voice by the combining unit 130.
[0079] The combined speech obtained by combining input speech A and input speech B is compared with registered speech, and a matching score, score AB, is calculated. In addition, the combined speech obtained by combining input speech C and input speech B is also compared with registered speech, and a matching score, score CB, is calculated. The re-determination unit 140 determines whether the average value of scores AB and CB (i.e., (score AB + score CB) / 2) is greater than the average value of score A, which is the matching score for input speech A, and score C, which is the matching score for input speech C (i.e., (score A + score C) / 2). If the average value of scores AB and CB is greater than the average value of scores A and C, input speech B is determined to be the user's voice.
[0080] (Technical Effects) Next, technical effects obtained by the fourth information processing apparatus 10 will be described.
[0081] 10 and 11 , in the fourth information processing device 10, multiple combined voices are generated from multiple first input voices, and based on the results of comparing the multiple combined voices with the registered voice, it is determined again whether the second input voice is the voice of the registered user. In this way, multiple first input voices are taken into consideration, so it is possible to perform the re-determination with higher accuracy than when performing the re-determination using only one combined voice. Furthermore, when multiple combined voices are used, it is possible to perform a more appropriate comparison by calculating and using the average matching score.
[0082] The scope of each embodiment also includes a processing method in which a program that operates the configuration of each embodiment to realize the functions of the above-described embodiments is recorded on a recording medium, the program recorded on the recording medium is read as code, and the program is executed on a computer. In other words, a computer-readable recording medium is also included in the scope of each embodiment. Furthermore, each embodiment includes not only a recording medium on which the above-described program is recorded, but also the program itself.
[0083] Examples of recording media that can be used include floppy disks, hard disks, optical disks, magneto-optical disks, CD-ROMs, magnetic tapes, non-volatile memory cards, and ROMs. Furthermore, the scope of each embodiment is not limited to programs that execute processes by themselves, but also includes programs that execute processes by operating on an OS in conjunction with other software or expansion board functions. Furthermore, the program itself may be stored on a server, and part or all of the program may be downloadable from the server to a user terminal. The program may be provided to the user in, for example, a SaaS (Software as a Service) format.
[0084] <Supplementary Notes> The above-described embodiment may be further described as in the following supplementary notes, but is not limited to the following.
[0085] (Supplementary Note 1) The information processing device described in Supplementary Note 1 is an information processing device that includes: a first acquisition means for acquiring a first input voice that is longer than a predetermined value and is determined to be the voice of the registered user in speaker verification that determines whether an input voice is the voice of the registered user corresponding to the registered voice; a second acquisition means for acquiring a second input voice that is shorter than the predetermined value and is determined not to be the voice of the registered user in the speaker verification; a combination means for combining the first input voice and the second input voice to generate a combined voice; and a re-determination means for re-determining whether the second input voice is the voice of the registered user based on a result of comparing the combined voice with the registered voice.
[0086] (Appendix 2) The information processing device described in Appendix 2 is the information processing device described in Appendix 1, further comprising: a determination means for comparing the input voice with the registered voice to determine whether the input voice is the voice of a registered user corresponding to the registered voice; and an output means for, when the voice determined by the determination means to not be the voice of the registered user includes a voice shorter than the predetermined value, outputting the voice shorter than the predetermined value to the second acquisition means as the second input voice.
[0087] (Appendix 3) The information processing device described in Appendix 3 is the information processing device described in Appendix 2, in which the output means outputs, to the first acquisition means, as the first input voice, a voice that is determined by the determination means to be the voice of the registered user and that is longer than the predetermined value.
[0088] (Supplementary Note 4) In the information processing device according to Supplementary Note 4, the re-determination means re-determines whether the second input speech is the voice of the registered user by comparing a first matching score calculated by matching the first input speech with the registered speech and a second matching score calculated by matching the combined speech with the registered speech.
[0089] (Supplementary Note 5) The information processing device described in Supplementary Note 5 is the information processing device described in Supplementary Note 4, in which the re-determination means determines that the second input voice is the voice of the registered user if the second matching score is higher than the first matching score.
[0090] (Appendix 6) The information processing device described in Appendix 6 is the information processing device described in Appendix 4, in which the re-determination means determines that the second input voice is the voice of the registered user when a value obtained by subtracting the first matching score from the second matching score is greater than a first predetermined threshold value.
[0091] (Appendix 7) The information processing device described in Appendix 7 is the information processing device described in Appendix 1 or 2, wherein the first acquisition means acquires a plurality of the first input voices, the combination means combines each of the plurality of first input voices with the second input voice to generate a plurality of the combined voices, and the re-determination means re-determines whether the second input voice is the voice of the registered user based on a result of comparing the plurality of the combined voices with the registered voice.
[0092] (Appendix 8) The information processing device described in Appendix 8 is the information processing device described in Appendix 7, in which the re-determination means re-determines whether the second input voice is the voice of the registered user by comparing a first average score, which is the average value of the matching scores calculated by matching multiple first input voices with the registered voices, with a second average score, which is the average value of the matching scores calculated by matching each of multiple combined voices with the registered voices.
[0093] (Appendix 9) The information processing device described in Appendix 9 is the information processing device described in Appendix 8, in which the re-determination means determines that the second input voice is the voice of the registered user if the second average score is higher than the first average score.
[0094] (Appendix 10) The information processing device described in Appendix 10 is the information processing device described in Appendix 8, in which the re-determination means determines that the second input voice is the voice of the registered user if the value obtained by subtracting the first average score from the second average score is greater than a second predetermined threshold value.
[0095] (Appendix 11) The information processing method described in Appendix 11 is an information processing method in which, by at least one computer, speaker verification is performed to determine whether an input voice is the voice of a registered user corresponding to a registered voice, and a first input voice that is longer than a predetermined value and is determined to be the voice of the registered user is obtained, and a second input voice that is shorter than the predetermined value and is determined not to be the voice of the registered user is obtained in the speaker verification, and the first input voice and the second input voice are combined to generate a combined voice, and whether the second input voice is the voice of the registered user is re-determined based on a result of comparing the combined voice with the registered voice.
[0096] (Appendix 12) The recording medium described in Appendix 12 is a recording medium having recorded thereon a computer program for causing at least one computer to execute an information processing method, the method comprising: acquiring a first input speech longer than a predetermined value that is determined to be the voice of the registered user in speaker verification that determines whether an input speech is the voice of the registered user corresponding to the registered speech; acquiring a second input speech shorter than the predetermined value that is determined not to be the voice of the registered user in the speaker verification; combining the first input speech and the second input speech to generate a combined speech; and re-determining whether the second input speech is the voice of the registered user based on the result of comparing the combined speech with the registered speech.
[0097] (Supplementary Note 13) The computer program described in Supplementary Note 13 is a computer program that causes at least one computer to execute an information processing method, in which speaker verification that determines whether an input voice is the voice of a registered user corresponding to a registered voice acquires a first input voice that is longer than a predetermined value and is determined to be the voice of the registered user, acquires a second input voice that is shorter than the predetermined value and is determined not to be the voice of the registered user in the speaker verification, combines the first input voice and the second input voice to generate a combined voice, and re-determines whether the second input voice is the voice of the registered user based on a result of comparing the combined voice with the registered voice.
[0098] This disclosure may be modified as appropriate within the scope that does not contradict the gist or idea of the invention that can be read from the claims and the entire specification, and information processing devices, information processing methods, and recording media that involve such modifications are also included in the technical idea of this disclosure.
[0099] REFERENCE SIGNS LIST 10 Information processing device 11 Processor 12 RAM 13 ROM 14 Storage device 15 Input device 16 Output device 110 First acquisition unit 120 Second acquisition unit 130 Combination unit 140 Re-determination unit 210 Audio acquisition unit 220 Determination unit 230 Output unit
Claims
1. In a speaker verification for determining whether an input voice is the voice of a registered user corresponding to a registered voice, a first acquisition means for acquiring a first input voice that is longer than a predetermined value determined to be the voice of the registered user; a second acquisition means for acquiring a second input voice that is shorter than the predetermined value determined not to be the voice of the registered user in the speaker verification; a combining means for combining the first input voice and the second input voice to generate a combined voice; and a re-determination means for re-determining whether the second input voice is the voice of the registered user based on a result of comparing the combined voice with the registered voice. An information processing apparatus comprising the above.
2. A determination means for comparing the input voice with the registered voice to determine whether the input voice is the voice of a registered user corresponding to the registered voice; and an output means for outputting, as the second input voice to the second acquisition means, a voice shorter than the predetermined value when the voice determined not to be the voice of the registered user by the determination means includes a voice shorter than the predetermined value. The information processing apparatus according to claim 1, further comprising the above.
3. The output means according to claim 2 outputs, as the first input voice to the first acquisition means, a voice longer than the predetermined value among the voices determined to be the voice of the registered user by the determination means. The information processing apparatus according to claim 2.
4. The re-determination means according to claim 1 or 2 re-determines whether the second input voice is the voice of the registered user by comparing a first verification score calculated by comparing the first input voice with the registered voice and a second verification score calculated by comparing the combined voice with the registered voice. The information processing apparatus according to claim 1 or 2.
5. The re-determination means according to claim 4 determines that the second input voice is the voice of the registered user when the second verification score is higher than the first verification score. The information processing apparatus according to claim 4.
6. The re-determination means according to claim 4 determines that the second input voice is the voice of the registered user when a value obtained by subtracting the first verification score from the second verification score is greater than a first predetermined threshold. The information processing apparatus according to claim 4.
7. The first acquisition means acquires a plurality of the first input voices, the combining means combines each of the plurality of the first input voices with the second input voice to generate a plurality of the combined voices, and the re-determination means re-determines whether the second input voice is the voice of the registered user based on the result of collating the plurality of the combined voices with the registered voice. The information processing apparatus according to claim 1 or 2.
8. The re-determination means re-determines whether the second input voice is the voice of the registered user by comparing a first average score, which is an average value of collation scores calculated by collating the plurality of the first input voices with the registered voice, with a second average score, which is an average value of collation scores calculated by collating each of the plurality of the combined voices with the registered voice. The information processing apparatus according to claim 7.
9. The re-determination means determines that the second input voice is the voice of the registered user when the second average score is higher than the first average score. The information processing apparatus according to claim 8.
10. The re-determination means determines that the second input voice is the voice of the registered user when a value obtained by subtracting the first average score from the second average score is greater than a second predetermined threshold. The information processing apparatus according to claim 8.
11. In speaker verification for determining whether an input voice is the voice of a registered user corresponding to a registered voice by at least one computer, a first input voice longer than a predetermined value determined to be the voice of the registered user in the speaker verification is acquired, a second input voice shorter than the predetermined value determined not to be the voice of the registered user in the speaker verification is acquired, the first input voice and the second input voice are combined to generate a combined voice, and based on the result of collating the combined voice with the registered voice, it is re-determined whether the second input voice is the voice of the registered user. An information processing method.
12. In a speaker verification for determining whether an input voice is the voice of a registered user corresponding to a registered voice in at least one computer, a first input voice longer than a predetermined value determined to be the voice of the registered user is acquired, in the speaker verification, a second input voice shorter than the predetermined value determined not to be the voice of the registered user is acquired, the first input voice and the second input voice are combined to generate a combined voice, and based on a result of comparing the combined voice with the registered voice, it is re-determined whether the second input voice is the voice of the registered user. A recording medium on which a computer program for executing an information processing method is recorded.
Citation Information
Patent Citations
Device and method for collating speech, and storage medium with speech collation processing program stored therein
JP2002023792A