Information processing method, information processing device, and program
Patent Information
- Application Number
- PCT/JP2024/042181
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-25
- Filing Date
- 2024-11-28
- Publication Date
- 2025-10-02
AI Technical Summary
Speaker recognition systems face misjudgments due to score fluctuations in real environments, where conditions affect similarity scores, leading to incorrect determinations of speaker identity.
A method involving at least two score correction processes: QMF processing to adjust scores based on speech quality indicators and CMF processing to account for voice feature variations over time, followed by ASnorm processing to normalize scores based on unspecified majority speaker matches.
Improves speaker verification accuracy by stabilizing scores, reducing calculation load, and enhancing matching performance even in varying environmental conditions.
Smart Images

Figure JP2024042181_02102025_PF_FP_ABST
Abstract
Description
Information processing method, information processing device, and program
[0001] The present disclosure relates to an information processing method, an information processing device, and a program.
[0002] Generally, in speaker recognition, a registered voice of a pre-registered speaker is compared with an input voice of an unknown speaker to be evaluated to calculate a similarity (score), and the score is compared with a pre-set threshold to determine whether the input voice is the voice of the person who claims to be the speaker (speaker verification) or which registered speaker it is (speaker identification).
[0003] Sturim DE and Reynolds DA, Speaker adaptive cohort selection for Tnorm in text-independent speaker verification, ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings, I, art. no.1415220, pp.741-744(2005).Mandasari, MI, Saeidi, R., McLaren, M. and van Leeuwen, DA, Quality measurefunctions for calibration of speaker recognition systems in various duration conditions, IEEE Transactions on Audio, Speech, and Language Processing, 21(11), pp.425-2438(2013).Yu Zheng, Yajun Zhang, Chuanying Niu, Yibin Zhan, Yanhua Long, and Dongxing Xu, Unisound system for voxceleb speaker recognition challenge 2023, eprint arXiv:2308.12526(2023).
[0004] However, in a real environment, the similarity (score) changes depending on the conditions under which the voice is recorded, and so there is a risk of misjudgment (misrecognition) occurring, such as the score falling below the threshold despite the voice being the person's own input, or the score exceeding the threshold despite the voice being someone else's input, due to changes in the conditions. For this reason, there is room for improvement in speaker recognition with regard to appropriate score correction.
[0005] One of the objects of the present disclosure is to appropriately correct speaker recognition scores that may vary in a real environment.
[0006] An information processing method according to the present disclosure is an information processing method executed by at least one processor in an information processing device including at least one processor, the information processing method including: storing enrollment speaker data relating to enrollment speech, which is speech of a registered enrollment speaker; acquiring evaluation speech, which is speech of a speaker to be evaluated; calculating a score indicating the similarity between an enrollment speaker expression vector, which is a feature of the enrollment speech, and an evaluation speaker expression vector, which is a feature of the evaluation speech; correcting the score by combining at least two score correction processes; comparing the corrected score with a predetermined threshold to determine whether the speaker to be evaluated matches the enrollment speaker; and outputting information indicating the determination result. The at least two score correction processes include a first process of correcting the score using a model of interaction between speeches in metadata including quality indicators of the enrollment speech and the evaluation speech; and a second process of correcting a speaker matching score based on a result of matching an unspecified majority speaker speech, which is speech of an unspecified majority speaker, with an unspecified majority speaker expression vector. At least the first process is applied before the second process.
[0007] FIG. 1 is a diagram illustrating an example of the configuration of a speaker verification score correction system according to the first embodiment. FIG. 2 is a diagram illustrating an example of the functional configuration of a speaker verification device according to the first embodiment. FIG. 3 is a diagram illustrating an example of the hardware configuration of an information processing device that realizes the speaker verification device according to the first embodiment. FIG. 4 is a flowchart illustrating an example of the flow of score correction processing executed in the speaker verification device according to the first embodiment. FIG. 5 is a diagram illustrating an example of QMF processing in FIG. 4. FIG. 6 is a diagram illustrating an example of CMF processing in FIG. 4. FIG. 7 is a diagram illustrating an example of ASnorm processing in FIG. 4. FIG. 8 is a diagram illustrating improvement in accuracy of speaker verification by score correction according to the first embodiment. FIG. 9 is a diagram illustrating an example of ASnorm processing according to the second embodiment. FIG. 10 is a diagram illustrating improvement in accuracy of speaker verification by score correction according to the second embodiment. FIG. 11 is a diagram illustrating an example of ASnorm processing according to the third embodiment.
[0008] (1) An information processing method according to the present disclosure is an information processing method executed by at least one processor in an information processing device including at least one processor. The information processing method includes: storing enrollment speaker data related to enrollment speech, which is speech of a registered enrollment speaker; acquiring evaluation speech, which is speech of a speaker to be evaluated; calculating a score indicating the similarity between an enrollment speaker expression vector, which is a feature of the enrollment speech, and an evaluation speaker expression vector, which is a feature of the evaluation speech; correcting the score by combining at least two score correction processes; comparing the corrected score with a predetermined threshold to determine whether the speaker to be evaluated matches the enrollment speaker; and outputting information indicating the determination result. The at least two score correction processes include a first process of correcting the score using a model of interaction between speeches in metadata including quality indicators of the enrollment speech and the evaluation speech; and a second process of correcting a speaker matching score based on a result of matching an unspecified majority speaker speech, which is speech of an unspecified majority speaker, with an unspecified majority speaker expression vector. At least the first process is applied before the second process.
[0009] This configuration improves speaker verification performance by applying QMF processing before ASnorm processing, thereby appropriately correcting speaker recognition scores that may fluctuate in a real environment.
[0010] (2) In the information processing method described in (1) above, the at least two score correction processes further include a third process of correcting the score based on a variation in a speaker expression vector for each time period of each voice.
[0011] According to this configuration, even when CMF processing is further applied, QMF processing is applied at least before ASnorm processing, thereby improving speaker verification performance.
[0012] (3) In the information processing method described in (1) or (2) above, the metadata includes an average value of the scores based on an unspecified multi-speaker expression vector used in the first processing of the unspecified multi-speaker speech.
[0013] According to this configuration, QMF processing is applied prior to ASnorm processing, so that speaker matching performance can be improved even when the score average with unspecified number of speaker data 4 is used as metadata in addition to quality indicators such as utterance length and SN.
[0014] (4) In the information processing method described in (1) or (2) above, the first processing calculates the scores of each of the registered speaker expression vector and the evaluated speaker expression vector and an unspecified majority speaker expression vector used in the first processing, selects the score when the unspecified majority speaker expression vector used in the first processing that has a high score with the registered speaker expression vector and the evaluated speaker expression vector is compared with the evaluated speaker expression vector, selects the score when the unspecified majority speaker expression vector used in the first processing that has a high score with the evaluated speaker expression vector is compared with the registered speaker expression vector, and calculates the average score of the unspecified majority speaker expression vector used in the first processing for each of the registered speaker expression vector and the evaluated speaker expression vector.
[0015] According to this configuration, QMF processing is applied prior to ASnorm processing, so that speaker matching performance can be improved even when the score average with unspecified number of speaker data 4 is used as metadata in addition to quality indicators such as utterance length and SN.
[0016] (5) In the information processing method according to any one of (1) to (4), the second processing includes the first processing to which the registered speaker expression vector and an unspecified majority speaker expression vector used in the second processing are input, and the second processing to which the evaluation speaker expression vector and an unspecified majority speaker expression vector used in the second processing are input, and the scores of each of the registered speaker expression vector and the evaluation speaker expression vector and the unspecified majority speaker expression vector used in the second processing are calculated, and the second processing having a high score with the registered speaker expression vector is selected. The scores of the unspecified majority speaker expression vectors used in the second processing are selected when compared with the evaluation speaker expression vectors, the scores of the unspecified majority speaker expression vectors used in the second processing that have high scores with the evaluation speaker expression vectors are selected when compared with the registered speaker expression vectors, the average value and variance of the scores of the unspecified majority speaker expression vectors used in the second processing for each of the registered speaker expression vectors and the evaluation speaker expression vectors are calculated, and the scores corrected in at least the first processing are normalized using the calculated average value and variance.
[0017] According to this configuration, QMF processing is applied prior to ASnorm processing, so that speaker matching performance can be improved even when the score average with unspecified number of speaker data 4 is used as metadata in addition to quality indicators such as utterance length and SN.
[0018] (6) In the information processing method described in any one of (1) to (4) above, the second processing calculates the scores of the registered speaker expression vectors and unspecified majority speaker expression vectors used in the second processing, calculates the scores of the evaluation speaker expression vectors and unspecified majority speaker expression vectors used in the second processing, selects the scores of the unspecified majority speaker expression vectors used in the second processing that have a high score with the registered speaker expression vector when compared with the evaluation speaker expression vector, selects the scores of the unspecified majority speaker expression vectors used in the second processing that have a high score with the evaluation speaker expression vector when compared with the registered speaker expression vector, calculates the average value and variance of the scores of the unspecified majority speaker expression vectors used in the second processing for the registered speaker expression vectors and the evaluation speaker expression vectors, and normalizes at least the scores corrected in the first processing using the calculated average value and variance.
[0019] According to this configuration, by applying ASnorm processing after at least QMF processing, it is possible to improve matching accuracy while reducing the amount of calculation required for ASnorm processing that is applied after at least QMF processing.
[0020] (7) In the information processing method described in (6) above, the second processing includes the first processing on the calculated mean value and variance value, and normalizes at least the score corrected by the first processing using the mean value and the variance value to which the first processing has been applied.
[0021] With this configuration, whether the mean and variance are calculated after QMF processing and CMF processing or before QMF processing and CMF processing, the calculation results, i.e., the values of the corrected scores, are approximately the same, and the amount of calculation for the ASnorm processing applied after QMF processing can be reduced at least.
[0022] (8) In the information processing method described in any one of (2) to (7) above, which cites at least (2) above, the first processing, the third processing, and the second processing are applied in this order.
[0023] According to this configuration, the QMF processing is applied before the ASnorm processing, thereby improving the performance of speaker verification.
[0024] (9) An information processing device according to the present disclosure includes a memory and at least one processor. The memory stores enrollment speaker data related to enrollment speech, which is speech of a registered enrollment speaker. The at least one processor is configured to acquire evaluation speech, which is speech of a speaker to be evaluated, calculate a score indicating the similarity between an enrollment speaker expression vector, which is a feature of the enrollment speech, and an evaluation speaker expression vector, which is a feature of the evaluation speech, correct the score by combining at least two score correction processes, compare the corrected score with a predetermined threshold to determine whether the speaker to be evaluated matches the enrollment speaker, and output information indicating the determination result. The at least two score correction processes include a first process that corrects the score using a model of interaction between speeches in metadata including quality indicators of the enrollment speech and the evaluation speech, and a second process that corrects a speaker matching score based on a matching result of an unspecified majority speaker speech, which is speech of an unspecified majority speaker, with an unspecified majority speaker expression vector. At least the first process is applied before the second process.
[0025] This configuration improves speaker verification performance by applying QMF processing before ASnorm processing, thereby appropriately correcting speaker recognition scores that may fluctuate in a real environment.
[0026] (10) A program according to the present disclosure causes a computer to store enrollment speaker data relating to an enrollment speech, which is the speech of a registered enrollment speaker; acquire an evaluation speech, which is the speech of a speaker to be evaluated; calculate a score indicating the similarity between an enrollment speaker expression vector, which is a feature of the enrollment speech, and an evaluation speaker expression vector, which is a feature of the evaluation speech; correct the score by combining at least two or more score correction processes; compare the corrected score with a predetermined threshold to determine whether the speaker to be evaluated matches the enrollment speaker; and output information indicating the determination result. The at least two or more score correction processes include a first process of correcting the score using a model of interaction between speeches in metadata including quality indicators of the enrollment speech and the evaluation speech; and a second process of correcting a speaker matching score based on a matching result of an unspecified multi-speaker speech, which is the speech of an unspecified multi-speaker, with an unspecified multi-speaker expression vector. At least the first process is applied before the second process.
[0027] This configuration improves speaker verification performance by applying QMF processing before ASnorm processing, thereby appropriately correcting speaker recognition scores that may fluctuate in a real environment.
[0028] Hereinafter, with reference to the drawings, embodiments of a speaker recognition method (information processing method), a speaker recognition device (information processing device), a program, and a recording medium according to the present disclosure will be described.
[0029] In the description of the present disclosure, components having the same or substantially the same functions as those described above with respect to the previously-mentioned drawings may be given the same reference numerals, and descriptions thereof may be omitted as appropriate. Furthermore, even when the same or substantially the same parts are shown, the dimensions and proportions may be different depending on the drawing. Furthermore, for example, in order to ensure the visibility of the drawings, reference numerals may be given to only the main components in the description of each drawing, and reference numerals may not be given to components having the same or substantially the same functions as those described above with respect to the previously-mentioned drawings.
[0030] In the description of the present disclosure, components having the same or substantially the same functions may be distinguished by adding an alphanumeric character to the end of the reference symbol. Alternatively, when multiple components having the same or substantially the same functions are not distinguished, they may be collectively described by omitting the alphanumeric character at the end of the reference symbol.
[0031] In the following description, an example is given in which the score correction according to the present disclosure is applied to speaker verification as speaker recognition, but the application is not limited to this. The score correction according to the present disclosure may be applied to speaker identification instead of or in addition to speaker verification. In other words, the speaker verification score correction system according to the present disclosure may be realized as a speaker identification score correction system or a speaker recognition score correction system. Similarly, the speaker verification device according to the present disclosure may be realized as a speaker identification device or a speaker recognition device.
[0032] The speaker recognition according to the present disclosure may be speech content-dependent speaker recognition in which the words to be spoken (uttered) are determined in advance, or speech content-independent speaker recognition in which any words may be uttered, or a combination of these.
[0033] The speaker verification score correction system 1 according to the present disclosure may be applied to a moving object such as a vehicle. For example, the speaker recognition device according to the present disclosure may be mounted on a moving object configured to be capable of executing control in response to a user's voice operation, such as for autonomous driving. Here, the moving object may be any type of vehicle, such as an electric vehicle (EV) driven by a motor as a power source, a vehicle such as an automobile driven by an engine (internal combustion engine) as a power source, or a hybrid vehicle driven by both an engine and a motor as a power source. Furthermore, the moving object may be, for example, an automobile (vehicle) such as a passenger car, truck, or motorcycle, but may also be an electric bicycle, an electric kick scooter, an electric wheelchair, construction machinery, agricultural machinery, a ship, a train, an airplane (aircraft), or the like. Furthermore, the moving object may be configured to be capable of autonomous movement, or may be configured to be capable of movement in response to a user's direct or remote operation. Here, "voice operation" of a mobile object may be an operation to control the movement of the mobile object, such as driving operation or setting a destination or route in a navigation system, or an operation to control other functions of the mobile object, such as music playback, video playback, or internet search. Furthermore, the "movement" of a mobile object is realized by autonomous or heteronomous control (driving control), and may be expressed as "driving." Furthermore, a user who operates a mobile object by voice may be a driver, a passenger, or other occupant of the mobile object, or an operator who remotely controls the mobile object from outside the mobile object.
[0034] First Embodiment FIG. 1 is a diagram showing an example of the configuration of a speaker verification score correction system 1 according to a first embodiment.
[0035] The speaker verification score correction system 1 compares the speech of a registered speaker registered in advance (registered speech) with the speech of an unknown speaker to be matched, i.e., an evaluation target speech (evaluation speech), to calculate a similarity, and determines whether the unknown speaker matches the registered speaker based on the similarity. For example, the speaker verification score correction system 1 determines whether the unknown speaker matches the registered speaker by comparing the similarity with a preset threshold. Note that, in this embodiment, a use case in which the speech of the registered speaker is registered in advance as the registration speech is illustrated, but this is not limited thereto. The speaker verification score correction system 1 according to this embodiment may also be applied to a use case in which the registered speaker is not registered "in advance." As an example, when applied to a use case such as taking minutes of a meeting, the registration of the registered speaker can be performed "concurrently" rather than "in advance." For example, the speaker verification score correction system 1 may simultaneously acquire the speech of the registered speaker to be registered and the speech of the unknown speaker to be evaluated, and compare the acquired speeches to calculate a similarity. Furthermore, in this embodiment, a use case in which the speaker to be evaluated is an unknown speaker is illustrated, but this is not limited thereto. The speaker to be evaluated may be a known speaker, for example, previously enrolled or evaluated.
[0036] For example, in a real environment (e.g., a noisy environment), the similarity (score) changes depending on the conditions under which the voice is recorded. Therefore, for example, a change in the conditions may cause the score to fall below the threshold even though the input voice is the person's own. Also, for example, a change in the conditions may cause the score to exceed the threshold even though the input voice is from another person. In other words, in a real environment, the score changes depending on the conditions under which the voice is recorded, which may result in an erroneous determination (mismatch). For these reasons, there has been a demand for a technology that can appropriately correct the speaker matching score, which can change depending on the recording conditions in a real environment, to improve matching accuracy.
[0037] In this context, the speaker verification score correction system 1 according to the embodiment is configured to correct the speaker verification similarity (score), which indicates the similarity between speaker expression vectors (x-vectors) extracted from the speech data to be compared, by combining at least two or more score correction techniques.
[0038] As an example, the speaker verification score correction system 1 according to the embodiment is configured to perform a quality measure function (QMF) process (first process) that adds or subtracts (corrects) a speaker verification score using a QMF model that models the interaction of quality indicators (metadata) such as speech lengths between speeches and signal-to-noise ratios (SN). In other words, the speaker verification score correction system 1 according to the embodiment is configured to perform a QMF process that corrects a speaker verification score using a QMF model that has learned how much to correct the speaker verification score with respect to the quality indicators (metadata). Here, model training refers to determining or updating parameters of at least one function that defines the model. Details of the QMF process will be described later.
[0039] As an example, the speaker verification score correction system 1 according to the embodiment is configured to perform a CMF (Consistency Measure Factor) process (third process) that penalizes (corrects) the speaker verification score based on the variation in the feature values (e.g., voiceprints) of each voice over time. Specifically, the speaker verification score correction system 1 according to the embodiment calculates a CMF value for each voice individually, reflecting the degree of consistency or variance of the feature values over time of the voice, and scales the speaker verification score using the calculated CMF value as a correction coefficient. For example, a larger CMF value indicates a more concentrated distribution of speaker expression vectors, thereby increasing the speaker verification score. Details of the CMF process will be described later.
[0040] As an example, the speaker verification score correction system 1 according to the embodiment is configured to perform adaptive symmetric normalization (ASnorm) processing (second processing) that normalizes (corrects) the speaker verification score based on the verification results with an unspecified number of speakers (imposters). Details of the ASnorm processing will be described later.
[0041] The speaker verification score correction system 1 according to the embodiment then compares the similarity (corrected score) corrected using a combination of at least two or more score correction techniques with a preset threshold to determine (match) whether the unknown speaker matches a registered speaker. The speaker verification score correction system 1 also outputs information indicating the determination result (matching result). For example, if the corrected score is equal to or greater than the preset threshold, the speaker verification score correction system 1 determines that the unknown speaker is a registered speaker (the person in question) and outputs a matching result indicating that the unknown speaker is the registered speaker (the person in question). On the other hand, if the corrected score is less than the preset threshold, the speaker verification score correction system 1 determines that the unknown speaker is a person (other person) other than the registered speaker (the person in question) and outputs a matching result indicating that the unknown speaker is an unregistered speaker (other person).
[0042] As shown in Fig. 1, a speaker verification score correction system 1 includes a speaker verification device 10. Fig. 2 is a diagram showing an example of the functional configuration of the speaker verification device 10 according to the first embodiment. As shown in Fig. 2, the speaker verification device 10 has functions as an input / output unit 101, an execution unit 102, and a storage unit 103.
[0043] The input / output unit 101 acquires the evaluation speech uttered by the unknown speaker and outputs information indicating the determination result (matching result).
[0044] The execution unit 102 performs speaker verification processing, score correction processing, and judgment processing.
[0045] The storage unit 103 stores programs, parameters, data being processed, data resulting from processing, and the like, relating to each process executed by the speaker verification device 10 .
[0046] As an example, the storage unit 103 stores enrolled speaker data 3, unspecified majority speaker (Imposter) data 4, and trained QMF model 5.
[0047] The enrollment speaker data 3 is information relating to the speech of a previously enrolled speaker (enrollment speech). The enrollment speaker data 3 includes enrollment speech 3a, enrollment speaker expression vector 3b, and enrollment metadata 3c. The enrollment speech 3a is the speech of a previously enrolled speaker. The enrollment speaker expression vector 3b is a speaker expression vector (x-vector) of the enrollment speech 3a, i.e., a feature extracted from the enrollment speech 3a. The enrollment metadata 3c is quality indicators (metadata) such as the speech length and signal-to-noise ratio (SN) of the enrollment speech 3a.
[0048] The unspecified multi-speaker data 4 is information relating to the speech of an unspecified multi-speaker (unspecified multi-speaker speech). The unspecified multi-speaker data 4 includes unspecified multi-speaker speech for ASnorm 4a, unspecified multi-speaker expression vectors for QMF 4b, unspecified multi-speaker expression vectors for ASnorm 4c, and unspecified multi-speaker metadata for ASnorm 4d. The unspecified multi-speaker speech for ASnorm 4a is speech of an unspecified multi-speaker that is subjected to ASnorm processing. The unspecified multi-speaker expression vectors for QMF 4b and the unspecified multi-speaker expression vectors for ASnorm 4c are each speaker expression vectors (x-vectors) of the unspecified multi-speaker speech 4a, i.e., feature quantities extracted from the unspecified multi-speaker speech 4a. The unspecified multi-speaker metadata 4d for ASnorm is quality indicators (metadata) such as speech length and signal-to-noise (SN) ratio in the unspecified multi-speaker speech 4a. In the description of the present disclosure, "X" for "A" means at least "X" used for "A," and does not prevent "X" for "A" from being used for "B" other than "A," or "X" for "C" other than "A" from being used for "A." Furthermore, the "X" for "A" and the "X" for "D" other than "A" may be partially or entirely the same.
[0049] The trained QMF model 5 is a machine learning model or at least one function whose parameters are determined to output information indicating the degree of correction depending on input of similarity (score) and metadata, or a corrected score. As this machine learning model, any machine learning model such as a CNN (Convolutional Neural Network) can be used as appropriate depending on the processing. Note that the trained QMF model 5 may be stored outside the speaker verification device 10.
[0050] Note that two or more of the functions of the speaker verification device 10 according to the embodiment may be integrated into one function. Also, some of the functions of the speaker verification device 10 according to the embodiment may be implemented by an information processing device provided outside the speaker verification device 10 in the speaker verification score correction system 1.
[0051] FIG. 3 is a diagram showing an example of the hardware configuration of the information processing device 8 that realizes the speaker verification device 10 according to the first embodiment.
[0052] 2, the information processing device 8 has a processor 81, a main storage device 82, an auxiliary storage device 83, and an I / F (interface) 84. The processor 81, the main storage device 82, the auxiliary storage device 83, and the I / F 84 are interconnected by a bus or the like, and have a hardware configuration using a typical computer. Note that each component of the information processing device 8 may be realized by a combination of two or more components.
[0053] The processor 81 is, for example, at least one CPU (Central Processing Unit). The processor 81 executes a program, for example, to comprehensively control the operation of the information processing device 8 and realize each function of the information processing device 8.
[0054] As an example, in an information processing device 8 that realizes the speaker verification device 10, a processor 81 loads a program stored in an auxiliary storage device 83 into a main storage device 82 and executes it, thereby realizing each function of the speaker verification device 10, including the execution unit 102 illustrated in Figure 2.
[0055] 2 illustrates only the functions necessary for explaining the main parts of this embodiment, but the functions of the speaker verification device 10 are not limited to these. Also, some or all of the functions of the speaker verification device 10 may be realized by a dedicated hardware circuit.
[0056] The processor 81 according to the embodiment is an example of at least one processor in the information processing device 8. As the at least one processor, at least one other processor may be used instead of or in addition to the CPU. As the other processor, various processors such as a CPU, a GPU (Graphics Processing Unit), a DSP (Digital Signal Processor), or a dedicated arithmetic circuit realized by an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array) can be used as appropriate.
[0057] The main storage device 82 is, for example, a RAM (Random Access Memory). The main storage device 82 temporarily stores data necessary for various processes performed by the processor 81. The main storage device 82 according to the embodiment is an example of at least one memory in the information processing device 8.
[0058] The auxiliary storage device 83 is, for example, a read-only memory (ROM). The auxiliary storage device 83 stores programs, parameters, and the like that realize various processes performed by the processor 81. The auxiliary storage device 83 according to the embodiment is an example of at least one memory in the information processing device 8. Note that, as the auxiliary storage device 83, various storage media and storage devices such as a hard disk drive (HDD), a solid state drive (SSD), and a flash memory can be used as appropriate instead of or in addition to the ROM.
[0059] As an example, in the information processing device 8 that realizes the speaker verification device 10 , the main storage device 82 and the auxiliary storage device 83 realize the storage unit 103 .
[0060] The I / F 84 is an interface for implementing input / output functions, connecting external devices, and / or communicating with the outside. The I / F 84 may be an output interface that connects or realizes an output device that outputs audio, images, or video, an input interface that connects or realizes an input device that acquires user input, or an interface that functions as these devices. As output devices, various displays and speakers such as liquid crystal displays (LCDs), organic electroluminescence (EL) displays, and projectors can be used as appropriate. As input devices, keyboards, mice, touch panels, microphones, and the like can be used as appropriate. As interfaces for communicating with the outside, communication circuits for wired or wireless communication can be used as appropriate. As communication circuits for wireless communication, communication circuits compatible with various standards such as 3G, 4G, 5G, 6G, Wi-Fi (registered trademark), Bluetooth (registered trademark), and infrared communication can be used as appropriate.
[0061] As an example, in an information processing device 8 that realizes a speaker verification device 10, the I / F 84 realizes an input interface (input / output unit 101) that connects or realizes an input device, and an output interface (input / output unit 101) that connects or realizes an output device.
[0062] When the speaker verification score correction system 1 according to the embodiment is applied to a vehicle (mobile object), the information processing device 8 that realizes the speaker verification device 10 may be realized by an in-vehicle computer such as an ECU (Electronic Control Unit) provided inside the vehicle, a DCU (Domain Control Unit) such as a CDC (Cockpit Domain Controller) that integrates multiple ECUs, or an OBU (On Board Unit). Alternatively, the information processing device 8 that realizes the speaker verification device 10 may be an external computer installed near the dashboard of the vehicle. Furthermore, the information processing device 8 that realizes the speaker verification device 10 may be realized by an information processing device 8 shared with other in-vehicle devices, or may be realized by different information processing devices 8. For example, the information processing device 8 that realizes the speaker verification device 10 may be integrated with an in-vehicle car navigation device.
[0063] Furthermore, when the speaker verification score correction system 1 according to the embodiment is applied to a vehicle (mobile object), the information processing device 8 realizing the speaker verification device 10 may transmit and receive information to and from another computer installed in the vehicle via an in-vehicle network including a CAN (Controller Area Network), Ethernet (registered trademark), USB (Universal Serial Bus (registered trademark)), etc. within the vehicle, or may communicate with an information processing device outside the vehicle via a network such as the Internet. As an example, the information processing device 8 realizing the speaker verification device 10 outputs the verification result (determination result) of speaker verification to another in-vehicle computer that controls the vehicle.
[0064] An example of the operation of the speaker verification score correction system 1 according to the embodiment will be described below with reference to the drawings. Note that the processing described below is an example, and it is possible to change the processing order, delete some processing, or add other processing.
[0065] First, with reference to FIG. 1, the overall processing flow for speaker verification executed in a speaker verification score correction system 1 according to the embodiment will be described.
[0066] In the speaker verification score correction system 1, the storage unit 103 of the speaker verification device 10 stores registered speaker data 3 including the voice of a registered speaker (registered voice) that has been registered in advance.
[0067] The input / output unit 101 of the speaker verification device 10 acquires the speech of an unknown speaker to be evaluated (evaluation speech). The execution unit 102 then performs speaker verification processing, comparing the registered speech with the evaluation speech to calculate a similarity (score) (S1). Specifically, in the speaker verification processing, the execution unit 102 calculates a speaker verification similarity (score) indicating the similarity between speaker expression vectors (x-vectors) extracted from the registered speech and the evaluation speech. The execution unit 102 also performs a score correction processing, combining at least two or more score correction techniques to correct the speaker verification score (S2). The execution unit 102 also performs a determination (matching) processing, comparing the speaker verification similarity (corrected score) corrected by the score correction processing with a preset threshold value to determine whether the unknown speaker matches the registered speaker (S3). The input / output unit 101 then outputs the determination result (matching result).
[0068] 4 is a flowchart showing an example of the flow of the score correction process executed in the speaker verification device 10 according to the first embodiment. The flow in FIG. 4 corresponds to the processes in S1 and S2 in FIG.
[0069] First, the execution unit 102 performs QMF processing (first processing) and adds or subtracts (corrects) the speaker verification score using quality indicators (metadata) such as the speech length between voices and signal-to-noise ratio (SN) (S101).
[0070] Thereafter, the execution unit 102 performs CMF processing (third processing) and penalizes (corrects) the speaker verification score corrected by QMF processing based on the variation in the features (e.g., voiceprints) of each voice over time (S102).
[0071] Then, the execution unit 102 performs ASnorm processing (second processing) and normalizes (corrects) the speaker matching scores corrected by QMF and CMF processing based on the matching results with an unspecified number of speakers (Imposter) (S103).
[0072] FIG. 5 is a diagram for explaining an example of the QMF process of FIG.
[0073] In the QMF processing, the execution unit 102 calculates the Cosine similarity between the registered speaker expression vector 3b extracted from the registered speech 3a and the evaluation speaker expression vector 6b extracted from the evaluation speech 6a as a score for speaker matching (S201).
[0074] This cosine similarity is a measure of the similarity between speaker expression vectors, and is, for example, the cosine distance. This cosine distance is calculated, for example, as the inner product of the speaker expression vectors divided by the product of the magnitudes (L2 norms) of the speaker expression vectors. For example, if the cosine similarity between speaker expression vectors is "1", the two speaker expression vectors are completely similar. For example, if the cosine similarity between speaker expression vectors is "0", the two speaker expression vectors are not related to whether they are similar or not. For example, if the cosine similarity between speaker expression vectors is "-1", the two speaker expression vectors are completely dissimilar.
[0075] The execution unit 102 also calculates the cosine similarity between the registered speaker expression vector 3b and the unspecified majority speaker expression vector for QMF 4b (S202). Similarly, the execution unit 102 calculates the cosine similarity between the evaluation speaker expression vector 6b and the unspecified majority speaker expression vector for QMF 4b (S203).
[0076] Then, the execution unit 102 performs data selection to select the Cosine similarity when comparing the unspecified majority speaker expression vector 4b for QMF, which is similar to the registered speaker expression vector 3b, with the evaluation speaker expression vector 6b (S204). Similarly, the execution unit 102 performs data selection to select the Cosine similarity when comparing the unspecified majority speaker expression vector 4b for QMF, which is similar to the evaluation speaker expression vector 6b, with the registered speaker expression vector 3b (S205).
[0077] For example, the execution unit 102 selects unspecified majority speaker expression vectors 4b for QMF that have a high Cos similarity to the evaluation speaker expression vector 6b from among the unspecified majority speaker expression vectors 4b for QMF that have a high Cos similarity to the registered speaker expression vector 3b. For example, the execution unit 102 selects unspecified majority speaker expression vectors 4b for QMF that have a high Cos similarity to the registered speaker expression vector 3b from among the unspecified majority speaker expression vectors 4b for QMF that have a high Cos similarity to the evaluation speaker expression vector 6b. Here, the speaker expression vectors with a high Cos similarity may be speaker expression vectors whose Cos similarity is higher than a predetermined threshold, or may be a predetermined number of speaker expression vectors sorted in descending order of Cos similarity.
[0078] The execution unit 102 also calculates the average value of the Cosine similarity of the unspecified majority speaker expression vectors 4b for QMF after the data selection with respect to the registered speaker expression vectors 3b (S206). Similarly, the execution unit 102 calculates the average value of the Cosine similarity of the unspecified majority speaker expression vectors 4b for QMF after the data selection with respect to the evaluation speaker expression vectors 6b (S207).
[0079] Then, the execution unit 102 determines parameters for QMF using the average value of the Cosine similarity of the unspecified majority speaker expression vector 4b for QMF to the registered speaker expression vector 3b and the average value of the Cosine similarity of the unspecified majority speaker expression vector 4b for QMF to the evaluation speaker expression vector 6b (S208).
[0080] The QMF parameters may be the calculated average values themselves, or may be values calculated by calculation based on the average values. These QMF parameters are also treated as metadata in the QMF processing. That is, the metadata used in the QMF processing according to the embodiment includes the enrollment metadata 3c, the unspecified majority speaker metadata 4d for ASnorm, the evaluation metadata 6c, and the QMF parameters, which are the score averages of the enrollment speaker expression vectors 3b and the evaluation speaker expression vectors 6b with the unspecified majority speakers.
[0081] Thereafter, the execution unit 102 calculates a corrected score by correcting the Cos similarity as the speaker verification score calculated in the process of S201 (S209). Specifically, the execution unit 102 inputs the Cos similarity and metadata including the QMF parameters calculated in the process of S208 to the trained QMF model 5. The execution unit 102 also acquires the output of the trained QMF model 5 in response to the input of the Cos similarity and the metadata as a corrected score (similarity).
[0082] 6 is a diagram illustrating an example of the CMF processing of FIG. 4. In the CMF processing, the execution unit 102 calculates a CMF value indicating the variation of the feature of the registered speech 3a over time, i.e., the registered speaker expression vector 3b over time (S301). Similarly, the execution unit 102 calculates a CMF value indicating the variation of the feature of the evaluation speech 6a over time, i.e., the evaluation speaker expression vector 6b over time (S302). Then, the execution unit 102 calculates a corrected score (similarity) by scaling and correcting the similarity (corrected score) corrected by the QMF processing using the calculated CMF value as a correction coefficient.
[0083] In addition, when CMF processing is performed prior to QMF processing or ASnorm processing, the execution unit 102 calculates the Cos similarity between the registered speaker expression vector 3b and the evaluation speaker expression vector 6b as the speaker matching score, and then corrects it using the CMF value.
[0084] FIG. 7 is a diagram for explaining an example of the ASnorm process in FIG.
[0085] In the ASnorm process, the execution unit 102 calculates the Cosine similarity (speaker matching score) between the enrollment speaker expression vector 3b extracted from the enrollment speech 3a and the evaluation speaker expression vector 6b extracted from the evaluation speech 6a, and corrects the calculated Cosine similarity through QMF and CMF processes (S401). This process of S401 corresponds to the processes of S101 and S102 in Fig. 4. In other words, the process of this step can be a process of obtaining the similarity (corrected score) corrected through QMF and CMF processes.
[0086] The execution unit 102 also executes QMF processing and CMF processing using the registered speaker expression vector 3b and the ASnorm unspecified multi-speaker expression vector 4c as inputs, and calculates each similarity (S402). That is, the execution unit 102 executes the QMF processing of Fig. 5 using the ASnorm unspecified multi-speaker expression vector 4c instead of the evaluation speaker expression vector 6b. The execution unit 102 also executes the CMF processing of Fig. 6 using the ASnorm unspecified multi-speaker speech 4a instead of the evaluation speech 6a.
[0087] Similarly, the execution unit 102 executes QMF processing and CMF processing using the evaluation speaker expression vector 6b and the ASnorm unspecified multi-speaker expression vector 4c as input, and calculates each similarity (S403). That is, the execution unit 102 executes the QMF processing of Fig. 5 using the ASnorm unspecified multi-speaker expression vector 4c instead of the registered speaker expression vector 3b. Also, the execution unit 102 executes the CMF processing of Fig. 6 using the ASnorm unspecified multi-speaker speech 4a instead of the registered speech 3a.
[0088] The execution unit 102 then performs data selection to select the Cosine similarity when comparing the unspecified majority speaker expression vector 4c for ASnorm, which is similar to the registered speaker expression vector 3b, with the evaluation speaker expression vector 6b (S404), for example, in the same manner as the processing of S204 in Fig. 5. The execution unit 102 also performs data selection to select the Cosine similarity when comparing the unspecified majority speaker expression vector 4c for ASnorm, which is similar to the evaluation speaker expression vector 6b, with the registered speaker expression vector 3b (S405), for example, in the same manner as the processing of S205 in Fig. 5.
[0089] The execution unit 102 also calculates the average value and variance of the Cosine similarity of the unspecified majority speaker expression vector 4c for ASnorm after the data selection with respect to the registered speaker expression vector 3b (S406). Similarly, the execution unit 102 calculates the average value and variance of the Cosine similarity of the unspecified majority speaker expression vector 4c for ASnorm after the data selection with respect to the evaluation speaker expression vector 6b (S407).
[0090] Then, the execution unit 102 normalizes (corrects) the Cos similarity (corrected score) of the registered speaker expression vector 3b and the evaluation speaker expression vector 6b corrected by the QMF processing and the CMF processing using the average value and variance value of the Cos similarity of the unspecified majority speaker expression vector 4c for ASnorm for each of the registered speaker expression vector 3b and the evaluation speaker expression vector 6b, thereby calculating the corrected score (S408).
[0091] As described above, the speaker verification device 10 according to the embodiment corrects the speaker verification score by combining a plurality of score correction techniques.
[0092] FIG. 8 is a diagram for explaining improvement in accuracy of speaker verification by score correction according to the first embodiment.
[0093] In Figure 8, "BE1" indicates the first post-processing (BE). Similarly, "BE2" and "BE3" indicate the second and third post-processing, respectively. Furthermore, "minC" and "EER" are indices for measuring speaker verification performance, with smaller values indicating higher performance. "minC", also known as minDCF, is an index used to evaluate systems in speaker verification competitions held by the Speaker Recognition Evaluation (SRE) of the National Institute of Standards and Technology (NIST) in the United States. Furthermore, "EER" is an index called the equivalent error rate, used to evaluate biometric authentication systems. This "EER" is the value at which the false rejection rate (FRR), which indicates the rate at which a registered person (the person himself / herself) is mistakenly determined to be someone other than the registered person during authentication, is equal to the false acceptance rate (FAR), which indicates the rate at which a stranger is mistakenly determined to be the registered person during authentication.
[0094] Specifically, the speaker verification device 10 according to the embodiment is configured to use, as metadata, the score average with the unspecified large number of speaker data 4 in addition to quality indicators such as utterance length and SN in the QMF processing. The speaker verification device 10 according to the embodiment executes a combination of three score correction processes (post-processing: BE), namely, QMF processing, CMF processing, and ASnorm processing, in an order in which QMF processing is applied at least prior to ASnorm processing.
[0095] As an example, as shown in the first row of FIG. 8, the speaker verification apparatus 10 according to the embodiment applies and combines three score correction methods in the order of QMF processing, CMF processing, and ASnorm processing.
[0096] As an example, as shown in the second row of FIG. 8, the speaker verification apparatus 10 according to the embodiment applies and combines three score correction methods in the order of QMF processing, ASnorm processing, and CMF processing.
[0097] As an example, as shown in the third row of FIG. 8, the speaker verification apparatus 10 according to the embodiment applies and combines three score correction methods in the order of CMF processing, QMF processing, and ASnorm processing.
[0098] These configurations reduce the values of "minC" and "EER" compared to when QMF processing is applied after ASnorm processing (lines 4 to 6), thereby improving speaker verification performance. Therefore, speaker recognition scores that may fluctuate in a real environment can be appropriately corrected.
[0099] Other embodiments of the present disclosure will be described below with reference to the drawings. In the following description of each embodiment, differences will be mainly described, and descriptions of content that overlaps with the above-described content will be omitted as appropriate.
[0100] Second Embodiment Fig. 9 is a diagram for explaining an example of ASnorm processing according to a second embodiment. Here, differences from the ASnorm processing exemplified in Fig. 7 will be mainly explained.
[0101] In the ASnorm processing of this embodiment, after processing S401, the execution unit 102 calculates the Cos similarity (speaker matching score) between the registered speaker expression vector 3b and the unspecified majority speaker expression vector 4c for ASnorm, instead of performing QMF processing and CMF processing using the registered speaker expression vector 3b and the unspecified majority speaker expression vector 4c for ASnorm as input (S501).
[0102] Similarly, in the ASnorm processing of this embodiment, instead of performing QMF processing and CMF processing using the evaluation speaker expression vector 6b and the unspecified majority speaker expression vector 4c for ASnorm as input, the execution unit 102 calculates the Cos similarity (speaker matching score) between the evaluation speaker expression vector 6b and the unspecified majority speaker expression vector 4c for ASnorm (S502).
[0103] Then, based on the Cosine similarity calculated in the process of S501, the execution unit 102 performs data selection to select the Cosine similarity when the unspecified majority speaker expression vector 4c for ASnorm, which is similar to the registered speaker expression vector 3b, is compared with the evaluated speaker expression vector 6b (S404).Furthermore, based on the Cosine similarity calculated in the process of S502, the execution unit 102 performs data selection to select the Cosine similarity when the unspecified majority speaker expression vector 4c for ASnorm, which is similar to the evaluated speaker expression vector 6b, is compared with the registered speaker expression vector 3b (S405).
[0104] 9 illustrates a case where, in ASnorm processing applied after at least QMF processing, QMF processing and CMF processing are excluded when calculating a score with the unspecified majority speaker expression vector 4c, but this is not limiting. In ASnorm processing applied after at least QMF processing, it is sufficient that at least QMF processing is excluded when calculating a score with the unspecified majority speaker expression vector 4c, and CMF processing does not have to be excluded.
[0105] As described above, in the QMF processing, when the score average with the unspecified multi-speaker data 4 is used as metadata in addition to quality indicators such as speech length and SN, applying the QMF processing prior to the ASnorm processing can improve the performance of speaker verification. In this situation, the unspecified multi-speaker (imposter) data may contain tens of thousands of pieces of utterance data. Furthermore, in the ASnorm processing, the QMF processing and the CMF processing are applied to calculate the similarity between the registration / evaluation data and the imposter data. Therefore, when the QMF processing is applied prior to the ASnorm processing, the ASnorm processing involves a similarity calculation between the unspecified multi-speaker expression vector 4b for QMF and the unspecified multi-speaker expression vector 4c for ASnorm during the QMF processing. This means that similarity calculations on the scale of tens of thousands by tens of thousands are performed, resulting in a problem of an enormous amount of computation.
[0106] In contrast, the speaker verification device 10 according to the present embodiment is configured to exclude at least the QMF process when calculating the score with the unspecified majority speaker expression vector 4c in the ASnorm process. This configuration makes it possible to reduce the amount of calculation required for the ASnorm process that is applied at least after the QMF process.
[0107] 10 is a diagram for explaining improvement in accuracy of speaker verification by score correction according to the second embodiment. As shown in Fig. 10, even when the amount of calculation is reduced by excluding at least QMF processing when calculating the score with the unspecified many speaker expression vector 4c in ASnorm processing, the performance of speaker verification can be improved by applying QMF processing, which uses the score average with the unspecified many speaker data 4 as metadata, in addition to quality indicators such as utterance length and SN, prior to ASnorm processing.
[0108] As an example, as shown in the first row of FIG. 10, the speaker verification device 10 according to this embodiment applies and combines three score correction methods in the order of QMF processing, CMF processing, and ASnorm processing.
[0109] As an example, as shown in the second row of FIG. 10, the speaker verification device 10 according to this embodiment applies and combines three score correction methods in the order of QMF processing, ASnorm processing, and CMF processing.
[0110] As described above, according to the configuration of this embodiment, by applying ASnorm processing after at least QMF processing, it is possible to improve matching accuracy while reducing the amount of calculation required for ASnorm processing that is applied after at least QMF processing.
[0111] 11 is a diagram for explaining an example of ASnorm processing according to a third embodiment. Here, differences from the ASnorm processing exemplified in FIG. 10 will be mainly explained.
[0112] In the ASnorm processing according to this embodiment, the execution unit 102 applies QMF processing and CMF processing to the mean value and variance of the Cos similarity of the unspecified majority speaker expression vector 4c for ASnorm with respect to the registered speaker expression vector 3b, calculated in the processing of S406 (S601). Similarly, the execution unit 102 applies QMF processing and CMF processing to the mean value and variance of the Cos similarity of the unspecified majority speaker expression vector 4c for ASnorm with respect to the evaluation speaker expression vector 6b, calculated in the processing of S407 (S602).
[0113] Then, the execution unit 102 calculates the corrected score by normalizing (correcting) the Cos similarity (corrected score) of the registered speaker expression vector 3b and the evaluation speaker expression vector 6b corrected by the QMF processing and the CMF processing using the average value and variance value of the Cos similarity corrected by the QMF processing and the CMF processing (S408).
[0114] 11 illustrates an example in which, in ASnorm processing applied after at least QMF processing, the amount of calculation is reduced by calculating and approximating the mean and variance of the Cos similarity for the unspecified majority speaker expression vector 4c before the QMF processing and the CMF processing. However, this is not limiting. In ASnorm processing applied after at least QMF processing, it is sufficient to calculate and approximate the mean and variance of the Cos similarity for the unspecified majority speaker expression vector 4c at least before the QMF processing, and the CMF processing may be performed before calculating the mean and variance of the Cos similarity for the unspecified majority speaker expression vector 4c.
[0115] In the ASnorm processing according to this embodiment, the number of data selected in the data selection processing is equal to the number of original data. Also, in the ASnorm processing according to this embodiment, the metadata and score values used in the QMF processing are independent of each other.
[0116] As described above, the speaker verification device 10 according to this embodiment is configured to perform QMF processing and CMF processing on the unspecified majority speaker expression vector 4c and the mean and variance of its metadata. Here, both QMF processing and CMF processing are linear transformations. Therefore, whether the mean and variance are calculated after QMF processing and CMF processing or before QMF processing and CMF processing, the calculation results, i.e., the corrected score values, are approximately the same. Therefore, even with the configuration according to this embodiment, it is possible to reduce the amount of calculation for ASnorm processing, which is applied at least after QMF processing.
[0117] In each of the above-described embodiments, "is it A?" refers to at least one of "is A" and "is not A." In other words, in each of the above-described embodiments, the determination of "is A" may be realized by determining "is A," or by determining "is not A," or by determining both of these.
[0118] The program executed by the speaker verification device 10 of each of the above-described embodiments may be provided by being recorded in an installable or executable file format on a computer-readable recording medium (Computer Program Product) such as a CD-ROM, FD, CD-R, or DVD.
[0119] The programs executed by the speaker verification device 10 of each of the above-described embodiments may be stored on a computer connected to a network such as the Internet and provided by being downloaded via the network.The programs executed by the speaker verification device 10 of each of the above-described embodiments may be provided or distributed via a network such as the Internet.
[0120] Furthermore, the programs executed by the speaker verification device 10 of each of the above-described embodiments may be provided by being pre-installed in a ROM or the like.
[0121] According to at least one of the embodiments described above, it is possible to appropriately correct speaker recognition scores that may fluctuate in a real environment.
[0122] Although several embodiments of the present invention have been described, these embodiments are presented as examples and are not intended to limit the scope of the invention. These embodiments can be implemented in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their modifications are included within the scope and spirit of the invention, as well as within the scope of the invention described in the claims and their equivalents.
[0123] 1 Speaker verification score correction system 10 Speaker verification device 101 Input / output unit 102 Execution unit 103 Storage unit 3 Enrolled speaker data 3a Enrolled speech 3b Enrolled speaker expression vector (xvector) 3c Enrolled metadata 4 Unspecified multi-speaker (Imposter) data 4a Unspecified multi-speaker speech for ASnorm 4b Unspecified multi-speaker expression vector for QMF 4c Unspecified multi-speaker expression vector for ASnorm 4d Unspecified multi-speaker metadata for ASnorm 5 Trained QMF model 6a Evaluation speech 6b Evaluation speaker expression vector 6c Evaluation metadata 8 Information processing device 81 Processor 82 Main memory device 83 Auxiliary memory device 84 I / F
Claims
1. An information processing method executed by at least one processor in an information processing device having at least one processor, comprising: storing registration speaker data relating to registration speech, which is the speech of a registered registration speaker; acquiring evaluation speech, which is the speech of a speaker to be evaluated; calculating a score indicating the similarity between a registration speaker expression vector, which is a feature of the registration speech, and an evaluation speaker expression vector, which is a feature of the evaluation speech; correcting the score by combining at least two or more score correction processes; comparing the corrected score with a predetermined threshold to determine whether the speaker to be evaluated matches the registration speaker; and outputting information indicating the determination result; the at least two or more score correction processes including: a first process that corrects the score using a model of interaction between speeches in metadata including quality indicators of the registration speech and the evaluation speech; and a second process that corrects a speaker matching score based on a matching result of an unspecified multi-speaker speech, which is the speech of an unspecified multi-speaker, with an unspecified multi-speaker expression vector; and at least the first process is applied before the second process.
2. The information processing method according to claim 1, wherein the at least two score correction processes further include a third process of correcting the score based on the variability of the speaker expression vector for each time period of each speech.
3. The information processing method according to claim 1, wherein the metadata includes an average value of the scores based on an unspecified multi-speaker expression vector used in the first processing of the unspecified multi-speaker speech.
4. The information processing method of claim 3, wherein the first processing comprises: calculating the scores of each of the registered speaker expression vector and the evaluated speaker expression vector and an unspecified majority speaker expression vector used in the first processing; selecting the scores when the unspecified majority speaker expression vector used in the first processing that has a high score with the registered speaker expression vector is compared with the evaluated speaker expression vector; selecting the scores when the unspecified majority speaker expression vector used in the first processing that has a high score with the evaluated speaker expression vector is compared with the registered speaker expression vector; and calculating an average value of the scores of the unspecified majority speaker expression vector used in the first processing for each of the registered speaker expression vector and the evaluated speaker expression vector.
5. The second processing includes the first processing which receives as input the registered speaker expression vector and an unspecified majority speaker expression vector used in the second processing, and the second processing which receives as input the evaluation speaker expression vector and an unspecified majority speaker expression vector used in the second processing, calculating the scores of each of the registered speaker expression vector and the evaluation speaker expression vector and the unspecified majority speaker expression vector used in the second processing, selecting the scores when the unspecified majority speaker expression vector used in the second processing which has a high score with the registered speaker expression vector is compared with the evaluation speaker expression vector, selecting the scores when the unspecified majority speaker expression vector used in the second processing which has a high score with the evaluation speaker expression vector is compared with the registered speaker expression vector, calculating the average value and variance of the scores of the unspecified majority speaker expression vector used in the second processing for each of the registered speaker expression vector and the evaluation speaker expression vector, The information processing method according to claim 1 , further comprising: normalizing the score corrected in at least the first process using the calculated mean value and variance value.
6. An information processing method according to any one of claims 1 to 4, wherein the second processing comprises: calculating the score between the registered speaker expression vector and an unspecified majority speaker expression vector used in the second processing; calculating the score between the evaluation speaker expression vector and an unspecified majority speaker expression vector used in the second processing; selecting the score when the unspecified majority speaker expression vector used in the second processing, which has a high score with the registered speaker expression vector, is compared with the evaluation speaker expression vector; selecting the score when the unspecified majority speaker expression vector used in the second processing, which has a high score with the evaluation speaker expression vector, is compared with the registered speaker expression vector; calculating the average value and variance of the scores of the unspecified majority speaker expression vector used in the second processing for each of the registered speaker expression vector and the evaluation speaker expression vector; and normalizing at least the score corrected in the first processing using the calculated average value and variance.
7. The information processing method of claim 6, wherein the second processing includes the first processing on the calculated mean value and variance value, and the score corrected by at least the first processing is normalized using the mean value and variance value to which the first processing has been applied.
8. The information processing method according to claim 2, wherein the first process, the third process, and the second process are applied in this order.
9. An information processing device comprising: a memory for storing registration speaker data relating to registration speech, which is the speech of a registered registration speaker; and at least one processor configured to: acquire evaluation speech, which is the speech of a speaker to be evaluated; calculate a score indicating the similarity between a registration speaker expression vector, which is a feature of the registration speech, and an evaluation speaker expression vector, which is a feature of the evaluation speech; correct the score by combining at least two or more score correction processes; compare the corrected score with a predetermined threshold to determine whether the speaker to be evaluated matches the registration speaker; and output information indicating the determination result; wherein the at least two or more score correction processes include a first process for correcting the score using a model of interaction between speeches in metadata including quality indicators of the registration speech and the evaluation speech, and a second process for correcting a speaker matching score based on a matching result of an unspecified multi-speaker speech, which is the speech of an unspecified multi-speaker, with an unspecified multi-speaker expression vector; and at least the first process is applied before the second process.
10. A program for causing a computer to execute the following steps: store registration speaker data regarding registration speech, which is the speech of a registered registration speaker; acquire evaluation speech, which is the speech of a speaker to be evaluated; calculate a score indicating the similarity between a registration speaker expression vector, which is a feature of the registration speech, and an evaluation speaker expression vector, which is a feature of the evaluation speech; correct the score by combining at least two or more score correction processes; compare the corrected score with a predetermined threshold to determine whether the speaker to be evaluated matches the registration speaker; and output information indicating the determination result, wherein the at least two or more score correction processes include a first process that corrects the score using a model of interaction between speech in metadata including quality indicators of the registration speech and the evaluation speech, and a second process that corrects the speaker matching score based on the result of matching an unspecified multi-speaker speech, which is the speech of an unspecified multi-speaker, with an unspecified multi-speaker expression vector, and at least the first process is applied before the second process.