Voiceprint comparison method and device based on log-likelihood ratio calibration and readable medium

By using a mixed Gaussian model to fit the same-original and non-homologous data in the voiceprint alignment technology and calculating the log-likelihood ratio calibration expression, the problem of failure of the log-likelihood ratio calibration algorithm in the prior art when the Gaussian distribution is largely different from the actual data distribution is improved, and the accuracy of voiceprint alignment is improved.

CN119993167APending Publication Date: 2025-05-13XIAMEN KUAISHANGTONG TECH CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510114492.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the existing voiceprint comparison technology, assuming that the variance of the Gaussian distribution is the same and the data distribution conforms to the single Gaussian function, the log-likelihood ratio calibration algorithm fails and the calculated LLR is inaccurate.

Method used

The vocalprint alignment method based on log-likelihood ratio calibration is used to fit the homologous data and non-homologous data through a mixed Gaussian model, and the homologous distribution and non-homologous distribution are calculated, and the sampling scores are calculated by the maximum and minimum values ​​to form a sampling score set, and the log-likelihood ratio is calculated, and the log-likelihood ratio calibration expression is fitted.

Benefits of technology

It effectively avoids the problem of excessive differences in the distribution of the fitted data and the actual data, resulting in the failure of the log-likelihood calibration expression, and improves the accuracy and scope of application of the log-likelihood calibration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993167A_ABST
    Figure CN119993167A_ABST
Patent Text Reader

Abstract

The invention discloses a voiceprint comparison method and device based on log-likelihood ratio calibration and a readable medium. The voiceprint comparison method comprises the steps of collecting voice data of multiple persons in a target scene and forming a voice data set; performing voiceprint comparison on every two pieces of voice data of the same person or different persons in the voice data set, and constructing homologous data and non-homologous data; respectively calculating homologous distribution and non-homologous distribution based on the homologous data and the non-homologous data, and obtaining a sampling score set and a log-likelihood ratio set; calculating a log-likelihood ratio calibration expression based on the sampling score set and the log-likelihood ratio set; and obtaining to-be-detected voice data and sample voice data in the target scene, performing voiceprint comparison to obtain a corresponding comparison score, and performing calculation according to the log-likelihood ratio calibration expression to obtain a corresponding log-likelihood ratio so as to determine the similarity between the to-be-detected voice data and the sample voice data. According to the invention, the application range and accuracy of the original calibration algorithm are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of voiceprint comparison, and in particular to a voiceprint comparison method, device and readable medium based on log-likelihood ratio calibration. Background Art

[0002] Likelihood ratio (LR for short) is expressed as f is the probability density function, H s H stands for homology d Represents different sources. For example, if we calculate the likelihood ratio of a three-legged cow, then x represents three legs, Hs is the total number of cows, and Hd is the total number of non-cows, that is, f(x|H s ) is the probability of a three-legged cow, assuming it is 0.2, f(x|H d ) is the probability that a cow does not have three legs, assuming it is 0.4, then The log-likelihood ratio (LLR) is generally used to measure the probability of an event occurring. This value can be used in international courts to determine the credibility of evidence, such as detecting whether a speech and a sample speech are of the same source (i.e., whether they come from the same person).

[0003] The calibration algorithm of the log-likelihood ratio generally adopts the expression LLR=αx+β, which is a linear regression expression derived based on a special double Gaussian expression, that is, where g(x|u s ,σ) represents the homologous Gaussian distribution, g(x|u d ,σ) are non-homologous Gaussian distributions, and the variances σ of the two Gaussian distributions are the same. The above calibration algorithm has two problems. One is that it assumes that the variances of the two Gaussian distributions are the same, and the other is that it assumes that the distribution of the data conforms to a single Gaussian function. Therefore, if the Gaussian distribution and the actual data distribution are too different, the calibration algorithm (LLR = αx + β) will fail, and the calculated LLR will be inaccurate. Summary of the invention

[0004] The purpose of this application is to propose a voiceprint comparison method, device and readable medium based on log-likelihood ratio calibration to address the above-mentioned technical problems.

[0005] In a first aspect, the present invention provides a voiceprint comparison method based on log-likelihood ratio calibration, comprising the following steps:

[0006] Collecting voice data of multiple people in the target scene and forming a voice data set, wherein at least two voice data are collected for each person;

[0007] Perform voiceprint comparison on every two voice data of the same person in the voice data set to obtain a first comparison score, and all the first comparison scores constitute homologous data; perform voiceprint comparison on every two voice data of different people in the voice data set to obtain a second comparison score, and all the second comparison scores constitute non-homologous data;

[0008] Based on the homologous data and the non-homologous data, the homologous distribution and the non-homologous distribution are calculated respectively; N sampling scores are calculated according to the maximum and minimum values ​​in the homologous data and the non-homologous data to form a sampling score set; the log-likelihood ratio corresponding to each sampling score in the sampling score set is calculated according to the homologous distribution and the non-homologous distribution to form a log-likelihood ratio set; the log-likelihood ratio calibration expression is calculated based on the sampling score set and the log-likelihood ratio set;

[0009] The speech data to be detected and the sample speech data in the target scenario are obtained and voiceprint comparison is performed to obtain the corresponding comparison score. The log-likelihood ratio value corresponding to the speech data to be detected and the sample speech data is calculated according to the log-likelihood ratio calibration expression. The similarity between the speech data to be detected and the sample speech data is determined according to the log-likelihood ratio value corresponding to the speech data to be detected and the sample speech data.

[0010] Preferably, the homologous distribution and the non-homologous distribution are calculated based on the homologous data and the non-homologous data, respectively, specifically including:

[0011] The homologous data are input into the EM algorithm to obtain the parameters of the first mixed Gaussian model, and the first mixed Gaussian model with known parameters is recorded as the homologous distribution G s ;

[0012] The non-homologous data are input into the EM algorithm to obtain the parameters of the second mixed Gaussian model, and the second mixed Gaussian model with known parameters is recorded as the non-homologous distribution G d .

[0013] Preferably, N sampling scores are calculated according to the maximum and minimum values ​​in the homologous data and the non-homologous data to form a sampling score set, which specifically includes:

[0014] The nth sampling score is calculated using the following formula:

[0015]

[0016] Among them, x n represents the nth sampling score, n∈[1,N], x max Indicates the maximum value between homologous data and non-homologous data, x min represents the minimum value in homologous data and non-homologous data, N represents the total number of sampled scores; the sampled score set is represented by S = {x1, x2, ..., xN}.

[0017] Preferably, the log-likelihood ratio corresponding to each sampling score in the sampling score set is calculated according to the homologous distribution and the non-homologous distribution to form a log-likelihood ratio set, which specifically includes:

[0018] The log-likelihood ratio corresponding to the nth sampling score is calculated using the following formula:

[0019]

[0020] Among them, x n Represents the nth sampling score, n∈[1,N], m n Represents the log-likelihood ratio corresponding to the nth sampling score, G s represents homologous distribution, G d represents non-homologous distribution, G s (x n ) represents the value corresponding to the nth sampling score in the homologous distribution, G d (x n ) represents the value corresponding to the nth sampling score in the non-homologous distribution; the set of log-likelihood ratios is denoted as M = {m1, m2, m3, ..., m N}.

[0021] Preferably, the log-likelihood ratio calibration expression is calculated based on the sampling score set and the log-likelihood ratio set, specifically including:

[0022] Each sampling score x in the sampling score set is used as the independent variable in the log-likelihood ratio calibration expression, and each log-likelihood ratio LLR in the log-likelihood ratio set is used as the dependent variable in the log-likelihood ratio calibration expression. The slope α and intercept β in the log-likelihood ratio calibration expression LLR=αx+β are calculated according to the least squares method.

[0023] Preferably, the value range of the log-likelihood ratio is from negative infinity to positive infinity. When the log-likelihood ratio LLR corresponding to the speech data to be detected and the sample speech data is less than 0, it indicates the degree to which the speech data to be detected and the sample speech data are from different sources; when the log-likelihood ratio LLR corresponding to the speech data to be detected and the sample speech data is greater than 0, it indicates the degree to which the speech data to be detected and the sample speech data are from the same source.

[0024] In a second aspect, the present invention provides a voiceprint comparison device based on log-likelihood ratio calibration, comprising:

[0025] A data acquisition module is configured to collect voice data of multiple persons in a target scene and form a voice data set, wherein at least two voice data are collected for each person;

[0026] The voiceprint comparison module is configured to compare the voiceprints of every two voice data of the same person in the voice data set to obtain a first comparison score, and all the first comparison scores constitute homologous data; compare the voiceprints of every two voice data of different people in the voice data set to obtain a second comparison score, and all the second comparison scores constitute non-homologous data;

[0027] The expression calculation module is configured to calculate homologous distribution and non-homologous distribution respectively based on homologous data and non-homologous data; calculate N sampling scores according to the maximum and minimum values ​​in the homologous data and non-homologous data to form a sampling score set; calculate the log-likelihood ratio corresponding to each sampling score in the sampling score set according to the homologous distribution and non-homologous distribution to form a log-likelihood ratio set; calculate the log-likelihood ratio calibration expression based on the sampling score set and the log-likelihood ratio set;

[0028] The similarity determination module is configured to obtain the voice data to be detected and the sample voice data in the target scenario and perform voiceprint comparison to obtain the corresponding comparison score, calculate the log-likelihood ratio value corresponding to the voice data to be detected and the sample voice data according to the log-likelihood ratio calibration expression, and determine the similarity between the voice data to be detected and the sample voice data according to the log-likelihood ratio value corresponding to the voice data to be detected and the sample voice data.

[0029] In a third aspect, the present invention provides an electronic device comprising one or more processors; a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner in the first aspect.

[0030] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in any implementation manner in the first aspect.

[0031] In a fifth aspect, the present invention provides a computer program product, comprising a computer program, which, when executed by a processor, implements the method described in any implementation manner in the first aspect.

[0032] Compared with the prior art, the present invention has the following beneficial effects:

[0033] (1) The voiceprint comparison method based on log-likelihood ratio calibration proposed in the present invention uses a mixed Gaussian model to fit homologous data and non-homologous data respectively to obtain homologous distribution and non-homologous distribution, thereby avoiding a large difference between the fitted data and the actual data distribution, which would cause the log-likelihood ratio calibration expression to fail.

[0034] (2) The voiceprint comparison method based on log-likelihood ratio calibration proposed in the present invention calculates multiple sampling scores through maximum and minimum values ​​to form a sampling score set, and obtains corresponding log-likelihood ratios through the sampling score set, homologous distribution and non-homologous distribution to form a log-likelihood ratio set. The log-likelihood ratio calibration expression is obtained by fitting the sampling score set and the log-likelihood ratio set, so that the LLR finally calculated by the log-likelihood ratio calibration expression is more effective, thereby improving the scope of application and accuracy of the original calibration algorithm. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0036] Figure 1 A flow chart of a voiceprint comparison method based on log-likelihood ratio calibration according to an embodiment of the present application;

[0037] Figure 2 A schematic diagram of a voiceprint comparison device based on log-likelihood ratio calibration according to an embodiment of the present application;

[0038] Figure 3 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0039] In order to make the purpose, technical scheme and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0040] Figure 1 A voiceprint comparison method based on log-likelihood ratio calibration provided by an embodiment of the present application is shown, comprising the following steps:

[0041] S1, collect voice data of multiple people in the target scene and form a voice data set, where at least two voice data are collected for each person.

[0042] Specifically, the embodiment of the present application takes the target scenario as the voiceprint comparison scenario in the case as an example, and collects voice data of some people according to the voice source of the case, and collects at least two voice data of each person. During the collection process, the identity label of the corresponding person is marked for each voice data, so as to distinguish whether the two voice data of the same person are being compared with the voiceprint or the two voice data of different people are being compared with the voiceprint in step S2.

[0043] S2, compare the voiceprints of every two voice data of the same person in the voice data set to obtain a first comparison score, and all the first comparison scores constitute homologous data; compare the voiceprints of every two voice data of different people in the voice data set to obtain a second comparison score, and all the second comparison scores constitute non-homologous data.

[0044] Specifically, the process of constructing homologous data is as follows: compare the voiceprints of any two voice data of the same person to obtain the scoring result, that is, the first comparison score. The set constructed by all the first comparison scores is the homologous data.

[0045] The process of constructing non-homologous data is as follows: perform voiceprint comparison on any two voice data of different people to obtain the scoring result, that is, the second comparison score. The set constructed by all the second comparison scores is the non-homologous data.

[0046] It should be noted that the voiceprint comparison method mentioned in the embodiments of the present application can adopt the existing voiceprint comparison method, such as first extracting the voice features of the two voice data respectively, and then calculating the score of the similarity index of the voice features of the two voice data. The larger the score, the greater the possibility that the two voice data belong to the same person. Conversely, the greater the possibility that the two voice data belong to different people. However, the score of the similarity index needs to be further calculated through the calibration algorithm of the log-likelihood ratio to obtain the corresponding log-likelihood ratio. The log-likelihood ratio (LLR) is generally used to measure the possibility of something happening. This value can be used to judge the credibility of evidence in international courts. For example, the detection voice and the sample voice are judged to be of the same origin (that is, whether they come from the same person). The specific method and process of voiceprint comparison are not limited in the embodiments of the present application.

[0047] S3, based on the homologous data and non-homologous data, respectively calculate the homologous distribution and non-homologous distribution; calculate N sampling scores according to the maximum and minimum values ​​in the homologous data and non-homologous data to form a sampling score set; calculate the log-likelihood ratio corresponding to each sampling score in the sampling score set according to the homologous distribution and non-homologous distribution to form a log-likelihood ratio set; calculate the log-likelihood ratio calibration expression based on the sampling score set and the log-likelihood ratio set.

[0048] In a specific embodiment, the homologous distribution and the non-homologous distribution are calculated based on the homologous data and the non-homologous data, respectively, including:

[0049] The homologous data are input into the EM algorithm to obtain the parameters of the first mixed Gaussian model, and the first mixed Gaussian model with known parameters is recorded as the homologous distribution G s ;

[0050] The non-homologous data are input into the EM algorithm to obtain the parameters of the second mixed Gaussian model, and the second mixed Gaussian model with known parameters is recorded as the non-homologous distribution G d .

[0051] Specifically, the expression of the mixed Gaussian model is Among them, g(x|u i ,σ i ) is the i-th one-dimensional Gaussian function, u i and σ i Respectively represent the mean and variance of the i-th one-dimensional Gaussian function, w i ∈[0,1] represents the function corresponding to the i-th one-dimensional Gaussian function.

[0052] Based on homologous data and non-homologous data, the EM algorithm is used to calculate the homologous distribution G s and non-homologous distribution G d . Homologous distribution G s and non-homologous distribution G d The above Gaussian mixture models are used to calculate the parameters of the two Gaussian mixture models through the EM algorithm, so the homologous distribution G can be obtained. s and non-homologous distribution G d .

[0053] In a specific embodiment, N sampling scores are calculated according to the maximum and minimum values ​​in the homologous data and the non-homologous data to form a sampling score set, which specifically includes:

[0054] The nth sampling score is calculated using the following formula:

[0055]

[0056] Among them, x n represents the nth sampling score, n∈[1,N], x max Indicates the maximum value between homologous data and non-homologous data, x min represents the minimum value in homologous data and non-homologous data, N represents the total number of sampled scores; the sampled score set is represented by S = {x1, x2, ..., x N}.

[0057] Specifically, get the maximum value x between homologous data and non-homologous datamax and the minimum value x min Here we will select the maximum value x in the mixed data consisting of homologous data and non-homologous data. max and the minimum value x min In the embodiment of the present application, N is 1000 as an example, and 1000 sampling scores are calculated according to the maximum and minimum values ​​to form a sampling score set S = {x1, x2, ..., x 1000}, where the nth sampling score n∈[1,1000].

[0058] In a specific embodiment, the log-likelihood ratio corresponding to each sampling score in the sampling score set is calculated according to the homologous distribution and the non-homologous distribution to form a log-likelihood ratio set, which specifically includes:

[0059] The log-likelihood ratio corresponding to the nth sampling score is calculated using the following formula:

[0060]

[0061] Among them, x n Represents the nth sampling score, n∈[1,N], m n Represents the log-likelihood ratio corresponding to the nth sampling score, G s represents homologous distribution, G d represents non-homologous distribution, G s (x n ) represents the value corresponding to the nth sampling score in the homologous distribution, G d (x n ) represents the value corresponding to the nth sampling score in the non-homologous distribution; the set of log-likelihood ratios is denoted as M = {m1, m2, m3, ..., m N}.

[0062] Specifically, according to the sampling score set S, homologous distribution and non-homologous distribution, the corresponding log-likelihood ratio (LLR) set M = {m1, m2, m3, ..., m 1000}, where the log-likelihood ratio corresponding to the nth sampling score is First, the values ​​corresponding to the nth sampling score in the homologous distribution and the values ​​corresponding to the nth sampling score in the non-homologous distribution are calculated using the homologous distribution and the non-homologous distribution, and then the corresponding log-likelihood ratio is calculated.

[0063] In a specific embodiment, the log-likelihood ratio calibration expression is calculated based on the sampling score set and the log-likelihood ratio set, specifically including:

[0064] Each sampling score x in the sampling score set is used as the independent variable in the log-likelihood ratio calibration expression, and each log-likelihood ratio LLR in the log-likelihood ratio set is used as the dependent variable in the log-likelihood ratio calibration expression. The slope α and intercept β in the log-likelihood ratio calibration expression LLR=αx+β are calculated according to the least squares method.

[0065] Specifically, according to the least squares method, the corresponding data in the sampling score set S and the log-likelihood ratio M are used to calculate α and β in the log-likelihood ratio calibration expression LLR=αx+β, thereby determining the log-likelihood ratio calibration expression. s and non-homologous distribution G d The variances in are not necessarily the same, and the data distribution conforms to the mixed Gaussian model, not the single Gaussian model, so the homologous distribution G is used here. s and non-homologous distribution G d When calculating the value corresponding to the nth sampling score in the homologous distribution and the value corresponding to the nth sampling score in the non-homologous distribution respectively, the final value will be more consistent with the actual data distribution, so the final fitted log-likelihood ratio calibration expression LLR=αx+β will also be more accurate.

[0066] S4, obtain the voice data to be detected and the sample voice data in the target scenario and perform voiceprint comparison to obtain the corresponding comparison score, calculate the log-likelihood ratio value corresponding to the voice data to be detected and the sample voice data according to the log-likelihood ratio calibration expression, and determine the similarity between the voice data to be detected and the sample voice data according to the log-likelihood ratio value corresponding to the voice data to be detected and the sample voice data.

[0067] In a specific embodiment, the value range of the log-likelihood ratio is from negative infinity to positive infinity. When the log-likelihood ratio LLR corresponding to the speech data to be detected and the sample speech data is less than 0, it indicates the degree to which the speech data to be detected and the sample speech data are from different sources; when the log-likelihood ratio LLR corresponding to the speech data to be detected and the sample speech data is greater than 0, it indicates the degree to which the speech data to be detected and the sample speech data are from the same source.

[0068] Specifically, the voice data to be tested and the sample voice data in the case are compared with each other to obtain the corresponding comparison score x. The log-likelihood ratio LLR corresponding to the comparison score x is calculated according to the log-likelihood ratio calibration expression LLR=αx+β, and the log-likelihood ratio LLR is used to evaluate the similarity of the two voice data. The voice data to be tested is one of the voice data in the sample database, and the sample voice data is one of the voice data in the sample database. The two are compared with each other, and the comparison score is converted into an LLR value according to the log-likelihood ratio calibration expression obtained in the above steps. The value range of LLR is positive and negative infinity. LLR<0 represents the degree of different sources of the two, and LLR>0 represents the degree of common sources of the two. As one example, it can be set as follows: when LLR is in the range of 3 to 4, it can be judged as homologous; when LLR is in the range of 2 to 3, it is more likely to be homologous; when LLR is in the range of 1 to 2, it tends to be homologous; when LLR is in the range of 0 to 1, there is a certain possibility of homology; when LLR=0, it is uncertain; when LLR is in the range of 0 to -1, there is a certain possibility of non-homologous; when LLR is in the range of -1 to -2, it tends to be non-homologous; when LLR is in the range of -2 to -3, it is more likely to be non-homologous; when LLR is in the range of -3 to -4, it can be judged as non-homologous.

[0069] Further references Figure 2 As an implementation of the methods shown in the above figures, the present application provides an embodiment of a voiceprint comparison device based on log-likelihood ratio calibration. Figure 1 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0070] The embodiment of the present application provides a voiceprint comparison device based on log-likelihood ratio calibration, comprising:

[0071] A data acquisition module 1 is configured to collect voice data of multiple persons in a target scene and form a voice data set, wherein at least two voice data are collected for each person;

[0072] The voiceprint comparison module 2 is configured to compare the voiceprints of every two voice data of the same person in the voice data set to obtain a first comparison score, and all the first comparison scores constitute homologous data; compare the voiceprints of every two voice data of different people in the voice data set to obtain a second comparison score, and all the second comparison scores constitute non-homologous data;

[0073] The expression calculation module 3 is configured to calculate homologous distribution and non-homologous distribution respectively based on homologous data and non-homologous data; calculate N sampling scores according to the maximum and minimum values ​​in the homologous data and non-homologous data to form a sampling score set; calculate the log-likelihood ratio corresponding to each sampling score in the sampling score set according to the homologous distribution and non-homologous distribution to form a log-likelihood ratio set; calculate the log-likelihood ratio calibration expression based on the sampling score set and the log-likelihood ratio set;

[0074] The similarity determination module 4 is configured to obtain the voice data to be detected and the sample voice data in the target scene and perform voiceprint comparison to obtain the corresponding comparison score, calculate the log-likelihood ratio value corresponding to the voice data to be detected and the sample voice data according to the log-likelihood ratio calibration expression, and determine the similarity between the voice data to be detected and the sample voice data according to the log-likelihood ratio value corresponding to the voice data to be detected and the sample voice data.

[0075] Figure 3 Schematic diagram of the hardware structure of the electronic device provided by the embodiment of the present invention. Figure 3 As shown, the electronic device of this embodiment includes: a processor 301 and a memory 302; wherein the memory 302 is used to store computer-executable instructions; the processor 301 is used to execute the computer-executable instructions stored in the memory to implement the various steps performed by the electronic device in the above embodiment. For details, please refer to the relevant description in the above method embodiment.

[0076] Optionally, the memory 302 may be independent or integrated with the processor 301 .

[0077] When the memory 302 is independently provided, the electronic device further includes a bus 303 for connecting the memory 302 and the processor 301 .

[0078] The embodiment of the present invention further provides a computer storage medium, in which computer execution instructions are stored. When the processor 301 executes the computer execution instructions, the above method is implemented.

[0079] The embodiment of the present invention further provides a computer program product, including a computer program. When the computer program is executed by the processor 301, the above method is implemented.

[0080] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of modules is only a logical function division, and there may be other division methods in actual implementation, such as multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or modules, which can be electrical, mechanical or other forms.

[0081] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to implement the solution of this embodiment.

[0082] In addition, each functional module in each embodiment of the present invention may be integrated into one processing unit, each module may exist physically separately, or two or more modules may be integrated into one unit. The unit formed by the above modules may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0083] The above-mentioned integrated module implemented in the form of a software function module can be stored in a computer-readable storage medium. The above-mentioned software function module is stored in a storage medium, including a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor 301 to perform some steps of the methods of various embodiments of the present application.

[0084] It should be understood that the processor 301 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor may be a microprocessor or the processor 301 may be any conventional processor 301, etc. The steps of the method disclosed in the invention may be directly embodied in the execution of the hardware processor 301, or may be executed by a combination of hardware and software modules in the processor 301.

[0085] The memory 302 may include a high-speed RAM memory, and may also include a non-volatile storage NVM, such as at least one disk memory, and may also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk, or an optical disk.

[0086] The bus 303 may be an Industry Standard Architecture (ISA), a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus 303 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the bus 303 in the drawings of the present application is not limited to only one bus 303 or one type of bus 303.

[0087] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The storage medium can be any available medium that can be accessed by a general or special purpose computer.

[0088] An exemplary storage medium is coupled to the processor 301, so that the processor 301 can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor 301. The processor 301 and the storage medium can be located in an application specific integrated circuit (Application Specific Integrated Circuits, referred to as ASIC). Of course, the processor 301 and the storage medium can also exist as discrete components in an electronic device or a main control device.

[0089] Those skilled in the art can understand that all or part of the steps of implementing the above-mentioned method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, the steps of the above-mentioned method embodiments are executed; and the aforementioned storage medium includes: ROM, RAM, disk or optical disk and other media that can store program codes.

[0090] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A voiceprint comparison method based on log-likelihood ratio calibration, characterized in that: The following steps are involved: Collecting voice data of multiple people in the target scene and forming a voice data set, wherein at least two voice data are collected for each person; Perform voiceprint comparison on every two voice data of the same person in the voice data set to obtain a first comparison score, and all the first comparison scores constitute homologous data; perform voiceprint comparison on every two voice data of different people in the voice data set to obtain a second comparison score, and all the second comparison scores constitute non-homologous data; Based on the homologous data and the non-homologous data, respectively, a homologous distribution and a non-homologous distribution are calculated; N sampling scores are calculated according to the maximum and minimum values ​​in the homologous data and the non-homologous data to form a sampling score set; based on the homologous distribution and the non-homologous distribution, the log-likelihood ratio corresponding to each sampling score in the sampling score set is calculated to form a log-likelihood ratio set; based on the sampling score set and the log-likelihood ratio set, a log-likelihood ratio calibration expression is calculated; Acquire the speech data to be detected and the sample speech data in the target scenario and perform voiceprint comparison to obtain the corresponding comparison score, calculate the log-likelihood ratio value corresponding to the speech data to be detected and the sample speech data according to the log-likelihood ratio calibration expression, and determine the similarity between the speech data to be detected and the sample speech data according to the log-likelihood ratio value corresponding to the speech data to be detected and the sample speech data.

2. The voiceprint comparison method based on log-likelihood ratio calibration according to claim 1, characterized in that: Calculating homologous distribution and non-homologous distribution based on the homologous data and non-homologous data respectively includes: The homologous data is input into the EM algorithm to obtain the parameters of the first mixed Gaussian model, and the first mixed Gaussian model with known parameters is recorded as the homologous distribution G s ; The non-homologous data are input into the EM algorithm to obtain the parameters of the second mixed Gaussian model, and the second mixed Gaussian model with known parameters is recorded as the non-homologous distribution G d .

3. The voiceprint comparison method based on log-likelihood ratio calibration according to claim 1, characterized in that: According to the maximum and minimum values ​​in the homologous data and the non-homologous data, N sampling scores are calculated to form a sampling score set, which specifically includes: The nth sampling score is calculated using the following formula: Among them, x n represents the nth sampling score, n∈[1,N], x max represents the maximum value among the homologous data and non-homologous data, x min represents the minimum value in the homologous data and non-homologous data, N represents the total number of sampled scores; the sampled score set is represented by S = {x1, x2, ..., x N }.

4. The voiceprint comparison method based on log-likelihood ratio calibration according to claim 1, characterized in that: Calculating the log-likelihood ratio corresponding to each sampling score in the sampling score set according to the homologous distribution and the non-homologous distribution to form a log-likelihood ratio set specifically includes: The log-likelihood ratio corresponding to the nth sampling score is calculated using the following formula: Among them, x n Represents the nth sampling score, n∈[1,N], m n Represents the log-likelihood ratio corresponding to the nth sampling score, G s represents homologous distribution, G d represents non-homologous distribution, G s (x n ) represents the value corresponding to the nth sampling score in the homologous distribution, G d (x n ) represents the value corresponding to the nth sampling score in the non-homologous distribution; the log-likelihood ratio set is shown as M = {m1, m2, m3, ..., m N }.

5. The voiceprint comparison method based on log-likelihood ratio calibration according to claim 1, characterized in that: Calculating a log-likelihood ratio calibration expression based on the sampling score set and the log-likelihood ratio set specifically includes: Each sampling score x in the sampling score set is used as an independent variable in the log-likelihood ratio calibration expression, and each log-likelihood ratio LLR in the log-likelihood ratio set is used as a dependent variable in the log-likelihood ratio calibration expression, and the slope α and intercept β in the log-likelihood ratio calibration expression LLR=αx+β are calculated according to the least squares method.

6. The voiceprint comparison method based on log-likelihood ratio calibration according to claim 1, characterized in that: The value range of the log-likelihood ratio is from negative infinity to positive infinity. When the log-likelihood ratio LLR corresponding to the speech data to be detected and the sample speech data is less than 0, it indicates the degree to which the speech data to be detected and the sample speech data are from different sources; when the log-likelihood ratio LLR corresponding to the speech data to be detected and the sample speech data is greater than 0, it indicates the degree to which the speech data to be detected and the sample speech data are from the same source.

7. A voiceprint comparison device based on log-likelihood ratio calibration, characterized in that: include: A data acquisition module is configured to collect voice data of multiple persons in a target scene and form a voice data set, wherein at least two voice data are collected for each person; The voiceprint comparison module is configured to compare the voiceprints of every two voice data of the same person in the voice data set to obtain a first comparison score, and all the first comparison scores constitute homologous data; compare the voiceprints of every two voice data of different people in the voice data set to obtain a second comparison score, and all the second comparison scores constitute non-homologous data; The expression calculation module is configured to calculate homologous distribution and non-homologous distribution respectively based on the homologous data and non-homologous data; calculate N sampling scores according to the maximum and minimum values ​​in the homologous data and non-homologous data to form a sampling score set; calculate the log-likelihood ratio corresponding to each sampling score in the sampling score set according to the homologous distribution and non-homologous distribution to form a log-likelihood ratio set; calculate the log-likelihood ratio calibration expression based on the sampling score set and the log-likelihood ratio set; The similarity determination module is configured to obtain the voice data to be detected and the sample voice data in the target scene and perform voiceprint comparison to obtain the corresponding comparison score, calculate the log-likelihood ratio value corresponding to the voice data to be detected and the sample voice data according to the log-likelihood ratio calibration expression, and determine the similarity between the voice data to be detected and the sample voice data according to the log-likelihood ratio value corresponding to the voice data to be detected and the sample voice data.

8. An electronic device, comprising: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.