An authentication system and method
Patent Information
- Application Number
- GB2024000894
- Authority / Receiving Office
- GB · GB
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-23
- Publication Date
- 2025-07-30
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
FIELD Embodiments described herein relate to an authentication system and an authentication method. BACKGROUND Authentication may comprise a process of verifying that an assertion, such as a user identity or other information such as a user location, is true. Authentication systems and methods are used in various fields, including healthcare and banking for example. For example, an authentication system can be used to verify a user location. In healthcare applications, authentication systems can be used to authenticate that both a carer and a patient are in the patient’s home for example, as part of a fraud detection system. The patient and carer location could be verified as being the patient’s home using GPS data from a smartphone for example. There is a continuing need for improved authentication methods and systems. BRIEF DESCRIPTION OF FIGURES Systems and methods in accordance with non-limiting embodiments will now be described with reference to the accompanying figures in which: Figure 1 shows a schematic illustration of a system in accordance with an embodiment; Figure 2(a) shows a schematic illustration of an audio stream; Figure 2(b) shows a flow chart of an authentication method in accordance with an embodiment; Figure 3 shows a schematic illustration of a method of determining the number of speakers which is used in the method of Figure 2(b) according to an embodiment; Figure 4 shows a schematic illustration of a method of speaker diarization which is used in the method of Figure 3 according to an embodiment; Figure 5 shows a schematic illustration of a method of determining acoustic condition information which is used in the method of Figure 2(b) according to an embodiment; Figure 6 shows a flow chart of a method of authentication in accordance with an embodiment; and Figure 7 shows a flow chart of a method of authentication in accordance with an embodiment. DETAILED DESCRIPTION According to a first aspect, there is provided an authentication system comprising: an input module configured to obtain an input audio signal; one or more processors, the one or more processors configured to: determine acoustic condition information for a first audio signal; compare the acoustic condition information for the first audio signal with acoustic condition information for a second audio signal to determine an indication of proximity; and an output module configured to provide an output based on the indication of proximity. The disclosed system provides an improvement to computer functionality by allowing computer performance of a function not previously performed by a computer. Specifically, the disclosed system provides for the authentication of the spatial proximity of a first speaker in relation to something, for example in relation to a second speaker or a location of the first speaker for a previous recording. The method achieves this by determining acoustic condition information from a first audio signal corresponding to the first speaker. The acoustic condition information is then compared with acoustic condition information for a second audio signal to determine an indication of proximity. The method enables the proximity of the first speaker to a location corresponding to the second audio signal to be authenticated, regardless of the location. Furthermore, the disclosed system addresses a technical problem tied to computer technology and arising in the realm of computer networks, namely the technical problem of resource utilization. The disclosed system solves this technical problem by performing authentication using an audio signal. In particular, the system performs authentication of the proximity of a user in relation to something using an audio signal. By using the audio signal, it is not required to transmit additional location information. The disclosed system may be used in the field of healthcare for example, for example to authenticate that a carer has visited a patient, regardless of the patient location. This can improve fraud detection processes. In one example, determining the indication of proximity comprises determining a difference between the acoustic condition information of the first signal and the acoustic condition information of the second signal, and determining whether the difference meets a pre-determined criteria. In one example, the first audio signal comprises speech from a first speaker; and the second audio signal comprises speech from a second speaker. In one example, the processor is configured to extract the first audio signal and the second audio signal from an input audio signal using a speaker diarization process. In one example, determining acoustic condition information comprises extracting a measure of at least one of: room reverberation, recording channel, and background noise. In one example, the first audio signal comprises speech from a first speaker; and wherein the one or more processors are further configured to: biometrically authenticate the first speaker voice against a stored template. In one example, the second audio signal comprises speech from a second speaker; and wherein the one or more processors are further configured to: biometrically authenticate the second speaker voice against a stored template. In one example, the one or more processors are further configured to: authenticate that the first audio signal is not a replay of a recording. In one example, the first audio signal comprises speech from a first speaker, and wherein the one or more processors are further configured to: authenticate that the speech is not computer generated. In one example, the speaker diarization process comprises performing a speech activity detection process and wherein the first audio signal comprises speech from a first speaker, wherein the one or more processors are further configured to: biometrically authenticate a first speaker voice against a stored template using an output of the speech activity detection process. In one example, the biometric authentication comprises a step of channel compensation, and wherein an output from the channel compensation step is used to determine the acoustic condition information. In one example, the biometric authentication comprises a step of background noise remove, and wherein an output from the background noise removal step is used to determine the acoustic condition information. According to another aspect, there is provided an authentication method comprising: obtaining, by way of an input, an input audio signal; determining acoustic condition information for a first audio signal; comparing the acoustic condition information for the first audio signal with acoustic condition information for a second audio signal to determine an indication of proximity; and providing an output based on the indication of proximity. According to another aspect, there is provided a carrier medium comprising computer readable code configured to cause a computer to perform any of the above methods. The methods are computer-implemented methods. Since some methods in accordance with embodiments can be implemented by software, some embodiments encompass computer code provided to a general purpose computer on any suitable carrier medium. The carrier medium can comprise any storage medium such as a floppy disk, a CD ROM, a magnetic device or a programmable memory device, or any transient medium such as any signal e.g. an electrical, optical or microwave signal. The carrier medium may comprise a non-transitory computer readable storage medium. According to a further aspect, there is provided a carrier medium comprising computer readable code configured to cause a computer to perform any of the above-described methods. Authentication may comprise a process of verifying that an assertion, such as a user identity or other information such as a user location, is true. Authentication systems and methods are used in various fields, including healthcare and banking for example. For example, an authentication system can be used to verify a user location. In healthcare applications, authentication systems can be used to authenticate that both a carer and a patient are in the patient’s home for example, as part of a fraud detection system. For example, in the US there exists healthcare legislation known as the 21st Century Cures Act. One of the requirements of the Act is Electronic Visit Verification (EVV), which is designed to stop Medicaid fraud. A fraud occurs when a carer submits a Medicaid claim for a home care visit which never actually occurred. EVV is therefore designed to verify that the carer does actually perform the visit by verifying their location at the patient’s home. Other requirements are verification of the identity of the carer and patient. Various systems attempt to verify the carer’s location using GPS information from smartphone apps, inbound phone calls to an Interactive Voice Response (IVR) system from the patient’s landline phone or mobile phone, or calls to an IVR system that require the caller to speak a passcode displayed on a passcode generator kept at the patient’s home. However, EVV is still required even if the care either commences or finishes in the patient’s home but continues or commences elsewhere, such as a community centre. IVR solutions that are configured for use only in the patient’s home therefore cannot be used in some instances. In one example, a method comprising obtaining an input audio signal, extracting a first input audio signal corresponding to a first speaker and a second input audio signal corresponding to a second speaker and comparing acoustic condition information for the first input audio signal with acoustic condition information for the second audio signal to determine an indication of proximity is performed. The method authenticates the physical location of the first and second speaker, who may be the carer and patient for example, as being in the same place. There is no requirement for knowing where that physical location is. The method may further biometrically authenticate one or both of the first and second speaker within the same authentication session, i.e. from an authentication transmission that contains audio from both parties, that can be determined to have been transmitted from the same physical location within the same timeframe. The biometric authentication is used to verify the actual identity of the speakers. Some alternative examples may authenticate that a user is in the same location as during a previous interaction. In such examples, the first signal and the second signal are separate audio signals. For example, the second signal may be a pre-recorded signal from a first speaker which is retrieved from memory, and the first signal may be an input audio signal received from a user device. An authentication method is then performed to verify that the input audio signal is received from the location corresponding to that in which the pre-recorded signal was recorded. The method comprises comparing the acoustic condition information for the first input audio signal with acoustic condition information for the second audio signal which is retrieved from memory to determine an indication of proximity. Again, the method authenticates the physical location of the user as being in the same place as the user was previously, without a requirement for knowing where that physical location is. In such applications, the location of a user in relation to something is required to be authenticated prior to executing further processes, such as validating an insurance claim for example. The location of a user in relation to something, for example another person, is also referred to here as the spatial proximity. The below described methods determine whether a first speaker is in the same location as something else, by performing acoustic proximity authentication, also referred to here as voice biometric proximity authentication. By the same location, it is meant that the first speaker and the something else (e.g. second speaker) are close to each other, for example in the same room. Figure 1 is a schematic illustration of a system 900 in accordance with an embodiment. The system comprises an input 901, a processor 905, a working memory 911, an output 903, and long term storage 907. In this example, the system 900 is a server device that receives an input signal originating from a user device 910. The user device 910 comprises a microphone (not shown) which generates an audio signal. The audio signal is then sent to the system 900 from the user device 910 through a communication network. The user device 910 may be a smart device, which transmits the audio signal via the Internet. The user device 910 may be a telephone, which transmits the audio signal through the telephone network to a third device, which in turn transmits the audio signal to the server device 900 via the Internet. The input signal may be received at the system 900 via one or more further intermediate devices or systems. The audio signal is received at the input 901 of the system 900. The input 901 is a receiver for receiving data from a communication network. The processor 905 accesses the input module 901. The processor 905 is coupled to the storage 907 and also accesses the working memory 911. The processor 905 may comprise logic circuitry that responds to and processes the instructions in code stored in the working memory 911. In particular, when executed, a program 909 is represented as a software product stored in the working memory 911. Execution of the program 909 by the processor 905 will cause embodiments as described herein to be implemented. The program 909 may comprise a proximity authentication module which embodies a proximity authentication method such as described in relation to the figures below. The program 909 may further comprise code implementing other methods as described herein. The processor 905 is also configured to communicate with the non-volatile storage 907. As illustrated, the storage 907 is local memory that is contained in the system 900. Alternatively however, the storage 907 may be wholly or partly located remotely from the system 900, for example, using cloud based memory that can be accessed remotely via a communication network (such as the Internet). The program 909 is stored in the storage 907. The program 909 is placed in working memory when executed. The processor 905 also accesses the output module 903. The output module 903 provides the response generated by the processor 905 to a communication network. The input and output modules 901, 903 may be a single component or may be divided into a separate input interface 901 and a separate output interface 903. As illustrated, the system 900 comprises a single processor. However, the program 909 may be executed across multiple processing components, which may be located remotely, for example, using cloud based processing. For example, the system 900 may comprise at least one graphical processing unit (GPU) and a general central processing unit (CPU), where various operations described in relation to the methods below are implemented by the GPU, and other operations are implemented by the CPU. Usual procedures for the loading of software into memory and the storage of data in the storage unit 907 apply. The program 909 can be embedded in original equipment, or can be provided, as a whole or in part, after manufacture. For instance, the program 909 can be introduced, as a whole, as a computer program product, which may be in the form of a download, or can be introduced via a computer program storage medium, such as an optical disk. Alternatively, modifications to existing software can be made by an update, or plug-in, to provide features of the below described embodiments. In the described example, the system comprises a server device which receives an audio signal from a user device 910. However, alternatively, the system 900 may be an end-user computer device, such as a laptop, tablet, smartwatch, or smartphone. In some examples, the program 909 comprising the proximity authentication module is executed on the same device which records the sound. In such a system 900, the processor 905 accesses the input module 901, which comprises a microphone. The output module 903 provides the response generated by the processor 905 to an output such as a speaker, a screen, or a transmitter for transmitting data to an external device through a network, for example. The output may comprise an audible message that is played on a speaker, a message that is displayed to the user on a screen, or data that is transmitted over a communication network. It will also be appreciated that in some the examples, parts of the program 909 may be executed on a user device 910 whilst other parts of the program may be executed on a server device 900, with data being transmitted between the two devices. While it will be appreciated that the below embodiments are applicable to any computing system, the example computing system 900 illustrated in Figure 1 provides means capable of putting an embodiment, as described herein, into effect. For example, the system in Figure 1 may perform the method of Figure 2(b), or the method of Figure 6. As is described in more detail below, in use, the system 900 receives, by way of input 901, an audio file. The proximity authentication module 909, executed on processor 905, determines acoustic condition information, compares acoustic condition information of a first and second signal, determines an indication of proximity, and provides an output based on the indication of proximity in the manner that will be described with reference to the following figures. The system 900 outputs data by way of the output 903. The analysis performed by the proximity authentication module 909 may be performed some time after the audio has been captured. For example, the analysis may be performed hours or days after the audio has been captured. Alternatively, the analysis is performed as the audio is captured, for example, the audio file is analysed at regular intervals as the audio stream is received. By performing analysis at regular intervals as the audio stream is received, the result of the analysis may be obtained in near realtime. For example, when the total captured audio has a duration of 5 minutes, the audio file may be analysed in sections of 30 seconds. Figure 2(a) shows a schematic illustration of an example audio file comprising an audio stream that is created when two speakers speak during a single capture session. The capture session might involve two speakers speaking during a single capture session of a software application, or during a single telephone call. One capture session results in a single continuous audio file. Where a user is trying to commit fraud, an audio file may be obtained by taking separate recordings from different capture sessions and merging them together. The recordings may be from the same device at different times or from different devices. The steps performed in the following method are intended to detect such audio signals. The audio file shown in Figure 2(a) comprises periods of speech from two speakers and a period of silence. Sound resulting from the acoustic conditions is also present throughout the audio signal, as shown in Figure 2(a). As shown in Figure 2(a), a first speaker (speaker 1) speech occurs from a time to to a time ti and a second speaker (speaker 2) speech occurs from a time t2 to a time to. From ti to t2, the audio stream comprises silence, i.e. no speech. The speaker 1 speech comprises the phrase “I am here” spoken by a first speaker while the second speaker speech comprises the phrase “I am here too” spoken by a second speaker. The first speaker and the second speaker are captured in a single capture session. The time frame of the capture session is selected to be suitable for proximity authentication purposes. For example, the capture session is set to be long enough that the two speakers have sufficient time to each utter a message. In one example, the capture session is 10 seconds or more. In another example, the capture session is 30 seconds or more. From the time to to the time to the audio stream comprises sound resulting from the acoustic conditions. The sound resulting from the acoustic conditions comprises one or more of: background noise or noise floor; sound from reverberation; and sound from channel characteristics. Background noise or noise floor may arise due to sources of noise other than the speakers, such as traffic, people, equipment such as computers or ventilation, or weather (e.g. rain). Reverberation sound arises due to the reflection of the speech from the surfaces of the environment, for example an enclosed room. Channel characteristic sounds arise from factors such as telephony or codec. For example, the acoustic characteristics of a signal obtained using a landline phone will be different to a signal obtained using a GSM (Global System for Mobile communications) phone. The segment of the audio signal from time to to ti is referred to here as the first signal. The first signal comprises the first speaker (speaker 1) speech together with the sound from the acoustic conditions that occurs from time to to ti. The segment of the audio signal that occurs from time t2 to to is referred to as the second signal. The second signal comprises the second speaker (speaker 2) speech together with the sound from the acoustic conditions that occurs from time t2 to ta. The audio stream of Figure 2(a) is collected by a user device. The user device comprises a microphone and is able to receive audio. The user device may be a telephone. Alternatively, the user device may be a smart device which runs a software application which records audio for example. Although the audio stream illustrated in Figure 2(a) comprises silence (from a time ti to tz), it will be understood that this period of silence may be absent from the audio file. It will also be understood that in some examples, the first speaker and the second speaker may speak at least partly at the same time, such that a time segment may comprise both first speaker and second speaker speech. Although Figure 2(a) illustrates an example with two speakers, it will be understood that in some examples, the audio file may comprise more than two speakers. Figure 2(b) shows a flow chart of an authentication method according to an embodiment. In step S11, an audio file is obtained. The audio file corresponds to that described in relation to Figure 2(a) in this example. As described above, the audio file may be obtained from a user device, such as a laptop, tablet, smartwatch, landline or mobile phone. In step S12, the audio file is analysed to determine the number of speakers. In this example there are two speakers. A first signal and a second signal are then extracted from the audio file. How the first and second signals are extracted from the audio file will be described further below in relation to Figures 3 and 4. In step S13, acoustic condition information of the first and second signals is then obtained. The first signal comprises the speech spoken by first speaker together with the sound from the acoustic conditions. Similarly, the second signal comprises the speech spoken by the second speaker together with the sound from the acoustic conditions. How the acoustic condition information is obtained will be described further below in relation to Figure 5. In step S15, the acoustic condition information of the first signal is compared with the acoustic condition information of the second signal. The comparison is described further below in relation to Figure 5. In step S17, an indication of the proximity of the two speakers that spoke during the capture session is determined from the comparison of the two acoustic conditions. In S17, an indication of the spatial proximity of where the first signal was recorded to where the second signal was recorded is determined. This indication is also referred to as a proximity parameter. The proximity parameter may have values such as True or False, or Yes or No, or “1” or “0”. For example, a proximity value of “True”, “Yes” or “1” indicates that the speakers are in the same location. In step S19, an output is provided based on the indication of the proximity determined in S17. The output may comprise transmitting a message to another computing device, for example, the output may comprise transmitting a message to the user device 910 or to another server device indicating whether the speakers are authenticated as being in the same location. The client device 910 or server device may execute further processes in response to the provided output. The output step of S19 may additionally or alternatively comprise displaying a message on a screen or outputting an audio message. The authentication system may be used in an insurance fraud detection system for example. If the proximity parameter generated in S17 indicates that the speakers are not in the same spatial location (e.g. FALSE) then the output step S19 may comprise sending a message to a fraud detection server indicating that an insurance claim should be queried. In the above described example, audio signal analysis is performed to determine that the acoustic conditions associated with each individual speaker are the same. In some examples, the same signal analysis is performed to determine that the acoustic conditions associated with each individual speaker and the silence segment are the same. This is done by checking changes on the recording of reverberation, recording channel and / or background noise between the speaker segments and the silence segment. The method described in relation to Figure 2(b) includes performing analysis of an audio signal in order to determine whether two or more speakers are in the same physical location, regardless of where that physical location is, through acoustic proximity authentication. Figure 3 shows a flow chart of a process performed in step S12 of Figure 2(b) in accordance with an embodiment. This process comprises the steps of speaker diarisation S131 and signal determination S133. In step S131, Speaker Diarization (SD) is performed. The purpose of SD is to determine which individual spoke at which times during the received audio signal. The audio signal obtained in S11 of Figure 2(b) comprises speech from a plurality of individual speakers that have spoken at different times. SD may be understood as a process of partitioning the input audio signal into groups of signal segments, such that signal segments corresponding to speech from the same speaker are grouped together. An example process of SD is described in relation to Figure 4 below. SD may be used to arrange the audio stream into speaker turns. A speaker turn refers to a period between the point at which an individual speaker starts to talk and the point at which said speaker stops. In this example, the output of the SD step in S131 comprises a sequence of start and end times, with speaker labels for each turn. For example, the output of S131 for the input audio signal shown in Figure 1 may comprise the following list of speaker turns: Start time Stop time Label to ti Speaker 1 t2 t3 Speaker 2 In Step 133, a first speaker signal comprising one or more speaker 1 turns, and a second speaker signal comprising one or more speaker 2 turns, are generated from the start and end times and the speaker labels output from S131. For example, referring to Figure 2(a), a speaker 1 turn occurs from time to to ti, while a speaker 2 turn occurs from t2 to to. From the timings of the speaker 1 and speaker 2 turns, the audio file is segmented into a first signal and a second signal. In this example, the first signal comprises the segment of the audio stream from to to ti, while the second signal comprises the segment of the audio stream from t2 to to. In an alternative example however, the first signal comprises segment of the audio stream from to to t2 (such that it comprises the speaker 1 signal as well as the silence), while the second signal comprises the segment of the audio stream from b to to. Although Figure 2(a) shows an example in which the input audio signal comprises only one turn for speaker 1 and one turn for speaker 2, it is possible that the audio stream comprises multiple turns for speaker 1 and / or speaker 2. In this case, in one example, the first signal comprises all of the segments of the audio stream corresponding to the first speaker turns, and the second signal comprises all of the segments of the audio steam corresponding to the second speaker turns. In other words, the audio segments corresponding to the first speaker turns are aggregated to form the first signal. The audio segments corresponding to the second speaker turns are aggregated to form the second signal. The aggregation of the speaker turns may provide higher confidence in the determined acoustic condition information for each of the first and second signal. Alternatively, the first signal comprises only one first speaker turn from the one or more first speaker turns, and the second speaker signal comprises only one second speaker turn from the one or more second speaker turns. A second speaker turn that is adjacent in time to the first speaker turn may be selected, so that the first and second signals comprise any two consecutive turns. Although Figure 1 shows an example where the first speaker and second speaker speak at different times, it is also possible that the audio stream comprises portions where the two speakers speak at the same time, such that the two speakers partly overlap. In this case, optionally, the overlapped part of the signal may be discarded and only instances where the two speakers do not speak at the same time are then used to generate the first signal and the second signal. Figure 5 shows an illustration of an example process for determining acoustic condition information for the first signal in S13 of Figure 2(b) in accordance with an embodiment. The same process can be used to determine the acoustic condition information for the second signal in S13. In this example, the acoustic condition information is generated based on information relating to: background noise, reverberation and channel characteristics. The first signal is taken as input to S501. A measure of the background noise is then determined in S501. In this example, a measure of the background noise is determined by applying at least a part of a speech enhancement algorithm. However, it will be appreciated that other methods of obtaining a measure of the background noise can be used. A speech enhancement algorithm takes as input a noisy speech signal, and outputs an estimated clean speech signal. A speech enhancement algorithm may generate an estimated noise magnitude spectrum, an estimated noise power spectrum or an estimated noise power spectral density as part of the speech enhancement process. The estimated noise magnitude spectrum, estimated noise power spectrum or estimated noise power spectral density may be taken as the measure of background noise in S501. The parts of the speech enhancement algorithm that are not relevant to estimation of the background noise (e.g. not relevant to the estimation of the estimated noise magnitude spectrum) need not be performed in S501. An example speech enhancement algorithm is based on spectral subtraction. As part of a spectral subtraction algorithm, a noise magnitude spectrum is estimated. In one example, a noise magnitude spectrum estimated using at least part of a spectral subtraction algorithm can be taken as the measure of background noise in S501. The parts of the spectral subtraction algorithm that are not relevant to estimation of the noise magnitude spectrum need not be performed in this step. An example method of estimating the noise magnitude spectrum will now be described. The first signal is windowed. An example window length may be in the range of 20-40ms. Overlapping windows may be used. A magnitude spectrum for each window is generated using a Fast Fourier Transform. Windows corresponding to speaker silence are then used to estimate a noise magnitude spectrum. Windows corresponding to speaker silence can be identified using a voice activity detection algorithm for example. In some examples, a statistical approach may be used to identify and / or combine the windows corresponding to speaker silence. For example, a minimum of the magnitude spectra is taken as the noise magnitude spectrum, or an average of multiple magnitude spectra which are identified as corresponding to pauses in speech is taken as the noise magnitude spectrum. The measure of background noise output in S501 in this example is an estimated noise magnitude spectrum, which may be in the form of a vector of magnitude values corresponding to a set of frequency components. An alternative speech enhancement algorithm that may be used in S501 is based on Weiner filtering. Alternatively, the measure of background noise can be obtained from the estimated clean signal output from the speech enhancement algorithm and the input noisy speech signal. For example, this may involve windowing the input and output signals, performing a Fast Fourier Transform for each window, subtracting the clean signal magnitude spectrum from the noisy signal magnitude spectrum for each window, and taking a statistical combination of the resulting magnitude spectra (for example an average) to give an estimated noise magnitude spectrum. The first signal is also taken separately as input to S502. In S502, a measure of the reverberation is determined. In this example, the measure of the reverberation is determined by determining information about the reverberation time using a blind estimation procedure. However, it will be appreciated that other methods of obtaining a measure of the reverberation can be used. Reverberation is caused by reflections of sound from various surfaces in a room. A difference in the geometry and objects in a room will lead to different reverberation times. In this example, the reverberation time is defined as the time taken for a sound to decay 60 dB below the initial level. The blind estimation procedure does not require information about the room geometry or sound source to be known. In the blind estimation framework, an audio signal is considered to comprise the following components: 1) the direct sound, 2) an early reverberation, and 3) a late reverberation. The late reverberation (also referred to as the reverberant tail) provides information about the acoustic environment. The decaying envelope of the late reverberation is characterised by a time constant t. In more detail, the signal is considered as being described by: y[n] = d[n]x[n] + q[n], n> 0, where n[n] is noise; d[n] is a multiplicative decay, wherein d[n] = exp(-n / T); n is the sample number; and x[n] is the reverberant tail. The time constant t is linearly proportional to the reverberation time, and is a measure of the reverberation. In this example, the measure of the reverberation determined in S502 is the value of the time constant t. In one example, the value of the time constant t is obtained using a maximumlikelihood estimation (MLE) procedure. An example of a MLE based procedure is described in Malik, H., 2013. Acoustic environment identification and its applications to audio forensics. IEEE Transactions on Information Forensics and Security, 8(11), pp. 1827-1837, the entire contents of which are incorporated by reference herein. Using a maximum-likelihood estimation (MLE) procedure, the time constant, t, of the decaying speech sound is obtained from the first signal. A probability density function, Py[n] is defined for y[n], A likelihood function (to measure the goodness of fit of the model to the sample data) is also defined. The MLE procedure comprises maximising the log-likelihood function i.e. setting the partial derivatives of the function to zero, and solving for t. Estimates for a value of t may be obtained using iterative methods. For example, a value of t is obtained when consecutive estimates differ by less than a predetermined threshold. In an alternative example, the value of the time constant t is obtained using a procedure based on automatic decaying tail selection. A procedure based on automatic decaying tail selection is described in Malik, H., 2013. Acoustic environment identification and its applications to audio forensics. IEEE Transactions on Information Forensics and Security, 8(11), pp. 1827-1837, the entire contents of which are incorporated by reference herein. The first signal is also taken as input to S503. In S503, a measure of the recording channel is determined. In step S503, a measure of the recording channel characteristics, relating to telephony codec, is determined. The codec is the device(s) and / or computer program(s) which encode and decode the audio signal. For example an audio signal captured using a carbon button handset will have different channel characteristics from a signal captured using an electret handset or a cell phone. In this example, the information extracted in S503 described below captures information relating to all variations between recordings. Thus the information extracted in S503 includes other acoustic condition information, such as environment conditions (such as reverberation), as well as the channel characteristics arising from the codec. Various methods can be used to extract a measure of the channel characteristics. For example, a first determination method comprises transforming the first signal to a k-dimensional representation from which the channel characteristics are determined. A second determination method comprises extracting a sequence of feature vectors from the first signal and classifying them. A third determination method uses an electric network frequency (ENF) based method. It will be appreciated that other methods of obtaining a measure of the channel characteristics may be used. The first determination method of determining channel characteristics of the first signal will now be described. In this method, the speech signal (the first signal) is transformed to a k dimensional representation, in other words to a vector of length k, where k is a positive integer. Various transformations may be used to transform the first signal to a k-dimensional representation. This vector representation may be a GMM mean supervector, an i-vector or an x-vectorfor example. An example process for generating a GMM mean supervector is described in “Advances in channel compensation for SVM speaker recognition”; Solomonoff et al.; Proceedings. (ICASSP '05). IEEE International Conference on Acoustics, Speech, and Signal Processing, 2005, the entire contents of which are incorporated by reference herein. In this process, a training corpus comprising utterances, each labelled with a speaker identity label and a channel label, is used. An example method for generating an i-vector is described in “Variance-spectra based normalization for i-vector standard and probabilistic linear discriminant analysis”, Bousquet et al, June 2012, In Speaker Odyssey, the entire contents of which are incorporated by reference herein. Another example method for generating an i-vector is described in Dehak, N., Kenny, P.J., Dehak, R., Dumouchel, P. and Ouellet, P., 2010. “Front-end factor analysis for speaker verification”. IEEE Transactions on Audio, Speech, and Language Processing, 19(4), pp.788-798, the entire contents of which are incorporated by reference herein. In these examples, the Total Variability matrix is trained such that the i-vector contains the channel information. An example method for generating an x-vector is described in X-Vectors: Robust DNN Embeddings for Speaker Recognition; Snyder et al.; 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), the entire contents of which are incorporated by reference herein. In these examples, the DNN training process is modified such that the channel information is kept in the x-vector. Once obtained, the vector representation (i.e. the GMM mean supervector, the i-vector or the x-vector) is decomposed into multiple components representing the channel information. For example, an i-vector is decomposed into multiple components, where the one or more components represent the channel characteristics. The one or more components are output from step S503 as the channel characteristics. In an example, a stored matrix is used to decompose the i-vector into one or more components corresponding to the channel characteristics. In this example, the first signal is transformed to an i-vector as has been described above. The i-vector corresponding to the first signal is represented by the symbol 0xi in the below example. A stored matrix S is then used to generate the channel characteristics from the i-vector. The stored matrix S may be generated in advance of the method performed in Figure 5 from a data set of stored recordings. An example of how the stored matrix S is generated will be first be described. Given a stored data set of recordings in their expansion form: 'Ur - -<Pl,sNs> - <PhNs,sNs from Ns different speakers, each with hi different sessions over different channels. Here, Ns is the number of speakers, where the subscript Si indicates the recording corresponds to the first speaker and the subscript Sns indicates the recording corresponds to the Nsth speaker. For speaker 1, there are hi different recordings (different sessions). The subscript 1, Si indicates the first recording session for the first speaker and the subscript hi,si indicates the hith recording session for the first speaker. For speaker Ns, there are hNs different recordings sessions. The subscript 1, Sns indicates the first recording session for the Nsth speaker and the subscript hNs.SNs indicates the IWh recording session for the Nsth speaker. 01S1 is the i-vector generated for the first recording session and for the first speaker. is the i-vector generated for the hith recording session and for the first speaker. 0i,SjVs. is the i-vector generated for the first recording session and for the Nsth speaker. (phNs,sNs is the i-vector generated for the IWh recording session for the Nsth speaker. For each speaker Si, the average expansion for each speaker is calculated and then removed from all the corresponding examples: ^1,Si= <Pl,Si (p Si for I e [1 hj, where is the average expansion for the speaker Sj. The following matrix: M = ... ^hl,S1,... ... hNs,sNs] is formed by taking each vector $ to form a column of the matrix. Each vector has length k, where k is also the length of the i-vector. The matrix M has dimension k*N, where N = hi + ... + hNs ■ The matrix M represents all the intersession variations from the average speaker positions. To identify the subspace of dimension K where the variations are the largest, the K eigenvectors with the highest eigenvalues of the corresponding covariance matrix MMT are then calculated. This eigenvalue problem corresponds to a principal component analysis (PCA) on supervector intersessions differences. The resulting eigenvectors are orthogonal and normalised. They form a base to the subspace where channel variations are the most pronounced and are concentrated into a matrix S of size k*K. This matrix S is the stored matrix that is used to generate the channel information from the i-vector corresponding to the first signal in S503. For a channel estimation performed in S503, this matrix S is used on utterance Xi (the first signal) to keep the component within the channel subspace: 0Chanx = SST <t>xi. Thus in S503, the i-vector <bx corresponding to the first signal and the stored matrix S are used to generate 0Chanxi - the component(s) corresponding to the channel characteristics. <t>Chanxi is output from step S503 as the channel characteristics of the first signal. Although a specific example has been described here, other methods, with the same objective to extract a vector characteristic of the channel, can be used for this estimation. Furthermore, examples of how the channel characteristics are determined from a vector representation of the first signal have been described. However, it will be appreciated that other methods of obtaining a measure of the channel characteristics can be used. In S503, a second determination of the recording channel may be performed in addition to or alternatively to the first determination of the channel characteristics. In the second determination, a sequence of feature vectors is extracted from the first signal. The feature vectors may be extracted in the same manner disclosed in relation to the first determination. A single vector representation of the first signal may be obtained using a first trained model which takes the sequence of feature vectors as input. The single vector representation is then taken as input to a second trained model which is a classifier. In some alternative examples, the sequence of feature vectors may be taken directly into the classifier model, for example in sequence or in a combined form. The second trained model is a multi-class classifier, where each output class corresponds to a recording channel type. For example, a first output class may correspond to a landline, a second output class may correspond to a GSM, and so on. The model(s) may be trained using training data comprising utterances from many different speakers, and labelled with the channel class. For example, the second model may be a DNN. In S503, a third determination using an electric network frequency (ENF) based method may be performed in addition to or alternatively to one or both of the first determination and the second determination. When equipment (e.g. a microphone) is used to record a conversation, the 50 / 60 Hz electric network frequency (ENF) is captured together with the intended conversation. The 50 / 60 Hz ENF signal that is captured on audio recordings is sometimes referred to as the background mains hum. The ENF at an instant in time may be written as f = [fo ± Af\ Hz. fo is the mains center frequency and is 50 or 60 Hz. The term Af arises due to variations in energy supplied by an electric network. The mains hum signal may be treated as a time-dependent digital watermark from which acoustic condition information may be derived. ENF based methods comprise the following steps: • Optionally, the signal is down sampled to 120 Hz. • Apply a band-pass filter. For example, when fo is 50 Hz, the band-pass filter is a Butterworth filter that has a lower cut-off frequency of 49 Hz, a higher cut-off frequency of 51 Hz, and a slope of 24dB / Octave. • Compute a spectrogram by applying a Fast Fourier Transform (FFT). For example, the FFT is a 4096-point FFT. The computed spectrogram may then be compared with a reference spectrogram. The reference spectrogram may be obtained from historical records kept by the electricity suppliers. The timing of the reference spectrogram corresponds to the timing of the computed spectrogram. By comparing the computed spectrogram with the reference, it can be determined whether the recording was made by a device connected to a particular electric network. Examples of how the channel characteristics are determined have been described. However, it will be appreciated that other methods of obtaining a measure of the channel characteristics can be used. Returning to Figure 2(b), in this example, the comparison of the acoustic condition information in step S15 comprises performing a separate comparison for each of the background noise, the reverberation and the channel characteristics. The comparison of the acoustic condition information in step S15 comprises the determination of the similarity between the acoustic condition information of the first signal and second signal, for example. The similarity may be determined using any suitable metric. Examples of suitable metrics include Pearson correlation, cosine similarity or distance based metrics such as Euclidean distance. The Pearson correlation takes as input two quantitative continuous variables and outputs a coefficient in the range of -1 to 1. Supposing the two inputs are vectors x, and y, the coefficient is [Z(Xi-x’)(yry’)] / [x Z(x,-x’)2x > / Z(yry’)2], where x’ = 1 / n x ZnXj and y’ = 1 / n x Znyi. Xi represents the ith element of vector x. The coefficient is compared to a predetermined threshold value and when the coefficient is higher or lower than said value, the two vectors x and y are considered to be similar. The cosine similarity takes as input two vectors x and y and an output equal to x y / (||x||x||y||). The output has values between -1 and 1. The coefficient is compared to a predetermined threshold value and when the coefficient is higher than said value, the two vectors x and y are considered to be similar. The Euclidean distance is a distance between two vectors, x and y. This can be computed as [Z(xi-yi)2]1 / 2. The smaller the distance, the more similar the vectors are. The distance is compared to a predetermined threshold value and when the coefficient is smaller than said value, the two vectors x and y are considered to be similar. For example, the measure of background noise output in S501 may be an estimated noise magnitude spectrum, which may be in the form of a vector of values. A noise magnitude spectrum is determined for the first signal and the second signal in S13. These vectors are compared to one another and their similarity determined using a similarity measure as described above. A pre-determined threshold is then applied to give an output. If the similarity score is greater than or equal to the threshold, an output 1 of is given, indicating that the background noise is the same between the first and second signal. If the similarity score is less than the threshold, an output 0 of is given, indicating that the background noise is different between the first and second signal. For the measure of reverberation, the estimated time constants t for the first and second speaker signals determined in S502 may be compared to one another. For example, the difference between the time constant estimated for the first signal and the time constant estimate for the second signal is calculated. A threshold is applied to give an output. The threshold may be based on a standard deviation value for example. The threshold may be a pre-determined value. If the difference is greater than or equal to the threshold, an output 0 of is given, indicating that the reverberation is different between the first and second signal. If the difference is less than the threshold, an output 1 of is given, indicating that the reverberation is the same between the first and second signal. Where the measure of the recording channel determined in S503 comprises a measure of the recording channel characteristics, the recording channel characteristics are compared to one another and their similarity determined using a similarity measure as described above. For example, where the channel characteristics of the first signal extracted in S503 comprise components generated from the i-vector 4>Chanxi as described above and the channel characteristics of the second signal comprise components generated from the i-vector 0ChanX2 as described above, the similarity of channel can be estimated with a cosine distance: IpchanXl ' 4>chanX2 ll^c / ianX! II \\^chanX2 II The larger this value, the more likely the first signal and second signal come from the same channel. A pre-determined threshold is then applied to give an output. If the similarity score is greater than or equal to the threshold, an output 1 of is given, indicating that the channel characteristics are the same between the first and second signal. If the similarity score is less than the threshold, an output 0 of is given, indicating that the channel characteristics are different between the first and second signal. Where the measure of the recording channel determined in S503 comprises a second determination of the recording channel, the recording channel class output for the first signal and the second signal are compared. If the first signal and the second signal are determined to be in the same class, an output 1 of is given, indicating that the channel is the same between the first and second signal. If the first signal and the second signal are determined to be in a different class, an output 0 of is given, indicating that the channel is different between the first and second signal. Where the measure of the recording channel determined in S503 comprises a third determination using ENF based methods, the computed spectrogram from the first speaker signal may be compared with a reference spectrogram (from the electric network company). Any difference from the reference is then noted. This difference is referred to as the first ENF difference. Similarly, the computed spectrogram from the second speaker signal may be compared with a reference spectrogram. Again any difference is noted. This difference is referred to as the second ENF difference. The first and second ENF difference are then compared to one another and their similarity obtained using a similarity measure as described above. A pre-determined threshold is then applied to give an output. If the similarity score is greater than or equal to the threshold, an output 1 of is given, indicating that both speaker 1 and speaker 2 signals were obtained on devices connected to the same electric network. If the similarity score is less than the threshold, an output 0 of is given, indicating that the speaker 1 and speaker 2 signals were not obtained from devices connected to the same electric network, indicating that the speakers were not in proximity. Alternatively, the computed spectrogram obtained from the first speaker and from the second speaker can be compared directly to each other to detect whether the recording was made from the same device. For example, mains hum appears in audio recordings because of imperfect shielding or imperfect electronic systems. Thus, differences in the computed spectrograms may indicate that the signals were obtained on different devices. Where two or more comparisons are made for the recording channel, the output of the comparisons may be combined, such that if at least one outputs indicates a difference in recording channel, then the output for the recording channel comparison is 0 for example. In Step 17, it is determined whether any of the comparisons in S15 indicate that the speakers are not in the same location. In this example, if at least one of the outputs from S15 is 0, then the speakers are considered not to be in the same location. An indication of proximity is therefore generated to indicate that the speakers are not in the same location. This might comprise a “0”, “NO”, “N” or “FALSE” value for example, or some other indication. If all of the three outputs are 1, then the speakers are considered to be in the same location. An indication of proximity is therefore generated to indicate that the speakers are in the same location. This might comprise a “1”, “YES”, “Y” or “TRUE” value for example, or some other indication. Alternatively, the step S17 further comprises combining the outputs from S15 according to a rule to determine the indication of proximity. For example, the majority of the outputs may be considered. Although in the above described example, the acoustic condition information comprises information indicating the background noise, the reverberation and the channel characteristics, the acoustic condition information may comprise any one of these, or some combination of these. Additional information may also be included in the acoustic condition information. Figure 4 shows a flow chart of a process of speaker diarization that may be performed in step S131 shown in Figure 3 according to an embodiment. In this example, the speaker diarization is performed based on a “bottom-up” clustering approach. However, it will be understood that various methods of speaker diarization may be used in S131. In one example, speaker diarization is performed using the procedure described in, for example, Anguera, X., Bozonnet, S., Evans, N., Fredouille, C., Friedland, G. and Vinyals, 0., 2012. Speaker diarization: A review of recent research. IEEE Transactions on Audio, Speech, and Language Processing, 20(2), pp.356-370, the entire contents of which are incorporated by reference herein. In one example, speaker diarization is performed using the procedure described in, for example, PYANNOTE.AUDIO: NEURAL BUILDING BLOCKS FOR SPEAKER DIARIZATION, Herve Bredin et al, arXiv:1911.01255v1, 4 Nov 2019, the entire contents of which are incorporated by reference herein. In step S401, the audio file is received and initial processing is performed. Initial processing comprises the following steps: (i) noise reduction; and (ii) detection of speech segments with a speech activity detection (SAD) algorithm. Noise reduction is optional and may be omitted from S401. The noise reduction step may comprise filtering the input audio signal to reduce the noise. For example, the noise reduction step may comprise applying a Wiener filter. Speech activity detection (SAD) differentiates between segments of the audio file that comprise speech, and those that comprise non-speech. Referring to Figure 2 (a), nonspeech segments correspond to the “Silence” portion between ti and t2. The nonspeech segments may include sound from acoustic conditions, such as background noise. In this example, the SAD separates speech and non-speech segments by considering changes in the energy, spectrum, noise floor, or pitch of the audio signal together with a predetermined threshold. Alternatively, however the SAD algorithm may use a model trained on speech and non-speech data. For example, an SAD algorithm may comprise a classifier such as a Linear Discriminant Analysis (LDA) coupled with Mel Frequency Cepstrum Coefficients (MFCC), or Support Vector Machines (SVM). The remaining steps of the Speaker Diarization S131 process, namely S403, S405, and S407, are performed on parts of the audio file that comprise speech. Optionally, parts of the audio file that are marked as ‘non-speech’ are removed from the audio file before these steps are performed. In S403, cluster initialisation is performed. A predetermined number of clusters is formed in this step. The predetermined number is set at a value higher than an expected maximum number of speakers. For example, the number may be set at 10. The audio file is then split into fixed length frames. The number of frames exceeds the number of clusters. Each frame is randomly assigned to a cluster. Each cluster is then modelled by a Gaussian Mixture Model (GMM). For each cluster, a single GMM is trained on the frames in the cluster. In S405, the number of clusters is iteratively reduced by merging, with the goal being to obtain one cluster for each speaker. The two closest clusters are merged on every iteration. The two closest clusters are determined based on a distance metric. An example of a suitable distance metric is a Kullback-Leibler (KL) divergence. Upon merging, a single new GMM is trained on the segments that were previously assigned to the two individual clusters. The frames are then reassigned to the clusters. In S405, a cluster distance is determined and a cluster merging based on the cluster distance is performed to reduce the number of clusters. S405 is performed repeatedly until a stopping criterion is reached in S407. Assessing the stopping criterion in S407 may comprise a comparison of the distance metric of the two closest clusters, such as the KL distance, with a predetermined threshold. If the distance metric is higher than the threshold, the stopping criterion is met. In the above described example, the audio signal is initially segmented into equal length frames in S403. These frames are then assigned between the clusters in S405. In alternative examples, the segmentation in S403 is performed using hypothesis testing based on two sliding consecutive windows applied to the audio signal. In each position, each window corresponds to a frame of the signal. The consecutive windows may overlap. Two hypothesises are tested. The first hypothesis (Ho) is that the two frames correspond to the same speaker model. The second hypothesis (Hi) is that the two frames correspond to two different speaker models. The speaker models may be estimated from the two frames. A measure of which of Ho or Hi is more suitable is obtained by considering a distance metric. An example of a suitable distance metric is the Kullback-Leibler (KL) divergence. The KL divergence estimates the distance between two random distributions and may be used to characterise the similarity between two audio frames in the two sliding consecutive windows. If the KL divergence results in a high similarity of two audio segments, compared to an empirically determined threshold, then Ho is more accurate, the two audio segments are marked as corresponding to the same speaker, and no speaker turn is indicated. Conversely, when KL divergence results in a low similarity of two audio segments, compared to an empirically determined threshold, then Hi is more accurate, the two audio segments are marked as corresponding to the different speaker, and a speaker turn (a change of speaker) is indicated at the time instant corresponding to the crossover between the two consecutive windows. This process is performed over the entire audio stream and the time instants corresponding to a change of speaker (or a speaker turn) are obtained. Thus, a sequence of speaker turns is extracted. The audio signal is then split into frames in S403, with the frames from each speaker turn being assigned to a single cluster. The merging procedure in S405 is then performed. As described above, after two clusters are merged, the frames are then reassigned to the clusters. A step of re-segmenting the audio stream based on the clustering may be performed at this point. The output of the speaker diarization step S131 comprises a sequence of times that correspond to instants where the speaker has changed, together with an indication of the speaker. Thus, with reference to Figure 1, the output of S131 comprises times instants to, ti, tz and to together with Speaker 1 and Speaker 2 labels: Start time Stop time Label to ti Speaker 1 t2 to Speaker 2 As has been described previously, in Step 133 of Figure 3, a first signal comprising one or more speaker 1 turns, and a second signal comprising one or more speaker 2 turns, are generated from the start and end times and the speaker labels output from the speaker diarization. In this step, the segments of the original audio signal are used to generate the first and second signals, i.e. including the noise. In the above described method, an input audio signal is split into a first signal comprising speech from a first speaker, and a second signal comprising speech from a second speaker. In alternative examples however, the first signal and the second signal are separate audio signals. For example, the second signal may be a prerecorded signal from a first speaker which is retrieved from memory, and the first signal may be an input audio signal received from a user device. The method of proximity authentication may be performed to verify that the input audio signal is received from the location corresponding to that in which the pre-recorded signal was recorded. Figure 6 shows a flow chart of an authentication method according to an alternative embodiment in which the second signal is a pre-recorded signal and the first signal is an input signal from the same user. In step S61, an input audio signal is obtained. The audio signal may be obtained from a user device as described in relation to Figure 1 for example. In step S63, the first signal is extracted from the input audio file. The first signal might comprise the entire input audio file, or might be extracted using a speech activity detection (SAD) algorithm for example. A second audio signal is then retrieved from the storage. Acoustic condition information of the first and second signals is then obtained, in the same manner as has been described previously. In an alternative example, the acoustic condition information for the second audio signal is stored, and is simply retrieved in this step. In step S65, the acoustic conditions of the first signal are compared with the acoustic conditions of the second signal, in the same manner as has been described above. In step S67, an indication of the proximity of the location of the speaker for the first signal and the location of the speaker for the second signal is determined from the comparison of the two acoustic conditions. In this step, an indication of the spatial proximity of where the first signal was recorded to where the second signal was recorded is determined. In step S69, an output is provided based on the indication of the proximity determined in S67. Additionally or alternatively, the authentication method of Figure 6 may be used to verify that the input audio signal is genuine. For example, it may be used to verify that the input audio signal is not a recording. In some examples, additional steps of authentication may be performed. Figure 7 shows a flow chart of a method for authenticating an audio signal according to an embodiment. In this method, the obtained audio file is analysed in parallel to: authenticate the audio in S605; authenticate the speaker(s) in S603; and authenticate the proximity in S700. The outputs of the steps S605, S603 and S700 are then used to provide a final authentication in S607. In step S601, an audio file is obtained. The obtained audio signal is inputted to S605, to authenticate that the audio is genuine. This step comprises one or both of: a process of verifying that the speech is not synthetically generated S605a; and a process of verifying that the speech is not a recording S605b. Optionally, before performing S605a and S605b, channel compensation is performed. Step S605a verifies that the audio signal comprises speech from a human speaker. In this step, audio signal analysis is performed to authenticate that no voices are synthetically generated, in other words computer generated. Various methods of determining whether an audio signal comprises computer generated speech can be used, such as a trained binary classifier model that has been trained to classify whether an audio stream comprises synthesised speech or whether it is provided by a human speaker. Examples of such methods are described in Imutairi, Zaynab &Elgibreen, Hebah. (2022). A Review of Modern Audio Deepfake Detection Methods: Challenges and Future Directions. Algorithms. 15. 19. 10.3390 / a15050155, the entire contents of which are incorporated by reference herein. Such a model may take as input a set of features extracted from the audio, and be trained on datasets comprising sets of features extracted from many audio signals generated by a text to speech algorithm and many audio signals corresponding to live human speech. In this example, step S605a returns a single binary output that indicates whether the audio file is from a human speaker. For example, a value of 1 is returned if the audio file is determined to be from a human speaker, and a value of 0 is returned if the audio file is determined to comprise synthetically generated speech. Step S605b comprises verifying whether the audio comprises a voice recording. In this step, it is verified that the first speaker did not simply replay a voice recording of the second speaker for example, in order to generate an audio steam similar to that shown in Figure 1. Various methods of determining whether an audio signal corresponds to a replayed recording may be used, such as a trained binary classification model that has been trained to classify whether an audio stream comprises a replay of a recording. Such a model may take as input a set of features extracted from the audio, and be trained on datasets comprising sets of features extracted from many audio signals generated by replaying a voice recording and many audio signals corresponding to live human speech. Step S605b comprises performing audio signal analysis to determine that no voices are replays of recordings. In this example, step S605b returns a single binary output that indicates whether the audio file comprises a replay of recorded speech. For example, a value of 1 is returned if the audio file is determined not to comprise a replay of a recording, and a value of 0 is returned if the audio file is determined to comprise a replay of a recording. In this example, step S605 then returns a single binary output that indicates whether the audio file has been successfully verified (i.e., it is genuine) or not. The value is based on a combination of the outputs of S605a and S605b. For example, where S605a and S605b both return a value of 1, S605 returns a value of 1, indicating that the audio file is authenticated as genuine. If either S605a or S605b return a value of 0, S605 returns a value of 0, indicating that the audio file is not authenticated. The audio signal is also inputted to S603, to authenticate the identity of one or more speakers. Where the audio signal comprises speech from a single speaker, the identity of the speaker is authenticated in S603 from the obtained audio signal. Where the obtained audio signal comprises speech from two or more speakers, as determined in S601, each speaker may be authenticated or only one or some of the speakers may be authenticated. Step S603 comprises biometrically authenticating one or more speaker voices. For example, authenticating a first speaker in S603 comprises extracting voice information from the audio signal, and then comparing the voice information to stored voice information of one or more authorised users. The stored voice information of a user is referred to here as a template or user template. The templates of one or more authorised users are stored in a database, which may comprise templates of users who have been registered for example. As a first step, the first signal is processed to remove background noise and to perform channel compensation. As a second step, speaker diarization may be performed to extract the audio segment corresponding to different speakers. Speaker diarization may be performed as described above. The segments corresponding to each speaker may be aggregated. Authentication of the speaker is then performed using voice biometric analysis. The authentication performed in S603 may be text dependent or text independent. In a text dependent method, the user speaks a predetermined phrase. The voice information extracted from the spoken phrase is then compared to the template. The template is generated by the user speaking the predetermined phrase during the registration process. In a text independent method, the user may speak any phrase. The voice information extracted from the spoken phrase is then compared to the template. The template is generated by the user speaking any phrase during the registration process. In other words, the template may be generated from a different spoken phrase to the input. The voice biometric analysis uses an algorithm that generates a digital representation of the distortion of sound caused by the speaker’s physiology from an audio signal. This representation comprises a series of values, representing voice information. The values may be represented as float values, which are stored in a vector, referred to here as a voice information vector. The voice information comprises information specifying various characteristics of the user voice, such as pitch, cadence, tone, pronunciation, emphasis, speed of speech, and accent. The unique vocal tract and behaviour of each user results in distinct voice information that allows verification of the user using a stored template. The stored template is a vector comprising a set of values which were previously extracted from speech received during the registration process. In this example, the user provides a SpeakerlD, which identifies a stored template. The SpeakerlD may be an alphanumeric sequence that identifies the stored template. The user may speak their SpeakerlD during the capture session. An automatic speech recognition system may extract the SpeakerlD text from the audio signal and retrieve the stored template. Alternatively, the user may enter the SpeakerlD as text before or after the audio capture session. In this example, the user provides their SpeakerlD to the client device, either by speaking it as part of the capture session or by providing a separate text input. The SpeakerlD is then provided to the server by the client device. The stored template corresponding to the user is then identified by the Speaker ID. For example, the user or users may indicate that the audio signal corresponds to Speaker ID = Alice and Speaker ID = Bob. In S603, the first signal is then authenticated using a stored template identified by Speaker ID = Alice, and the second signal using a template identified by Speaker ID = Bob. Alternatively, the speaker for the first signal is authenticated against a cluster of templates identified by a Cluster ID, and the speaker for the second signal is authenticated against a cluster of templates identified by a Cluster ID. The Cluster ID is provided by a user, either by speaking it as part of the capture session or by providing a separate text input, as described above. In S603, it is then confirmed that the identity of each speaker belongs to a biometric cluster of the provided Cluster ID. In this example, step S603 returns a binary output that indicates whether the user or users have been successfully authenticated or not. Alternatively, S603 returns a real number. For example, S603 returns a probability that the speaker(s) are authenticated. In S700, the proximity is authenticated. This step comprises: determining the acoustic condition information S713, comparing the determined acoustic condition information, and determining an indication of the proximity S717. Steps S713, S715, and S717 may correspond to S13, S15, and S17 described above in relation to Figure 2(b) or steps S63, S65 and S67 in relation to Figure 6 for example. In some examples, a step of determining how many individual speakers are contained within the audio file is performed as part of S700. For example, where the method is used to authenticate that two speakers are in the same location, it is expected that the input audio file will comprise speech from two speakers. Step S12 described above in relation to Figure 2(b) may be performed as part of S700 in such examples. In this step, audio signal analysis is performed to determine that the acoustic conditions associated with each individual speaker or silence segment is the same. This is done by checking changes on the recording of room reverberation, recording channel and background noise for example, as has been described previously. In this example, step S700 returns a binary output that indicates the proximity, as has been described previously. Alternatively, S700 returns a real number. For example, S700 returns a probability. In step S607, the outputs of S605, S603, and S700 are combined to provide a final authentication, and an output is generated based on the result. In this example, S605, S603, and S700 return binary outputs, and therefore the audio file is considered to be authenticated if each of S605, S603, and S700 returns an output that indicates that the audio file has passed the respective verification. Alternatively, in step S607, when S605, S603, or S700 return a real number or a probability, the three outputs are combined together. The combined value is compared to a predetermined condition (e.g. a predetermined threshold), and if the condition is met, then the authentication is considered successful. For example the outputs may be combined by averaging, or summing. In this example, S607 comprises sending an output based on the result to a server. The server may execute further processes in response to the output. Alternatively or additionally, S607 comprises a step of displaying a message on a screen or producing a sound. The method described in relation to Figure 7 can be used where a single audio file or audio stream that is created by two or more speakers speaking into the same app session (for example running on a client device which is a smartphone or tablet or PC) or same telephone call is collected, such that a single, continuous audio file or audio stream is created. This audio file or audio stream will then comprise two or more speaker signals, potentially silence, and sound from the acoustic conditions, including background noise or noise floor, room reverberation and channel characteristics. The method described in relation to Figure 7 then comprises determining how many individual speakers are contained within the file. A step of biometrically authenticating each voice against a template identified by a Speaker ID passed to the server by the app or IVR application that collected the audio, or alternatively biometrically identifying and authenticating each voice against a cluster of templates identified by a Cluster ID is then performed. The method further comprises performing audio signal analysis to biometrically determine that no voices are recordings i.e. replays. The method further comprises performing audio signal analysis to determine that no voices are synthetically generated, i.e. computer generated. The method further comprises performing audio signal analysis to determine that the acoustic conditions associated with each individual speaker or silence segment is the same. This is done by checking changes on the recording of: room reverberation; recording channel e.g. telephony, codec; and background noise. The combination of these measures allows confirmation that the identity of each speaker matches the biometric template of the provided Speaker ID or belongs to the biometric cluster of the provided Cluster ID, and that all speakers were in the same physical location when the continuous file was created. Although Figure 7 illustrates a method where steps S700, S603 and S605 are applied to an audio file to perform authentication, the method may instead comprise verifying proximity S700 and authenticating the speaker(s) S603, or verifying proximity S700 and authenticating the audio S605 for example. Although Figure 7 illustrates a method where steps S700, S603 and S605 are applied in parallel to an audio file to perform authentication, the method may instead comprise applying the steps sequentially or partly sequentially. For example, the method comprises first authenticating the speaker(s) in S603, and, if the speaker(s) is authenticated in S603, then performing S605 and / or S700. For example, as described above, S700 may comprise a step S12 of analysing the audio file to determine the number of speakers. As described in relation to Figure 4, this may comprise a step of detecting speech segments with a speech activity detection (SAD) algorithm. The output from this step can also be used in S603 for example. For example, by passing only segments of the audio signal that contain speech to S603, the authentication method is rendered more efficient since only these segments are processed. Thus in an example, a first part of S700 comprising determining the number of speakers is performed prior to at least a part of S603. In some examples, the biometric analysis performed in S603 comprises removing background noise and performing channel compensation, to improve accuracy of the voice recognition. The output from the noise removal step can also be used in the speaker diarization process described in relation to Figure 4, which is performed in S700. In particular, in some examples, the speaker diarization process described in relation to Figure 4 comprises a step of removing background noise S401. This step can be omitted by taking the output of the noise removal step performed as part of S603. Thus in some examples, a first part of the step S603 comprising background noise removal is performed prior to at least a part of S700. In some examples, the background noise removed from the signal in S603 is used as the measure of background noise in S713. In some examples, data generated in S603 to perform the channel compensation is used to determine the measure of the channel characteristics in S713. Thus in some examples, a first part of the step S603 comprising background noise removal and channel estimation is performed prior to at least a part of S700. The combination of the authentication steps described in relation to Figure 7 can be used in a method to confirm whether the identity of each speaker matches a biometric template of the provided Speaker ID or belongs to a biometric cluster of the provided Cluster ID, and also to confirm whether all speakers were in the same physical location when the continuous audio file was created. While certain embodiments have been described, these embodiments have been presented by way of example only, and are not intended to limit the scope of the inventions. Indeed the novel methods and apparatus described herein may be embodied in a variety of other forms; furthermore, various omissions, substitutions and 5 changes in the form of methods and apparatus described herein may be made.
Claims
1. An authentication system comprising:an input module configured to obtain an input audio signal;one or more processors, the one or more processors configured to:determine acoustic condition information for a first audio signal;compare the acoustic condition information for the first audio signal with acoustic condition information for a second audio signal to determine an indication of proximity; andan output module configured to provide an output based on the indication of proximity.
2. The system according to claim 1, wherein determining the indication of proximity comprises determining a difference between the acoustic condition information of the first signal and the acoustic condition information of the second signal, and determining whether the difference meets a pre-determined criteria.
3. The system according to claim 1 or 2, whereinthe first audio signal comprises speech from a first speaker; and the second audio signal comprises speech from a second speaker.
4. The system according to any preceding claim, wherein the processor is configured to extract the first audio signal and the second audio signal from an input audio signal using a speaker diarization process.
5. The system according to any preceding claim, wherein determining acoustic condition information comprises extracting a measure of at least one of: room reverberation, recording channel, and background noise.
6. The system according to any preceding claim, wherein the first audio signal comprises speech from a first speaker; and wherein the one or more processors are further configured to:biometrically authenticate the first speaker voice against a stored template.
7. The system according to claim 6, wherein the second audio signal comprises speech from a second speaker; and wherein the one or more processors are further configured to:biometrically authenticate the second speaker voice against a stored template.
8. The system according to any preceding claim, wherein the one or more processors are further configured to:authenticate that the first audio signal is not a replay of a recording.
9. The system according to any preceding claim, wherein the first audio signal comprises speech from a first speaker, and wherein the one or more processors are further configured to:authenticate that the speech is not computer generated.
11. The system according to claim 4, wherein the speaker diarization process comprises performing a speech activity detection process and wherein the first audio signal comprises speech from a first speaker, wherein the one or more processors are further configured to:biometrically authenticate a first speaker voice against a stored template using an output of the speech activity detection process.
12. The system according to claim 6 or 7, wherein the biometric authentication comprises a step of channel compensation, and wherein an output from the channel compensation step is used to determine the acoustic condition information.
13. The system according to claim 6 or 7, wherein the biometric authentication comprises a step of background noise remove, and wherein an output from the background noise removal step is used to determine the acoustic condition information.
14. An authentication method comprising:obtaining, by way of an input, an input audio signal;determining acoustic condition information for a first audio signal;comparing the acoustic condition information for the first audio signal with acoustic condition information for a second audio signal to determine an indication of proximity; andproviding an output based on the indication of proximity.
15. A carrier medium comprising computer readable code configured to cause a computer to perform the method of claim 14.
Citation Information
Patent Citations
Apparatus and method for determining co-location of services
US20160337803A1
Electronic device and control method thereof
US20210398544A1