USER IDENTIFICATION BASED ON VOICE INPUT

DE502021010956D1Active Publication Date: 2026-09-17DEUTSCHE TELEKOM AG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE502021010956
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-12-22
Publication Date
2026-09-17
Estimated Expiration
2041-12-22

AI Technical Summary

Technical Problem

Existing user identification methods in distributed applications require separate verification steps, which are inconvenient for users and can be computationally intensive, leading to suboptimal user experience and response times.

Method used

A distributed application captures a user's voice input, generates an audio file, and uses a backend to determine a subset of stored audio files based on similarity and metadata analysis, excluding incompatible files to identify the user without additional verification steps, utilizing metadata such as audio format, pitch, and confidence scores to ensure efficient matching.

Benefits of technology

This method provides a highly convenient and efficient user identification process with reduced computational load, ensuring quick authentication by narrowing down the comparison to a manageable subset of audio files, thus enhancing user experience and response time.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The invention relates to a method for identifying a user, in which a frontend of a distributed application executed on an end device captures input from a user of the end device, and a backend of the distributed application executed on a server identifies the user based on the captured input. The invention further relates to a distributed application for identifying a user and a computer program product.

[0002] Methods of the type mentioned above, in various forms, represent the state of the art and serve to restrict the use of a distributed application to identifiable users. For example, the distributed application identifies a user based on a username, email address, membership number, or customer number entered by the user.

[0003] Identifying a user requires prior user registration. To register a user, the distributed application stores the username, email address, or membership / customer number and identifies the user after registration based on the stored username, email address, or membership / customer number.

[0004] However, the distributed application must also verify whether a person who enters a username, email address, membership number, or customer number—that is, claims to be the registered user—is indeed the registered user. To verify the user, the distributed application prompts the identified user for further input that identifies them as the registered user. For example, the distributed application can verify the identified user using a password entered by the user. This further input can be provided using an electronic device in the possession of the registered user, such as a card readable by a digital reader.

[0005] US 2009 / 175424 A1 discloses a method for determining a user's group affiliation. Based on an audio file corresponding to a user's voice input, the user's age and / or gender is determined. If a user group is defined by age and / or gender, the user's group affiliation can be determined based on the identified age and / or gender.

[0006] US 2021 / 304774 A1 discloses a method for identifying a user by means of a user profile. The user profile comprises a voice profile with a plurality of audio files corresponding to the user's voice input. A distinction is made between explicit voice profiles, i.e., those assigned to known users, and anonymous voice profiles, i.e., those assigned to unknown users.

[0007] US 2020 / 395021 A1 discloses a method for distinguishing between users of a household. Demographic data such as age or gender of a user of the household are derived from an audio file corresponding to a voice input. The derived demographic data and, if necessary, the content of the voice input are used to distinguish the user from other users of the household.

[0008] An identified and verified user is considered authentic by the distributed application and can use it. However, users may find it inconvenient that identification requires separate verification before they can use the distributed application.

[0009] Therefore, one object of the invention is to propose a method for identifying a user that offers a high level of user convenience. A further object of the invention is to provide a distributed application for identifying a user and a computer program product.

[0010] An object of the invention is a method for identifying a user, in which a frontend of a distributed application, executed on an end device, captures input from a user of the end device, and a backend of the distributed application, executed on a server, identifies the user based on the captured input. The user operates the end device, in particular a mobile device such as a smartphone, tablet, notebook, or the like. The frontend captures the user's input via the end device. The backend is connected to the frontend via a network, in particular a mobile network; that is, both the end device and the server are connected to the network.

[0011] According to the invention, the frontend acoustically captures a user's voice input as the input, generates an audio file corresponding to the captured voice input, and transmits the generated audio file to the backend of the distributed application. The backend receives the transmitted audio file, uses it to determine a subset of stored audio files, and identifies the user based on the received audio file if the degree of similarity between the received audio file and an audio file from the determined subset is equal to or greater than a certain acceptance threshold. Voice input is already very convenient for the user. Moreover, voice input, like a fingerprint, is highly individual and difficult to imitate. Consequently, the distributed application can authenticate the user solely based on the identifying input.In other words, no further user input is required to authenticate the user, which results in even greater user convenience.

[0012] The specified subset reduces the computationally intensive task of determining the degree of match. Thanks to the specified subset, the backend doesn't need to compare the transmitted audio file with every stored audio file, resulting in a practically acceptable response time for user identification. The backend determines the subset solely based on the received audio file and the stored audio files. Determining the subset is equivalent to classifying the stored audio files in relation to the received audio file.

[0013] Preferably, the backend determines the subset by excluding stored audio files. When determining the subset, the backend starts with the majority, i.e., the total number, of stored audio files as the defined subset and excludes audio files complementary to the received audio file from this subset. This exclusion process reduces the size of the defined subset to such an extent that the degree of similarity between the received audio file and each stored audio file within the defined subset can be determined within a practically acceptable response time. Based on experimental experience, subsets of five to ten stored audio files are practical.

[0014] Furthermore, the backend preferentially retrieves multiple metadata for each audio file, assigns the retrieved metadata to the audio file, and excludes a stored audio file from the specified subset based on the assigned metadata if it detects an incompatibility between a metadata entry assigned to the stored audio file and a corresponding metadata entry assigned to the received audio file. The backend retrieves the metadata depending on the audio file. The mapping of audio file and metadata can be described as a speech profile of the user.

[0015] To determine incompatibility, the backend compares the metadata of the stored audio file with the corresponding metadata of the received audio file. Comparing metadata is significantly less computationally intensive than comparing audio files. The metadata of the received audio file can be compared with the corresponding metadata of all stored audio files within a practically acceptable processing time.

[0016] Advantageously, when determining incompatibility, the backend defines an order for the corresponding metadata based on the size of the subset to be identified. For each metadata element of the received audio file, the backend determines the size of the remaining subset when stored audio files with metadata incompatible with that metadata are excluded. In this way, the backend can prevent the identified subset from being too large for a practically acceptable response time or too small for successful user identification.

[0017] In particular, the backend can define the order of the corresponding metadata using feedback. If, after a multi-step exclusion of audio files, the remaining subset is too small or too large, the backend can undo the exclusion steps and define a different order for the corresponding metadata. Defining the different order corresponds to a different weighting of the corresponding metadata.

[0018] In particular, the backend can exclude one or more metadata elements of the received audio file from the incompatibility detection, for example, if the practical size of the subset has already been reached or the metadata leads to an empty subset.

[0019] Ideally, the backend derives each piece of metadata from the audio file and assigns the derived metadata to the audio file. The backend analyzes the form and / or content of the audio file and uses well-known tools for this analysis.

[0020] In one embodiment, the backend derives an audio format, a sampling rate, and / or a bit rate as technical metadata from the audio file. The audio format, sampling rate, and bit rate are contained as parameters in the audio file and are extracted from it by the backend. High-quality audio files have a sampling rate of 44 kHz and a bit rate of 96 kbps. Conversely, a sampling rate of 8 kHz (common in telephone networks) indicates poor audio quality. The backend can weight other metadata, different from the technical metadata, as described below.

[0021] Alternatively or additionally, the backend can derive pitch, timbre, and / or volume as vocal metadata from the audio file. The backend can calculate a frequency spectrum of the audio file using Fourier analysis and derive the pitch and timbre from this spectrum. The backend can calculate the volume from the amplitude of the audio file. Pitch and timbre are referred to as melodic parameters of the audio file. Volume is considered a dynamic parameter of the audio file.

[0022] Alternatively or additionally, the backend can derive language, speech rhythm, articulation, and / or intonation as linguistic metadata from the audio file. The backend can determine the language, for example, German, English, or French, using a known automatic speech recognition (ASR) algorithm. Speech rhythm comprises the ratio of speech to pauses, the number of words and / or syllables per unit of time. Speech rhythm is referred to as a temporal parameter of the audio file. Articulation comprises the change in amplitude per unit of time and is a measure of speech coherence and clarity. Articulation is referred to as an articulatory parameter of the audio file. Intonation comprises stress and is calculated from the amplitude of the audio file. Intonation is one of the dynamic parameters of the audio file.

[0023] In other implementations, the backend derives the user's gender and / or age as personal metadata from the audio file. The gender and age are derived from the voice and speech metadata. For example, an older person speaks more quietly and in a lower register than a younger person, or a woman speaks in a higher register than a man. It is understood that the attributions of young / old or female / male only classify the voice input, and not necessarily the user.

[0024] Ideally, the backend calculates a confidence score for each piece of vocal, linguistic, and / or personal metadata and assigns this score to the corresponding metadata. The confidence score can be calculated based on the technical metadata. Ranging from 0% to 100%, the confidence score indicates the likelihood that the derived metadata matches the audio file. This increased confidence booster raises the probability that the stored audio file identifying the user will not be excluded by the backend due to a detected incompatibility.

[0025] In advantageous embodiments, the backend determines the incompatibility based on the respective assigned confidence values. When determining the incompatibility, the backend considers confidence values ​​of the received audio file and / or confidence values ​​of the stored audio file. For example, the backend can ignore metadata in the received or stored audio file that has a low confidence value, i.e., is uncertain or has low confidence.

[0026] In another example of using confidence levels, the backend detects an incompatibility with any "female" voice input in a stored audio file if the received audio file has a "male" voice input and a confidence level above 70%. However, if the "male" voice input in the received audio file has a confidence level of 30%, an incompatibility is only detected for "female" voice inputs in stored audio files with a confidence level of 70%.

[0027] In short, the more secure the metadata of the received audio file, the more stored audio files will be excluded. Conversely, the less secure the metadata of the received audio file, the fewer stored audio files will be excluded.

[0028] If the backend defines the order of the corresponding metadata using feedback, it can undo any exclusion steps and repeat them with adjusted confidence thresholds. A threshold is increased if the remaining subset is too large and decreased if it is too small. Naturally, when repeating the exclusion steps, the backend can define a different order for the corresponding metadata and / or adjust the confidence thresholds.

[0029] This increases the probability that the stored audio file matching the received audio file belongs to the specified subset and that the user can be identified using the received audio file.

[0030] The backend can determine the acceptance threshold depending on a defined security level. For example, an acceptance threshold of over 90% may be required for a business transaction application, while an acceptance threshold of 50%-60% may be sufficient for a private gaming application.

[0031] In many implementations, the backend stores an initial audio file received from the user's device. This first audio file, corresponding to the user's voice input, is stored during user registration. It must be included in the specified subset to identify the user.

[0032] Another aspect of the invention is a distributed application for identifying a user, comprising a frontend executable from an end device and a backend executable from a server. Distributed applications for user identification are widely used. Consequently, the distributed application according to the invention is applicable in many ways.

[0033] According to the invention, the distributed application is configured to execute a method according to the invention. The distributed application offers a user great convenience in identification and authentication.

[0034] Another aspect of the invention is a computer program product comprising a computer-readable storage medium and program code stored on the storage medium. The computer can read the program directly from the storage medium and / or, after installing the program in the computer's main memory, from its working memory. Accordingly, the storage medium can be an external data carrier such as a CD-ROM or a USB flash drive, or an internal data storage device such as the computer's hard drive or main memory, or cloud storage.

[0035] According to the invention, the program code causes the computer, acting as an end device, to execute a frontend or, acting as a server, a backend of a distributed application according to the invention for identifying a user, when it is read and executed by a processor of the computer. The computer program product implements a distributed application for identifying a user, which offers the user a high degree of convenience.

[0036] A key advantage of the method according to the invention is that the user perceives identification and authentication via a distributed application as very convenient. A further advantage is that the distributed application is universally applicable and adaptable to various security requirements.

[0037] It is understood that the features mentioned above and those to be explained below can be used not only in the combinations specified, but also in other combinations or on their own, without leaving the scope of the present invention.

[0038] The invention is schematically illustrated in the drawings using an exemplary embodiment and is described in detail below with reference to the drawings. It shows Fig. 1 in a block diagram a distributed application according to an embodiment of the invention for identifying a user when executing a method according to a first embodiment of the invention; Fig. 2 in a block diagram the in Fig. 1 The distributed application shown is used when performing a method according to a second embodiment of the invention.

[0039] Fig. 1 Figure 1 shows a distributed application 1 according to an embodiment of the invention for identifying a user 4 when executing a method according to a first embodiment of the invention. The distributed application 1 for identifying the user 4 comprises a frontend 10 executable by an end device 2 and a backend 11 executable by a server 3 and is configured to execute a method described below.

[0040] The distributed application 1 can be implemented using a computer program product. The computer program product comprises a storage medium readable by a computer and program code stored on the storage medium, which, when executed by a processor of the computer, causes the computer, acting as the terminal device 2, to provide the frontend 10 or, acting as the server 3, the backend 11 of the distributed application 1 for the purpose of identifying a user 4.

[0041] In the process for identifying a user 4, the frontend 10 of the distributed application 1, executed by the terminal device 2, captures an input 40 from the user 4 of the terminal device 2, and the backend 11 of the distributed application 1, executed by the server 3, identifies the user 4 based on the captured input 40.

[0042] The frontend 10 captures a voice input from the user 4 as the input 40 audibly, generates an audio file 100 corresponding to the captured voice input and transfers the generated audio file 100 to the backend 11 of the distributed application 1.

[0043] Backend 11 receives the transmitted audio file 100 and, based on the received audio file 100, determines a subset 112 of audio files 110 from a plurality of stored audio files 110. Backend 11 determines the subset 112, in particular by excluding stored audio files 110.

[0044] For this purpose, backend 11 can determine a plurality of metadata 101, 111 for each audio file 100, 110, assign the determined metadata 101, 111 to the audio file 100, 110, and exclude a stored audio file 110 from the specified subset 112 based on assigned metadata 111 if backend 11 detects an incompatibility between a metadata 111 assigned to the stored audio file 110 and a corresponding metadata 101 assigned to the received audio file 100. The stored audio file 110 can be stored in a registry database 30 of server 3.

[0045] Backend 11 preferentially derives each metadata 100, 110 from the audio file 100, 110 and assigns the derived metadata 101, 111 to the audio file 100, 110.

[0046] Specifically, backend 11 can derive an audio format, a sampling rate, and / or a bit rate as technical metadata 1010, 1110 from audio file 100, 110. Backend 11 advantageously derives a pitch, timbre, and / or volume as vocal metadata 1011, 1111 from audio file 100, 110. Furthermore, backend 11 can derive a language, speech rhythm, articulation, and / or intonation as linguistic metadata 1012, 1112 from audio file 100, 110.

[0047] Ideally, the backend 11 derives a gender and / or age of user 4 as personal metadata 1013, 1113 from the audio file 100, 110.

[0048] Backend 11 can furthermore calculate a confidence value for each vocal, linguistic, and / or personal metadata element 1011, 1012, 1013, 1111, 1112, 1113 and assign the calculated confidence value to the respective metadata element 1011, 1012, 1013, 1111, 1112, 1113. Preferably, backend 11 determines the incompatibility based on the respective assigned confidence values.

[0049] When determining incompatibility, the backend 11 cleverly defines an order of the corresponding metadata 100, 110 depending on the size of the subset 112 to be determined.

[0050] Backend 11 identifies user 4 based on the received audio file 100 if the degree of similarity 113 between the received audio file 100 and an audio file 110 from the specified subset 112 is equal to or greater than a certain acceptance threshold. Backend 11 ideally determines the acceptance threshold based on a defined security level.

[0051] Fig. 2 shows in a block diagram the in Fig. 1 Distributed application 1 shown, when performing a method according to a second embodiment of the invention.

[0052] The frontend 10 of the distributed application 1, executed by the terminal device 2, captures an input 40 from the user 4 of the terminal device 2. The frontend 10 captures a voice input from the user 4 as input 40 audibly, generates an audio file 100 corresponding to the captured voice input, and transmits the generated audio file 100 to the backend 11 of the distributed application 1.

[0053] The backend 11 can determine a plurality of metadata 101 for the audio file 100 and assign the determined metadata 101 to the audio file 100.

[0054] Backend 11 preferentially derives each metadata element 101 from audio file 100 and assigns the derived metadata element 101 to audio file 100. The assignment from audio file 100 and metadata 101 can be referred to as a speech profile 102 of user 4.

[0055] Specifically, backend 11 can derive an audio format, a sampling rate, and / or a bit rate as technical metadata 1010 from audio file 100. Backend 11 advantageously derives pitch, timbre, and / or volume as vocal metadata 1011 from audio file 100. Furthermore, backend 11 can derive speech, speech rhythm, articulation, and / or intonation as linguistic metadata 1012 from audio file 100.

[0056] Ideally, the backend 11 derives a gender and / or age of the user 4 as personal metadata 1013 from the audio file 100.

[0057] The backend 11 can also calculate a confidence value for each vocal, linguistic and / or personal metadata 1011, 1012, 1013 and assign the calculated confidence value to the respective metadata 1011, 1012, 1013.

[0058] To register user 4, the backend 11 preferably stores the audio file 100 received from the end device 2 of user 4, especially with the associated metadata 101 and / or in a registration database 30 of the server 3, if it receives the audio file 100 as a first audio file 100 of user 4. Reference symbol list

[0059] 1 Distributed application 10 Frontend 100 Audio file 101 Metadata 1010 Technical metadata 1011 Vocal metadata 1012 Language metadata 1013 Personal metadata 102 Speech profile 11 Backend 110 Audio file 111 Metadata 1110 Technical metadata 1111 Vocal metadata 1112 Language metadata 1113 Personal metadata 112 Subset 113 Match rate 2 End device 3 Server 30 Registry database 4 User 40 Input

Claims

1. A method for identifying a user (4), wherein - a frontend (10) of a distributed application (1), which frontend (10) is executed by a terminal (2), acquires an input from a user (4) of the terminal (2), and a backend (11) of the distributed application (1), which backend (11) is executed by a server (3), identifies the user (4) based the acquired input; - the frontend (10) auditively acquires a voice input from the user (4) as the input, generates an audio file (100) corresponding to the acquired voice input, and transmits the generated audio file (100) to the backend (11) of the distributed application (1); - the backend (11) receives the transmitted audio file (100), determines a subset (112) of audio files (110) from a plurality of stored audio files (110) based on the received audio file (100); - the backend (11) identifies the user (4) on the basis of the received audio file (100) if a match level (113) of the received audio file (100) with an audio file (110) from the determined subset (112) is equal to or greater than a calculated acceptance threshold.

2. The method according to claim 1, wherein the backend (11) determines the subset (112) by excluding stored audio files (110).

3. The method according to claim 2, wherein the backend (11) calculates a plurality of metadata (101, 111) for each audio file (100, 110), assigns the calculated metadata (101, 111) to the audio file (100, 110), and excludes a stored audio file (110) from the calculated subset (112) on the basis of the assigned metadata (111) if the backend (11) calculates an incompatibility between a metadatum (111) assigned to the stored audio file (110) and a corresponding metadatum (101) assigned to the received audio file (100).

4. The method according to claim 3, wherein the backend (11), when calculating the incompatibility, defines an order of the corresponding metadata (101, 111) depending on the size of the subset (112) to be calculated.

5. The method according to claim 3 or 4, wherein the backend (11) derives each metadatum (100, 110) from the audio file (100, 110) and assigns the respectively derived metadatum (101, 111) to the audio file (100, 110).

6. The method according to claim 5, wherein the backend (11) derives an audio format, a sample rate, and / or a bit rate from the audio file (100, 110) as technical metadata (1010, 1110).

7. The method according to claim 5 or 6, wherein the backend (11) derives a pitch, a timbre, and / or a volume from the audio file (100, 110) as vocal metadata (1011, 1111).

8. The method according to any of claims 5 to 7, wherein the backend (11) derives a language, a speech rhythm, an articulation and / or a speech melody from the audio file (100, 110) as speech metadata (1012, 1112).

9. The method according to claim 7 or 8, wherein the backend (11) derives a gender and / or an age of the user (4) from the audio file (100, 110) as personal metadata (1013, 1113).

10. The method according to any of claims 7 to 9, wherein the backend (11) calculates a confidence value for each vocal, speech and / or personal metadatum (1011, 1012, 1013, 1111, 1112, 1113) and assigns the calculated confidence value to the respective metadatum (1011, 1012, 1013, 1111, 1112, 1113).

11. The method according to claim 10, wherein the backend (11) calculates the incompatibility depending on the respectively assigned confidence values.

12. The method according to any of claims 1 to 11, wherein the backend (11) calculates the acceptance threshold depending on a defined security level.

13. The method according to any of claims 1 to 12, wherein the backend (11) stores a first audio file (100) received from the terminal (2) of the user (4).

14. A distributed application (1) for identifying a user (4), having a frontend (10) executable by a terminal (2) and a backend (11) executable by a server (3), which application is configured to execute a method according to any of claims 1 to 13.

15. A computer program product comprising a storage medium readable by a computer and a program code stored by the storage medium, which program code causes the computer to provide, as a terminal (2), a frontend (10) or, as a server (3), a backend (11) of a distributed application (1) according to claim 14 for identifying a user (4) when being executed by a processor of the computer.