Method for identifying a speaker, method for recording a speaker and motor vehicle

By mixing a brief recognized speech signal with a predetermined extended signal in a motor vehicle to generate a complete signal, the problem of unreliable speaker recognition in the prior art is solved, achieving easy and reliable speaker recognition and system security.

CN115836345BActive Publication Date: 2026-08-04安培簡式股份有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
安培簡式股份有限公司
Filing Date
2021-03-02
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

In existing technologies, speech recognition methods based on wake-up phrases are difficult to reliably identify speakers, especially in motor vehicles. Phrases that are too long are inconvenient to use, while phrases that are too short are difficult to recognize reliably.

Method used

By using a reference speech signature recorded in computer memory, a complete signal is generated by mixing a short recognized speech signal with a predetermined extended signal, an extended speech signature is constructed, and comparisons are made to identify the speaker.

Benefits of technology

It achieves easy and reliable speaker recognition, improves system security and robustness to noise, and ensures user experience and system security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115836345B_ABST
    Figure CN115836345B_ABST
Patent Text Reader

Abstract

The invention relates to a method for identifying a specific speaker from a group of speakers by a computer comprising a computer memory in which speech signatures are stored, each speech signature being associated with one of the speakers of the group, the method comprising the steps of: - acquiring a speech signal produced by the specific speaker (S41), - constructing a new speech signature from the speech signal, - comparing the new speech signature with at least one of the speech signatures stored in the computer memory, and - identifying the specific speaker according to the result of the comparison. According to the invention, before the step of constructing, a step is provided of generating a complete signal (S4) comprising the speech signal and at least one predetermined extension signal (S31, S32), and in the step of constructing, it is provided that the new speech signature is also constructed from each extension signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention generally relates to the field of human recognition based on human speech.

[0002] Particularly advantageously, the present invention is applicable to identifying users of motor vehicles.

[0003] More specifically, the present invention relates to a method for identifying a specific speaker from a group of speakers by means of a computer including a computer memory, the computer memory storing at least one reference speech signature associated with one of the speakers in the group, the method comprising the steps of:

[0004] - Acquire the recognized speech signal generated by that specific speaker.

[0005] - Construct a recognized voice signature based on the recognized voice signal.

[0006] - Compare the identified voice signature with at least one reference voice signature recorded in the computer memory, and

[0007] - Identify the specific speaker based on the results of the comparison.

[0008] The present invention also relates to a method for recording a new speaker in a computer memory.

[0009] Finally, the present invention relates to a motor vehicle that includes the technical means required to implement one of the two methods and / or the other method. Background Technology

[0010] A known approach is to use wake-up phrases to bring electronic devices out of standby mode, thereby enabling control of specific functions. An example of a wake-up phrase is "Hello, Google." This phrase allows Android... ® The device leaves standby mode so that it can then perform specific actions (search for answers to questions, turn on lights, etc.).

[0011] The chosen wake-up phrases should be particularly short so that the speaker can pronounce them quickly.

[0012] One difficulty is that speakers often pronounce the phrase very quickly, and sometimes with truncation. This makes it difficult to detect the phrase using equipment.

[0013] Therefore, it should be understood that it is impossible to reliably identify the speaker based solely on the wake-up phrase.

[0014] Currently, particularly in the automotive industry, there is a desire to identify passengers who generate voice commands in order to, for example, ensure whether these passengers are authorized to generate those commands. For instance, it is desirable to ensure that a passenger ordering their car window to be fully opened is authorized to do so.

[0015] One known solution for identifying people in the field of voice biometrics involves requiring the speaker to produce a long phrase, such as "My voice is a password." The length of this phrase then serves as proof that the speaker can be identified from among the various speakers already recorded in the system.

[0016] The drawback of these phrases is that, due to their length, proving their pronunciation is too cumbersome to use frequently. Summary of the Invention

[0017] To overcome the aforementioned shortcomings of the prior art, the present invention proposes to use short phrases, and then enrich these phrases through calculation and in a way that is imperceptible to the user, so as to be able to identify anyone who generates the phrases with high reliability.

[0018] More specifically, according to the present invention, an identification method as defined in the introduction is proposed, wherein upstream, at least one reference speech signature recorded in a computer memory is determined based on a recorded speech signal and a predetermined extended signal, and wherein, prior to the construction step, a step is specified to generate a complete signal including the identified speech signal and the predetermined extended signal, and wherein, in the construction step, the identified speech signature is also constructed based on the extended signal.

[0019] The recorded speech signal can be recorded in a computer application. This signal is mixed with an extended signal and then processed to infer the recorded speech signature.

[0020] During the recognition process, the speaker generates a recognition speech signal again, which is mixed with the same extended signal and then processed to infer the recognition speech signature.

[0021] The identified voice signature is then compared with the voice signatures of all records stored in the application's memory in order to identify the speaker.

[0022] Therefore, a comparison is made between speech signatures enriched by extended signals.

[0023] In other words, with this invention, the speech signal used can be a short phrase, as long as the phrase is subsequently extended by an extended signal, which makes it a longer phrase, thereby ensuring better speaker identification from the speakers recorded in the system.

[0024] One advantage of this solution is that it is easier for users, as they, as before, prefer to speak shorter phrases.

[0025] Another advantage of this solution is its enhanced computer security. Specifically, if a hacker manages to obtain a user's recorded voice message, they will be unable to do anything with the recording because they are unaware of the extended signals that must be added to the voice signal for successful identification.

[0026] Another advantage is that the solution ensures better robustness to external parasitic noise because the added extended signal is noise-free and thus reduces the overall noise level of the complete signal used for identification.

[0027] The following are other advantageous and non-limiting features of the identification method according to the invention, which can be considered individually or in any technically feasible combination:

[0028] - The computer memory includes multiple reference speech signatures associated with multiple speakers in the group, the extended signal being associated with one of these speakers and different from the extended signals associated with the other speakers, the memory storing each extended signal for association with one of these speakers;

[0029] - In this generation step, the computer generates at least as many complete signals as the speakers in the group, each complete signal including the identified speech signal and one of the extended signals recorded in the memory;

[0030] - In this construction step, the computer builds a recognized speech signature for each complete signal;

[0031] - In this comparison step, the computer compares each identified voice signature with each reference voice signature recorded in the memory in order to infer a score from them;

[0032] - In this identification step, the inferred score is considered to identify the specific speaker;

[0033] - The complete signal is constructed by appending the extended signal before and / or after the recognized speech signal;

[0034] - The extended signal is a function of the sum of at least one sine wave with a frequency between 50 and 650 Hz, and preferably between 100 and 500 Hz;

[0035] - The extended signal is generated by the product of a parameterizable function and an observation window function, wherein the parameterizable function is preferably amplitude modulation and / or frequency modulation;

[0036] - The maximum amplitude of the extended signal is less than or equal to the maximum amplitude of the recognized speech signal, and preferably less than or equal to 80% of the maximum amplitude of the recognized speech signal;

[0037] - The maximum length of the at least one extended signal is less than or equal to one-third of the total length of the complete signal, and preferably equal to 20% of the total length of the complete signal;

[0038] - The recognized speech signal includes four or fewer syllables.

[0039] The present invention also relates to a method for recording a specific speaker using a computer including computer memory, the method comprising the following steps:

[0040] - Acquire the recorded speech signal produced by that specific speaker.

[0041] - Determine the extended signal,

[0042] - Generate a complete signal that includes the recorded speech signal and the extended signal.

[0043] - Determine the reference speech signature based on the complete signal of the record, and

[0044] - Store the reference voice signature in the memory so as to associate it with the specific speaker.

[0045] The present invention also relates to a motor vehicle including a passenger compartment, means for acquiring speech signals generated by a specific speaker located in the passenger compartment, and a computing unit programmed to implement one or more of the methods described above.

[0046] Of course, the various features, variations and embodiments of the present invention can be associated with each other in various combinations, as long as they are not mutually exclusive or incompatible. Attached Figure Description

[0047] Referring to the accompanying drawings, the following description, given by way of non-limiting examples, will make the scope of the invention and how it can be practiced quite clear.

[0048] In the attached diagram:

[0049] Figure 1 It is a graph showing the parameterizable functions that can be used in the context of the method according to the present invention;

[0050] Figure 2 It is a graph showing the observation window function that can be used in the context of the method according to the present invention;

[0051] Figure 3It is a graph showing the extended functions that can be used in the context of the method according to the present invention;

[0052] Figure 4 It showcases including Figure 3 A graph of the complete signal from the extended function;

[0053] Figure 5 This is a diagram illustrating one embodiment of the identification method according to the present invention. Detailed Implementation

[0054] This invention can be implemented on any type of device.

[0055] In the examples described herein, the invention will be implemented in a motor vehicle, and more particularly in a car that can accommodate several users (driver and passengers).

[0056] The vehicle will be in a traditional form.

[0057] Therefore, the vehicle includes a chassis that defines the passenger compartment for the user.

[0058] The vehicle also includes voice signal acquisition devices. These acquisition devices are arranged in the form of microphones in the vehicle to record phrases produced by individual passengers within the vehicle.

[0059] The motor vehicle also includes a computer connected to a microphone, and the computer forms an information processing system programmed in a specific manner to implement the present invention.

[0060] More specifically, the computer includes at least one processor, a memory, various input and output interfaces, and a human-machine interface.

[0061] The computer stores computer applications, consisting of computer programs including instructions, in its memory. The processor executes these instructions, enabling the computer to implement the methods described below.

[0062] The computer can read data acquired by the microphone through its input interface.

[0063] The computer can issue commands to perform certain functions of a motor vehicle through its output interface, such as, for example, opening windows or starting the engine.

[0064] Human-machine interfaces can take many forms. Here, we will consider one with a touchscreen and a speaker located in the vehicle's passenger compartment.

[0065] As will indeed be described in the remainder of this disclosure, the invention primarily relates to identifying a speaker based on phrases produced by the speaker's voice.

[0066] Here, "phrase" refers to a group of words that make up a fixed phrase. In practice, this refers to predefined keywords.

[0067] In the example to be considered, the speaker will be the driver of the vehicle, but as a variation, it could be any other passenger.

[0068] According to the present invention, the speaker can be identified as long as the speaker has been recorded in the information processing system in advance.

[0069] The process of identifying the speaker specifically involves determining which user is producing the phrase from a pre-recorded set of vehicle users.

[0070] Therefore, the first part of this disclosure will describe the methods by which drivers can be recorded on the system. The second part of this disclosure, in itself, will concern the identification of the driver.

[0071] The recording process is performed in several consecutive steps. Its purpose is to generate a voice signature associated with the speaker.

[0072] For the driver, the first step here involves initiating the process by selecting the corresponding menu in a computer application using a touchscreen.

[0073] Once the process has started, the computer generates a request through the human-machine interface, which includes asking the driver to pronounce or even preferably repeat the same predetermined phrase several times.

[0074] This phrase is preferably chosen when the computer application is designed to meet two criteria.

[0075] The first standard is understanding the standard.

[0076] For the computer to detect every moment the driver utters the phrase, the phrase must be audible. In other words, the phrase must include low-frequency tones. Therefore, the phrase will be selected such that it includes as many vowels as possible.

[0077] The second standard is the time standard.

[0078] Specifically, the phrase must be pronounced quickly so that drivers can say it easily and quickly without causing them trouble. This criterion is met when the phrase consists of three or four syllables. This allows the phrase to be pronounced in a time frame of less than one second.

[0079] The phrase chosen here is "Hello, Renault".

[0080] During the recording process, the computer records a long speech signal, which is then segmented into three speech signals corresponding to three moments when the phrase is spoken. These three speech signals are then combined into a single recorded speech signal S41, which is considered to form a characteristic example of the driver speaking the phrase.

[0081] The computer can infer the underlying voice signature from the recorded voice signal S41 using conventional processing procedures known to those skilled in the art, which will be referred to below as the "acoustic fingerprint generation process".

[0082] The process can be concisely described as follows.

[0083] The process begins with acoustic analysis, which involves extracting relevant and feature information from the recorded speech signal. To this end, multiple sets of acoustic coefficients are calculated over fixed-length signal blocks at regular time intervals (i.e., within consecutive observation windows). These sets of coefficients together form an acoustic matrix, which constitutes a digital signature representing the driver's speech.

[0084] For example, each set of coefficients is calculated using the discrete cosine transform of the logarithm of the signal energy spectral density. In particular, the cepstral coefficients produced by this analysis do indeed characterize the shape of the spectrum.

[0085] In this example, the cepstral coefficients used are MFCCs (Mel frequency cepstral coefficients). In particular, these cepstral coefficients have the advantage of being weakly correlated with each other.

[0086] Furthermore, this process is accomplished here through Mel filter bank filtering, which can highlight the richness of the pronunciation.

[0087] Therefore, the acoustic fingerprint generation process can generate a basic voice signature representing the driver's voice based on the recorded voice signal S41.

[0088] According to the present invention, once the basic voice signature has been obtained, the computer will seek to calculate another voice signature, referred to as the extended voice signature.

[0089] The idea is that the simple phrase "Hello, Renault" is too short for a robust speaker identification from a number of recorded users using only basic voice signatures. This is especially true when the driver is in a specific pathological state (illness, mood, fatigue, etc.), when the voice recording conditions are poor (ambient noise, etc.), or when the driver utters the phrase without understanding (truncated words, etc.).

[0090] To obtain an extended voice signature, the computer first determines the extended signal.

[0091] The extended signal is intended to be appended to the recorded speech signal to extend its length, thereby obtaining a complete signal that can be processed through the acoustic fingerprint generation process to generate an extended speech signature.

[0092] This extended signal is associated with the driver. Therefore, this extended signal was selected to be different from the extended signals already used by other speakers recorded in the system.

[0093] This extended signal is generated by the parameterizable function S1(t). Figure 1 An example of this parameterizable function is shown in the figure.

[0094] The parameterizable function S1(t) is preferably the sum of at least one sine wave with a frequency between 100 and 500 Hz.

[0095] In the embodiment described herein, the parameterizable function S1(t) is represented as follows:

[0096]

[0097] In this equation, the adjustable parameter is:

[0098] -M: Number of sine waves

[0099] -A i The amplitude of each sine wave.

[0100] -f i The frequency of each sine wave, and

[0101] - i The phase of each sine wave.

[0102] This function is preferably amplitude modulation (then A) i (a function of time t) and / or frequency modulation (then f) i (a function of time t).

[0103] The set of parameters chosen to create the extended signal is selected such that the extended signals associated with each speaker are distinct from one another.

[0104] It can be assumed that the two extended signals are different from each other in frequency when at least a 20 Hz step size separates each of the two frequencies.

[0105] It can be assumed that the two extended signals are phase-differentiated when a step size of at least π / 4 radians separates each of the two phases. The amplitude can be considered close to 1 in order to maximize the presence of the extended signal in terms of frequency (energy).

[0106] These multiple sets of parameters can be randomly selected by the computer, in which case the computer will then check whether these parameters actually meet the aforementioned difference conditions.

[0107] As a variation, multiple sets of parameters can be pre-defined and recorded in computer memory. In this case, the computer will be able to access its memory and search for a new set of parameters that has not yet been used each time a new speaker is recorded.

[0108] exist Figure 1 The example shown uses the following set of parameters:

[0109] M = 3

[0110] (A1, f1, 1) = (1, 127, 0)

[0111] (A2, f2, 2) = (1, 241, 0)

[0112] (A3, f3, 3) = (1, 353, 0)

[0113] Then, the obtained parameterizable function S1(t) is modified so that once it is attached to the recorded speech signal S41, there will be no discontinuity at the intersection of the curves.

[0114] For this purpose, the method for calculating the parameterizable function S1(t) and Figure 2 The product of the predetermined observation window function (S2(t)) is shown.

[0115] The observation window function (S2(t)) here is an apodization function. This ensures that the product of the parameterizable function S1(t) and the observation window function (S2(t)) takes zero values ​​at the beginning and end of the considered time window.

[0116] In the example described here, the equation for the observation window function (S2(t)) is as follows.

[0117]

[0118] In this equation:

[0119] -x is the time period normalized relative to the length of the time window under consideration, and

[0120] -r is the cosine weighting coefficient, which is chosen to be equal to 0.25 here.

[0121] Then, the extended signal S3 is chosen to be equal to the product of the parameterizable signal S1 and the observation window function S2. This is as follows: Figure 3 As shown.

[0122] At this stage, it should be noted that the extended signal S3 is parameterized such that its maximum amplitude is less than or equal to 80% of the maximum amplitude of the recorded speech signal, and that the total length of one or more extended signals attached to the recorded speech signal S41 does not exceed 20% of the total length of the complete signal.

[0123] The complete signal is then obtained by appending the extended signal S3 to the beginning and / or end of the recorded speech signal. Here, the extended signal is appended to the beginning and end of the speech signal.

[0124] Figure 4 The complete signal S4 thus obtained is shown in the figure. It can be observed that the complete signal includes two identical signals S31 and S32, which include the recorded speech signal S41 and correspond to the extended signal S3.

[0125] It can also be observed that the recorded speech signal S41 consists of four parts S42, S43, S44 and S45, which correspond to the four syllables of the phrase "hello, Reno".

[0126] At this stage, the complete signal S4 is processed through the acoustic fingerprint generation process to obtain the extended voice signature.

[0127] Then, the extended voice signature, the basic voice signature, and the extended signal S3 used are stored in the computer's memory for association with the driver.

[0128] This association can take various forms.

[0129] Therefore, these various elements can be simply stored in a record that stores the driver's access permissions (permission to open the window, permission to request to start the engine, etc.).

[0130] Here, more precisely, it will be considered that the basic voice signature, the extended voice signature, and the extended signal S3 are recorded in three fields of a single record in the database. This record further includes a fourth field storing the driver's name (pre-entered on the touchscreen) and a fifth field storing the driver's access permissions (selected by the driver from a menu displayed on the touchscreen). Any other variations are also conceivable.

[0131] In any case, at the end of several consecutive recording processes, the computer stores a closed set of N speech signature triples (each triple includes a basic speech signature, an extended speech signature associated with one of the N speakers recorded, and an associated extended signal S3). The extended speech signature stored in the computer memory at the end of the recording process is called the reference speech signature. The basic speech signature stored in the computer memory at the end of the recording process is called the reference speech signature.

[0132] Alternatively, in order to obtain space in the computer's memory, the basic voice signature and parameters can be stored, thereby allowing an extended voice signature to be reconstructed.

[0133] We can now describe how the methods used to identify drivers are implemented.

[0134] For this purpose, two different embodiments can be described.

[0135] Figure 5 The first embodiment is shown in the figure.

[0136] When the vehicle door is unlocked, the computer is powered and enters standby mode (step E1). In this mode, the computer needs to process the data received from the microphone.

[0137] Therefore, by initiating step E2 of the recognition method, the driver verbally utters the agreed-upon phrase (here, "Hello, Renault"), and the computer can detect this phrase. The computer then records the new speech signal, captured by the microphone and containing the phrase, into its memory. This new speech signal is the recognition speech signal.

[0138] The length of the new speech signal was adjusted to the time required to say the phrase.

[0139] In step E31, the computer appends the new speech signal to the first of N extended signals recorded in its memory, that is, the extended signal associated with the first speaker recorded and stored in the first record in its database. This operation is performed in the same manner as during the recording process, i.e., appending the extended signal before and after the new speech signal.

[0140] Then, in step E41, the computer determines a new extended voice signature. This new extended voice signature is a recognized voice signature. For this purpose, an acoustic fingerprint generation process is applied to the complete signal obtained in step E31.

[0141] Finally, in step E51, the computer compares the extended voice signature with the extended voice signature in the first record stored in its database. In other words, the computer compares the identified voice signature with the reference voice signature.

[0142] The comparison step is performed in a manner known per se, namely, by comparing multiple sets of acoustic coefficients for the two signatures. This comparison determines a score, where a higher score indicates that the multiple sets of acoustic coefficients for the two signatures are closer.

[0143] Using data from N records stored in a database associated with N speakers, these three steps E31, E41, and E51 are repeated N times here (see [link to documentation]). Figure 5 Steps E32 … E3 N E42 … E4 N and E52 …E5 N ).

[0144] The computer thus obtains as many scores as the speakers recorded in its memory.

[0145] Once these scores have been calculated, in step E6, the computer compares all of these scores and selects the highest score. This highest score is associated with one of the recorded speakers (hereinafter referred to as the selected speaker).

[0146] At this stage, the computer may conclude that the driver is consistent with the selected speaker.

[0147] However, for added security, the computer compares the maximum score to a predetermined threshold in step E7.

[0148] If the maximum score is below a predetermined threshold, then via step E8, the computer displays a message on the touchscreen or transmits it to the speaker, informing the driver that these speakers have not yet been identified. Specifically, the score is considered insufficient to reliably identify whether the selected speaker truly matches the driver. In this case, it is recommended that the driver record the message or have the driver repeat the phrase.

[0149] Conversely, through step E9, the computer deems the maximum score high enough to reliably identify the selected speaker as indeed the driver. In this case, the driver has indeed been identified. They can then generate instructions, such as commands to open the windows or start the engine. These instructions will then be followed, provided the driver has the necessary access permissions.

[0150] A second embodiment of the identification method can now be described.

[0151] In this second embodiment, steps E1 and E2 are the same as those described above, and refer to... Figure 5 Describe it.

[0152] However, at the end of step E2, it is stipulated that the computer continues to calculate the basic voice signature, taking into account the new voice signal just generated by the driver. This basic voice signature is the identification basic voice signature.

[0153] The computer then compares the identified underlying speech signature with each of the reference underlying speech signatures recorded in the computer's memory. For this purpose, the computer continues in the same manner as described above, allowing it to obtain N scores.

[0154] Then, if the maximum score obtained is higher than a first predetermined threshold, the computer can consider that the driver has been identified (step E9).

[0155] Conversely, if the maximum score is below the second predetermined threshold, the computer may assume that the driver has not yet been identified and the driver will not be identified (step E8).

[0156] If the maximum score falls between these two thresholds, the computer can attempt to identify the driver, that is, then proceed as in the first embodiment, but based on an extended voice signal instead of the basic voice signal. For this purpose, the computer can implement step E31 and the steps of the first embodiment described below.

[0157] This invention is by no means limited to the embodiments described and illustrated, but those skilled in the art will be able to add any variations thereto that conform to the invention.

[0158] Specifically, it can be specified that the signature associated with the speaker is not formed by a set of acoustic coefficients as described above, but by any other elements. By way of example, the speaker's voice signature can be formed by the recorded voice signal itself (either the original signal or a signal that may have been reprocessed, for example, to remove parasitic noise).

[0159] As another variation, the extended signal may not be directly appended to the beginning or end of the speech signal recorded by the microphone, but it can be specified that a gap time interval is left between the extended signal and the speech signal. It should be noted that, preferably, the two signals will not completely or partially overlap, as overlap would reduce the reliability of the results.

[0160] As another variation, the extended signals used by the various speakers recorded in the database may be the same, but here, the consequence of this is that the reliability of the results will be further reduced.

Claims

1. A method for identifying a specific speaker from a group of speakers by means of a computer including a computer memory, the computer memory storing at least one reference speech signature associated with one of the speakers in the group, the method comprising the steps of: - Acquire the recognized speech signal generated by the specific speaker (S41). - The step of constructing a recognized voice signature based on the recognized voice signal (S41) - A comparison step that compares the identified voice signature with at least one reference voice signature recorded in the computer memory, and -The identification step of identifying the specific speaker based on the results of the comparison. The feature is that the at least one reference voice signature recorded in the computer memory is determined based on the recorded voice signal and a predetermined extended signal (S31, S32). The feature is that, prior to the construction step, a generation step is specified to generate a complete signal (S4) including the recognized speech signal (S41) and the predetermined extended signals (S31, S32), and The feature is that, in this construction step, the recognized voice signature is also constructed based on the extended signals (S31, S32).

2. The method of claim 1, wherein, The computer memory includes multiple reference speech signatures associated with multiple speakers in the group, the extended signal (S31, S32) being associated with one of these speakers and different from the extended signals associated with the other speakers, and the memory stores each extended signal for association with one of these speakers.

3. The method of claim 2, wherein: - In this generation step, the computer generates at least as many complete signals (S4) as the number of speakers in the group, each complete signal (S4) including the recognized speech signal (S41) and one of the extended signals (S31, S32) recorded in the memory. - In this construction step, the computer constructs a recognized speech signature for each complete signal (S4). - In this comparison step, the computer compares each identified voice signature with each reference voice signature recorded in the memory in order to deduce a score from them, and - In this identification step, the inferred score is considered to identify the specific speaker.

4. The method of claim 1 or 2, wherein, The complete signal (S4) is constructed by appending the extended signal (S31, S32) before and / or after the recognized speech signal (S41).

5. The method of claim 1 or 2, wherein, The extended signal (S31, S32) is a function of the sum of at least one sine wave with a frequency between 50 and 650 Hz.

6. The method of claim 1 or 2, wherein, The extended signal (S31, S32) is generated by the product of the parameterizable function (S1) and the observation window function (S2).

7. The method as described in claim 1 or 2, wherein: - The maximum amplitude of the extended signal (S31, S32) is less than or equal to the maximum amplitude of the recognized speech signal (S41), and / or - The maximum length of the at least one extended signal (S31, S32) is less than or equal to one-third of the total length of the complete signal (S4).

8. The method of claim 1 or 2, wherein, The recognized speech signal (S41) includes four or fewer syllables.

9. The method of claim 5, wherein, The frequency is between 100 and 500 Hz.

10. The method of claim 6, wherein, The parameterizable function (S1) is amplitude modulation and / or frequency modulation.

11. The method of claim 7, wherein, The maximum amplitude of the extended signal (S31, S32) is less than or equal to 80% of the maximum amplitude of the recognized speech signal (S41).

12. The method of claim 7, wherein, The maximum length of the at least one extended signal (S31, S32) is equal to 20% of the total length of the complete signal (S4).

13. A method for recording a specific speaker using a computer including computer memory, the method comprising the steps of: - Acquire the recorded speech signal produced by that specific speaker. - Determine the extended signal, - Generate a complete signal that includes the recorded speech signal and the extended signal. - Determine the reference speech signature based on the complete signal of the record, and - Store the reference voice signature in the memory so as to associate it with the specific speaker.

14. A motor vehicle comprising a passenger compartment, means for acquiring speech signals generated by a specific speaker located in the passenger compartment, and a computing unit programmed to perform the method according to any one of claims 1-13.