METHOD FOR DETERMINING A VALUE OF A VOICE DESCRIPTOR OF AN INDIVIDUAL'S VOICE

A machine learning-based method for determining voice descriptors independently of recording equipment configuration addresses inefficiencies in current vocal assessments, enabling automated and accurate characterization for improved vocal health and performance.

FR3162906A1Pending Publication Date: 2025-12-05FADI DAHDOUH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
FR2024005601
Authority / Receiving Office
FR · FR
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-30
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

Current vocal assessments are inefficient and require expert intervention, making it difficult to accurately characterize an individual's voice and identify vocal pathologies for effective treatment and equipment configuration.

Method used

A computer-implemented method using machine learning models to determine intrinsic voice descriptors from audio signals, independent of recording equipment configuration, allowing characterization without expert intervention.

Benefits of technology

Enables efficient and accurate characterization of voice descriptors and physiological states, facilitating improved vocal health and performance through automated analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

METHOD FOR DETERMINING A VALUE OF A VOICE DESCRIPTOR OF AN INDIVIDUAL'S VOICE The invention relates to a computer-implemented method (100) for determining an intrinsic characteristic value of a voice descriptor of an individual's voice, said voice descriptor belonging to a set of reference voice descriptors, the method (100) comprising: receiving (REC) at least one input audio signal encoding a common sequence of at least one sound emitted by said individual's voice; implementing (IMP1) a first learned function generated from a first machine learning model previously trained from a first training domain. Figure for the abstract: Fig. 3
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: METHOD FOR DETERMINING A VALUE OF A VOICE DESCRIPTOR OF AN INDIVIDUAL'S VOICE Scope of the invention

[0001] The invention relates to the field of computer-assisted voice analysis. More specifically, the invention relates to a method for determining a value of a voice descriptor of an individual's voice. State of the art

[0002] In many fields, it is very valuable to gather in-depth analysis and a rich understanding of the human voice. For example, in the field of vocal arts, performers need vocal preparation to optimize their performance, or to ensure their performance while preserving their health, for example, by combating or protecting themselves against vocal strain and, more generally, by protecting their vocal instrument. In the field of health, a vocal assessment performed by a healthcare professional makes it possible to identify damaged areas of an individual's vocal tract and to diagnose a vocal pathology, in other words, one related to the vocal tract. Such vocal assessments are currently carried out by professionals or groups of professionals with expertise in the analysis of the human voice.Characterizing the human voice is, in general, a non-trivial problem and depends in particular on the person whose voice one wants to characterize, the equipment recording the voice, the recording configuration, the recording location, etc.

[0003] Finally, various functional elements constituting an aspect of the phonatory apparatus can have an influence on the voice, including the vocal cords, lungs, thorax, trachea, larynx, soft palate, hard palate, nasal cavity, pharynx, teeth, lips, tongue, oral cavity.

[0004] Characterizing an individual's voice and identifying the elements responsible for a pathology on the phonatory apparatus would allow the application of the most appropriate treatments or would allow the choice of equipment configurations and arrangement of said equipment, for example, during audio recordings to be made.

[0005] A characterization of the voices would allow us to: - to achieve better protection of the vocal cords in order to prevent damage due to intensive use without preparation. - improve performance through appropriate vocal warm-ups to enhance the flexibility and capacity of the vocal cords; - reduce tension which can affect voice quality and cause vocal fatigue.

[0006] There is currently a need to be able to perform such vocal assessments more efficiently and quickly by more accurately characterizing an individual's voice. The invention improves this situation by making it possible to objectify the characteristics of the voice. Summary of the invention

[0007] A first aspect of the invention relates to a computer-implemented method for determining an intrinsic characteristic value of a first voice descriptor of an individual's voice, said voice descriptor belonging to a set of reference voice descriptors, the method comprising:

[0008] - reception of at least one input audio signal, each input audio signal encoding a common sequence of at least one sound emitted by the individual's voice;

[0009] - implementation of a first learned function generated from a first machine learning model trained from a first training domain comprising at least two training audio signals encoding the same sequence called the training sequence and acquired by two different microphones respectively, the first learned function receiving as input at least one input audio signal, so as to generate as output, for each of the input audio signals, a corresponding value of at least one first voice descriptor relative to the microphone that acquired the individual's voice,

[0010] - determination of an intrinsic characteristic value of the first descriptor vocal of an individual's voice from the corresponding value(s) of the first voice descriptor relative to each microphone generated during the implementation step.

[0011] Thus, the invention makes it possible to obtain information on one or more vocal descriptors of an individual's voice, from audio signals constituting a representation of the individual's voice, without the intervention of a voice expert such as a vocal coach or expert voice practitioner, or a voice osteopath. Such information can advantageously be used in various applications, such as for the evaluation and coaching of artists, therapeutic applications for voice pathologies, or multimedia applications where voice samples are used.

[0012] An advantage of the invention is to take advantage of a hardware configuration and microphone arrangement that allows for the qualification and characterization of voice descriptors, so that when the voice is recorded by a single reference microphone, it will be possible to faithfully reproduce the value of the descriptor from a A single microphone or a few microphones can be used. This is possible thanks to a machine learning algorithm trained on a predefined microphone configuration, which will associate the actual value of the descriptor with the single reference microphone used. One advantage is the ability to accurately characterize the voice using just one or a few microphones.

[0013] One advantage is to determine an intrinsic value of a voice descriptor independently of the material used, its arrangement, for example its position or its type.

[0014] The characteristic value of such a descriptor made independent of the recording microphone is called an intrinsic characteristic value.

[0015] According to one embodiment, when a single input audio signal is acquired by means of a single microphone, the corresponding value of the first generated descriptor corresponds to the intrinsic characteristic value of the first descriptor.

[0016] The equipment used to implement the method can be an electronic terminal such as a personal computer, a smartphone, a tablet, or any other device containing a computer. For example, the method is performed on a remote server.

[0017] According to one embodiment and for all aspects of the invention, at least one vocal descriptor is selected from the list of descriptors: {a quantification of the attack of the sound, a quantification of the pitch, a quantification of air on the voice, a quantification of the termination of the sound, a quantification of projectivity, a quantification of harmonic mixing, a quantification of a degree of laterality, a quantification of a degree of vibrato}.

[0018] In certain embodiments, the at least one input audio signal comprises at least two input audio signals corresponding to recordings acquired by two microphones arranged in two different acquisition positions, said recordings corresponding to audio signals from the common sequence of at least one sound emitted (S) by the individual's voice, the method further comprising: • implementation of a function called global function receiving as input each corresponding value of the voice descriptor of said voice of said individual acquired by each microphone, so as to generate as output said intrinsic characteristic value of the voice descriptor of the individual's voice.

[0019] Thus, the global function makes it possible to determine the intrinsic characteristic value of the voice descriptor on the basis of at least two corresponding signal values input audio and therefore, with increased accuracy with the number of input audio signals.

[0020] In certain embodiments, the process includes, prior to the implementation step of the first learned function:

[0021] - receiving configuration information including, for each audio signal among at least one input audio signal, position information from an acquisition microphone associated with said input audio signal relating to a position of the individual when they emitted the sequence of at least one sound,

[0022] the first learned function also receiving as input all the position information of all the input audio signals.

[0023] Thus, the corresponding value(s) generated by the first learned function are obtained using information relating to the context of the acquisition of the input audio signals. More precisely, the corresponding value generated for a given input audio signal depends on the acquisition parameters of that input audio signal, such as the position of the microphone that acquired that input audio signal.

[0024] In some embodiments, the global function is one of the following: the maximum function, the minimum function, an averaging function, a weighted sum of the corresponding values ​​of the voice descriptor of each of the input audio signals.

[0025] Thus, the global function allows the contribution of each microphone, among the corresponding values ​​obtained during the implementation step of the first learned function, to be adjusted to the intrinsic characteristic value. In this way, when configuration information has also been received, the global function allows the contribution of different microphones that acquired the different input audio signals to be adjusted to the characteristic value of the voice descriptor.

[0026] In some embodiments, the method is a method for determining an intrinsic characteristic value of at least two different vocal descriptors of the individual's voice.

[0027] In some embodiments, the reference voice descriptors are normalised on the same scale of values.

[0028] For example, the intrinsic or non-intrinsic characteristic value of a voice descriptor is a score out of a predetermined number of points. In another example, the intrinsic or non-intrinsic characteristic value of a voice descriptor is a percentage.

[0029] Thus, it is possible, thanks to the normalization of voice descriptors, to compare the characteristic values ​​of different voice descriptors.

[0030] In certain embodiments, the reference vocal descriptors include: a sound attack, a note pitch, an amount of air over the voice, a sound termination, a harmonic mixture, a degree of laterality, a degree of projectivity, a degree of vibrato. In addition, reference vocal descriptors may include an indication of the location or proportion of the sound produced in the larynx, in the pharynx, for example at the level of the hypopharynx, oropharynx or nasopharynx.

[0031] In some embodiments, the first training domain comprises a set of training audio signals, each training audio signal being pre-labeled with at least one label, the at least one label comprising a value of at least one voice reference descriptor.

[0032] According to one example, each training audio signal was previously labeled with at least one label by a panel of experts.

[0033] Thus, by the nature of the training of the first machine learning model, the result provided by the combination of the first learned function and the global function makes it possible to reproduce values ​​of voice descriptors as close as possible to the estimation of these descriptors by experts.

[0034] According to one embodiment, the method comprises training the model including: • A plurality of recordings of a sequence of at least one sound by a plurality of microphones, each microphone being associated with a recording configuration and being arranged in a predefined position relative to a reference recording position; • An estimate of the corresponding value of the voice descriptor of said voice of said individual from each microphone; • An estimate of the intrinsic characteristic value of a vocal descriptor of said voice; • The association of said estimated value of the intrinsic characteristic value of a voice descriptor of said voice with each corresponding value of the voice descriptor of said voice of said individual of each microphone; • Learning the neural network by modifying the coefficients of said network based on associations of said values.

[0035] One advantage is having a model trained with numerous microphone configurations. This offers a twofold advantage. The first is having a model capable of performing well with an arbitrary recording configuration, that is, one comprising a given number of microphones, for example, a studio with 7 microphones. Thus, all the tracks can be used by the first learned function to produce corresponding values ​​for each microphone. The global function then generates the results in a second step. an intrinsic characteristic value. The second advantage is that even with a single recording of a track, for example by a reference microphone, the first learned function makes it possible to generate a corresponding value of the descriptor which is intrinsic to the voice and little dependent on the recording device.

[0036] In some embodiments, the at least one input audio signal consists of a single input audio signal.

[0037] Thus, advantageously, the invention makes it possible to determine one or more characteristic values ​​of vocal descriptors of an individual's voice from a single recording of it.

[0038] In certain embodiments, at least one input audio signal is / are acquired by means of several tracks, generally between 1 and 13 tracks. Each track allows the acquisition of one input audio signal. Therefore, there can be as many input audio signals as there are tracks provided for this purpose. According to one embodiment, between 3 and 4 tracks are configured to acquire input audio signals from 3 or 4 microphones. The invention relates to the case where a single microphone would be used to record one track, that is to say, a single input audio signal.

[0039] In some embodiments, the at least one input audio signal consists of thirteen input audio signals.

[0040] In some embodiments, the process includes, before the reception step, an acquisition step, by an acquisition system, of at least one input audio signal.

[0041] A second aspect of the invention relates to a programmable device configured to implement the process described above.

[0042] A third aspect of the invention relates to a system for determining an intrinsic characteristic value of a voice descriptor of an individual's voice and configured to implement the method described above, comprising:

[0043] - an acquisition system configured to implement the acquisition step, including at least one microphone, the at least one microphone being configured to simultaneously acquire the common sequence of at least one sound emitted by said voice of said individual, so as to generate at least one input audio signal;

[0044] - the programmable device described above.

[0045] In some embodiments, each microphone among the at least one microphone is identical to the others.

[0046] In certain embodiments, at least one microphone of the acquisition system comprises at least one of the following:

[0047] - a reference microphone positioned, in operation, facing and at the height of the mouth of the individual when the latter emits the sequence of at least one sound, at a distance between 1 cm and 40 cm from the mouth of individual I, the orientation of the The individual's body, from their back towards their torso, defines a direction of orientation and a line of orientation. A distance of between 5 cm and 30 cm between the microphone and the mouth is optimal for recording a track using the reference microphone. The reference microphone is preferably oriented towards the individual's mouth.

[0048] - a first harmonic measurement microphone and a second microphone The first harmonic measurement microphone is positioned in front of the individual along a vertical line containing the reference microphone and located at a distance of between 1 cm and 5 m from a vertical plane containing the individual. The first harmonic measurement microphone is located above the reference microphone. For example, the first harmonic measurement microphone is positioned above the reference microphone at a height greater than the reference microphone, such as 5 cm to 60 cm higher than the reference microphone. It is also called an "overhead" microphone in English terminology.

[0049] - The second harmonic measurement microphone is located below the The height of the reference microphone is, for example, lower than the height of the reference microphone, for example, a height of the reference microphone reduced by 5 cm to 60 cm. This microphone is advantageously placed at a similar distance to that of the overhead microphone, but oriented upwards towards the mouth.

[0050] The distance between the first harmonic microphone and the second harmonic microphone is preferably between 5 cm and 1.4 m and preferably between 30 cm and 1.4 m.

[0051] - A first microphone for measuring projectivity or directivity and a second A microphone measuring projection or directionality is positioned in a vertical plane in front of the individual. This vertical plane is advantageously located at a distance from the vertical plane containing the individual, ranging from a few centimeters to several meters depending on the room configuration in which the individual's voice is recorded. The height of the first microphone measuring projection or directionality is between 10 cm and 5 m from the height of the individual's mouth, and preferably between 30 cm and 2 m. The height of the second microphone measuring projection or directionality is also between 10 cm and 5 m from the height of the individual's mouth, and preferably between 30 cm and 2 m, to avoid excessive or muddy reverberation.

[0052] - a first microphone for measuring laterality and a second microphone for Laterality measurement: the microphones are positioned approximately in a vertical plane containing the reference microphone and perpendicular to the orientation direction, and at the height of the individual's mouth, on either side of the orientation line. "Approximately" means that a shift of a few centimeters can be considered as in the same plane. According to another embodiment, the measurement of laterality can be carried out in a plane parallel to the plane containing the reference microphone and located in the portion of space in front of the individual's face.

[0053] - a reflectivity measurement microphone positioned behind the individual along of the orientation line. Such a microphone can be positioned at different distances from the vertical plane containing the individual, depending on the dimensions of the room in which the recording is made. It can be located, for example, at a distance of between 30 cm and 10 m.

[0054] - a first ambient projectivity measurement microphone and a second Ambient projection microphones are positioned in a vertical plane in front of the individual. This vertical plane is located between 5 cm and 10 m from the vertical plane containing the individual. The first ambient projection microphone is positioned between 1 m and 10 m from the height of the individual's mouth, and the second ambient projection microphone is positioned between 1 cm and 10 m from the height of the individual's mouth. Ambient projection or directional microphones are preferably placed at a height of approximately 2 to 3 meters above the floor. This height allows for a good balance between direct and indirect reflections from the room, avoiding the capture of excessive reflections from the floor or ceiling.If the room is particularly tall with a high ceiling, the microphones can be placed even higher to capture more of the space's natural reverberation.

[0055] - a first ambient reflectivity measurement microphone and a second Ambient reflectivity measurement microphones are positioned in a vertical plane behind the individual. This vertical plane is located between 5 cm and 10 m from the vertical plane containing the individual, depending on the dimensions of the recording room. The first ambient projectivity measurement microphone, M8, is positioned between 1 cm and 5 m from the height of the individual's mouth, depending on the dimensions of the recording room. The second ambient projectivity measurement microphone is also positioned between 1 cm and 5 m from the height of the individual's mouth, depending on the dimensions of the recording room.

[0056] - a microphone called a laryngeal pickup microphone, for example a A stethoscope or laryngophone is positioned near the individual's throat. Such a microphone captures direct skin vibrations on the throat, where the larynx is located, without picking up ambient sounds from the mouth or other sources.

[0057] In certain embodiments:

[0058] - the vocal descriptor is the attack of the sound; this descriptor allows us to characterize how a note is initiated and produced by the voice.

[0059] More specific descriptors can also be defined, such as "soft attack," which is characterized by a gradual onset and generally a slight breath preceding the note. One advantage of the invention is its ability to characterize the gradation and amount of breath present, particularly through training different types of soft attacks found in a set of audio recordings of voices. Thus, the machine learning model makes it possible to deduce a characteristic output from a new input defining an individual's voice.

[0060] Similarly, a hard attack can be characterized by a sound that begins with a well-defined consonant and a rapid closure of the vocal cords, producing a clear and precise sound from the very beginning of the note. Quantifying the sound at the start, identifying the presence of a consonant, and the duration of vocal cord closure allows for the characterization and / or labeling of recordings in order to train a machine learning model.

[0061] Similarly, the glottal attack is characteristic of a glottal stop, where the vocal cords are abruptly closed before sound production, creating subglottal tension. Consequently, this descriptor can be characterized using a machine learning model trained with labeled data characterizing, in particular, the closure time of the vocal cords before sound production.

[0062] Vocal cord closure measurements can be performed using an imaging device positioned near the mouth or nose, for example using a laryngoscope.

[0063] Here are indicated some preferred associations between particular microphones and certain descriptors in order to train the model during the learning of the first learned function or to best characterize the descriptors during the exploitation phase of the first learned function. These associations are advantageously made during training to select the microphones of interest for learning the first function. During the exploitation of the learned function, only one microphone may be present for acquiring an individual's voice; however, some recording locations may use different microphones in order to refine the estimation of one or more vocal descriptor(s) of an individual.

[0064] Preferably, to characterize the descriptor relating to the attack of the sound, at least one microphone of the acquisition system includes for example the reference microphone and another microphone such as a harmonic microphone or a projection microphone.

[0065] In certain embodiments:

[0066] - the voice descriptor is a quantification of the pitch of the note,

[0067] - at least one microphone of the acquisition system is, for example, at least one harmonics microphone and possibly also the reference microphone.

[0068] In certain embodiments: - The vocal descriptor is a quantity of air over the voice, - at least one microphone in the acquisition system includes, for example, a reference microphone and a projection microphone.

[0069] In certain embodiments:

[0070] - the voice descriptor is a quantification of the sound ending,

[0071] - at least one microphone of the acquisition system includes a microphone of reference and a projection and / or harmonics microphone.

[0072] In certain embodiments:

[0073] - the vocal descriptor is a quantification of a harmonic mixture,

[0074] - at least one microphone of the acquisition system comprises, for example, two harmonic microphones or for example a reference microphone and a harmonic microphone.

[0075] In certain embodiments:

[0076] - the voice descriptor is a quantification of a degree of laterality,

[0077] - at least one microphone of the acquisition system includes at least one directional or lateral microphone and possibly also the reference microphone.

[0078] In certain embodiments:

[0079] - the voice descriptor is a quantification of a degree of projectivity,

[0080] - at least one microphone of the acquisition system includes at least one projectivity microphone and possibly also the reference microphone.

[0081] In certain embodiments:

[0082] - the vocal descriptor is a quantification of a degree of vibrato,

[0083] - at least one microphone of the acquisition system includes at least one laryngeal microphone, also called a "stethomicrophone" or laryngophone.

[0084] A fourth aspect of the invention relates to a device for characterizing at least one physiological descriptor of an individual's voice, comprising: - at least one input interface configured to receive at least one characteristic value of a voice descriptor of the individual's voice; - at least one configured processing unit implemented by computer:

[0085] - an implementation step of a second learned function generated from a second machine learning model trained from a second domain training, the second learned function receiving as input at least one characteristic value of the voice descriptor;

[0086] the implementation step of the second learned function allowing to generate as output at least one score characterizing at least one physiological descriptor of the individual's voice.

[0087] - at least one output interface configured to return at least one score characterizing at least one physiological descriptor of said voice of said individual.

[0088] In this case, the characteristic value of a voice descriptor of an individual's voice can be an intrinsic characteristic value or one measured by a microphone. It should be noted that the intrinsic characteristic value of a descriptor is calculated from a function learned from several microphones in such a way as to produce a characteristic value of the descriptor that is independent, or as independent as possible, of the microphone(s) that acquired the voice.

[0089] According to this fourth aspect of the invention, the characteristic value of a descriptor may be intrinsic or not. It is more generally called the characteristic value of the descriptor.

[0090] Thus, thanks to this device, it is possible, from characteristic values ​​of vocal descriptors, such as the attack of the sound, the air on the voice, or even the pitch of the note, to characterize a physiological descriptor of the voice of an individual, without the intervention of an expert voice practitioner, such as a voice osteopath.

[0091] In certain embodiments, at least one processing unit is further configured to implement by computer:

[0092] - an implementation step of a comparison function configured for compare at least one intrinsic or non-intrinsic characteristic value of the voice descriptor to a threshold value,

[0093] the implementation step of the comparison function allowing to generate as output at least one score characterizing at least one physiological descriptor of the individual's voice.

[0094] Thus, thanks to the invention, it is possible to characterize a physiological descriptor of an individual's voice from a single characteristic value of a vocal descriptor, and this without the intervention of an expert voice practitioner, such as a voice osteopath.

[0095] In certain embodiments, at least one physiological descriptor is included among: vocal strain, muscular strain, glottal stops. The physiological descriptor corresponds to a quantification of vocal strain, muscular strain, or one or more glottal stops.

[0096] In certain embodiments, the voice descriptor is included among a set of reference voice descriptors comprising: a quantification of the attack of the sound, a quantification of a pitch, a quantity of air on the voice, a quantification of a sound termination, a quantification of a harmonic mixture, a degree of laterality, a degree of projectivity, a degree of vibrato.

[0097] In certain embodiments, at least one processing unit is further configured to implement by computer, following the implementation step:

[0098] - an implementation step of a third learned function generated from of a third machine learning model trained from a third training domain, the third learned function receiving as input at least one score characterizing at least one physiological descriptor of the individual's voice, so as to generate as output at least one percentage value characterizing a contribution of at least one stage of the phonatory apparatus of said individual to said physiological descriptor.

[0099] In some embodiments, the at least stage is one of: the lung of the individual, the larynx of the individual, the pharynx of the individual, an articulator of the individual.

[0100] A fifth aspect of the invention relates to a method for characterizing at least one physiological descriptor of an individual's voice implemented by the device described above, comprising:

[0101] - a step of receiving at least one characteristic value of a voice descriptor of the said voice of the said individual;

[0102] - an implementation step of a second learned function generated from a second machine learning model trained from a second training domain, said second learned function receiving as input said at least one characteristic value of the voice descriptor;

[0103] the implementation step of the second learned function allowing to generate as output at least one score characterizing said at least one physiological descriptor of said voice of said individual.

[0104] According to this aspect of the invention, the method can be carried out with an intrinsic or non-intrinsic characteristic value of at least one descriptor.

[0105] A sixth aspect of the invention relates to a method for characterizing a contribution of at least one level of the vocal tract to a physiological descriptor of an individual's voice implemented by the device described above, comprising:

[0106] - the steps of the process for characterizing at least one physiological descriptor of the voice of the aforementioned individual previously described;

[0107] - a step called an additional step for implementing a third function learned generated from a third machine learning model trained from a third training domain, said third learned function receiving as input said at least one score characterizing said at least one physiological descriptor of said voice of said individual, so as to generate as output at least one percentage value characterizing a contribution of at least one level of the phonatory apparatus of said individual to said physiological descriptor.

[0108] A seventh aspect of the invention relates to a method for predicting a voice pathology in an individual, the method comprising:

[0109] - the steps of the process to characterize at least one physiological descriptor of the voice of the individual previously described,

[0110] - the steps of the process for characterizing a contribution of at least one stage of the vocal apparatus of said individual, as previously described, is characterized by a physiological descriptor.

[0111] - a step for predicting the pathology of the individual's voice based on the contribution from at least one level of the individual's vocal apparatus.

[0112] Different types of measurement microphones can be used and configured for recording an individual's voice. They can be used in combination when different types of measurement microphones are used.

[0113] Other types of microphones are presented here in addition to those already described.

[0114] A first type of measurement microphone provides a flat frequency response. These measurement microphones have a very flat frequency response. This means that they are designed to respond equally to all audible frequencies, without coloration or emphasis of any particular frequency. This characteristic makes it possible to obtain accurate measurements of the sound as it is, without distortion introduced by the microphone itself.

[0115] A second type of measurement microphone allows for a measurement of directivity. These allow for an estimation of sound dispersion.

[0116] A third type of measurement microphone with high sensitivity and good accuracy. These allow for the measurement of audio signals with low sound levels. This characteristic enables precise acoustic measurements, including noise level measurements.

[0117] Simple measurement microphones can also be used, such as studio microphones. Brief description of the figures

[0118] Other features and advantages of the invention will become apparent from the following detailed description, with reference to the accompanying figures, which illustrate:

[0119] [Fig-1]: an example of an acquisition system configured to acquire a set of audio signals that can be used in a process to determine a value intrinsic characteristic of a voice descriptor of an individual's voice according to the invention;

[0120] [Fig.2]: an example of a device called a descriptor device configured to implement the method for determining an intrinsic characteristic value of a voice descriptor of an individual's voice according to the invention;

[0121] [Fig.3]: an example of a set of steps that can be carried out to implement the method for determining an intrinsic characteristic value of a voice descriptor of an individual's voice according to the invention;

[0122] [Fig.4]: an example of a system configured to implement the method for determining an intrinsic characteristic value of a voice descriptor of an individual's voice according to the invention;

[0123] [Fig.5]: an example of a device called a pathology device configured to implement a method for characterizing at least one physiological descriptor of an individual's voice according to the invention;

[0124] [Fig.6] an example of a set of steps that can be carried out to implement the method for characterizing at least one physiological descriptor of an individual's voice according to the invention. Description of the invention

[0125] One aspect of the invention relates to a method 100, a device called a descriptor device 10, and a system 20 for determining a value D of a vocal descriptor of an individual's voice. One aspect of the invention also relates to a method 200 and a device called a pathology device 30 for characterizing at least one physiological descriptor of an individual's voice. One aspect of the invention also relates to a method for characterizing the contribution of at least one level of the vocal tract to a physiological descriptor of an individual's voice. A further aspect of the invention relates to a method for predicting a pathology of an individual's voice.

[0126] Fig. 1 represents an example of an acquisition system 15 configured to acquire and record a sequence of at least one sound S emitted by the voice of an individual I, so as to obtain at least one audio signal SAb SA2,... SAN encoding the sequence of at least one sound, with N an integer greater than or equal to 1.

[0127] In operation, the acquisition system 15 is placed in an indoor space, such as a recording room, a studio, an examination room.

[0128] The acquisition system 15 includes at least one microphone positioned to acquire the sequence of at least one sound S, the microphone being positioned at a predetermined acquisition position (or location).

[0129] When the acquisition system 15 includes at least two microphones, these are positioned at distinct acquisition positions.

[0130] Some embodiments of the acquisition system 15 will be described below, which can be combined together.

[0131] In some embodiments, the acquisition system 15 comprises between 1 and 13 microphones.

[0132] In some embodiments such as that illustrated in [Fig.1], the acquisition system 15 comprises 13 microphones (Mo, Mb M2... Mi2).

[0133] In some embodiments, the acquisition system 15 comprises a single microphone.

[0134] When individual I emits the sequence of at least one sound S for a recording by the acquisition system 15, the orientation of his body from back to front, defining a direction called orientation direction x and a line called orientation line passing through the center of gravity of the individual and parallel to the orientation direction x, and the body of individual I defining approximately a plane P perpendicular to the direction x.

[0135] Advantageously, a microphone Mo serving as a geometric reference and hereinafter referred to as the reference microphone Mo is positioned, in operation, in front of the mouth of the individual I whose sequence of at least one sound S is acquired and recorded, the individual being in a standing position when emitting the sequence of at least one sound S. The reference microphone Mo is located in a plane Pi distant from the plane P by a distance of between 1 cm and 40 cm and preferably between 5 cm and 30 cm.

[0136] In certain embodiments, the acquisition system 15 comprises a first harmonic measurement microphone Mi and a second harmonic measurement microphone M2, positioned at acquisition locations on the same vertical line positioned in front of the individual when he or she emits the sequence of at least one sound S. By "vertical line" is meant a line oriented along the vertical axis, i.e., the axis of gravity. For example, the first harmonic measurement microphone Mi is positioned on the vertical line below the reference microphone Mo and the second harmonic measurement microphone M2 is positioned on the vertical line above the reference microphone Mo.

[0137] The concepts of "below" and "above" are interpreted with respect to a vertical line extending from the floor to the ceiling of a room and passing through the reference microphone. The position on the z-axis gives the altitude and allows any object to be positioned along this axis and relative to the reference microphone. Since the altitude of the reference microphone is generally adjusted according to the size of the individual and the position From the mouth, any object in the room can also be positioned on the vertical axis opposite an individual's mouth.

[0138] Advantageously, the acquisition position of the first harmonic measurement microphone Mb placed above the mouth allows the acquisition of an audio signal comprising more information on the high harmonic components contained in the sequence of at least one sound emitted by individual I.

[0139] Advantageously, the acquisition position of the second harmonic measurement microphone M2, placed below the mouth, allows the acquisition of an audio signal comprising more information on low harmonic components contained in the sequence of at least one sound S emitted by the individual I.

[0140] According to one example, the first harmonic measurement microphone Mi and the second harmonic measurement microphone M2 are positioned at a distance of between 40 and 60 cm from the reference microphone Mo. Preferably, the first harmonic measurement microphone Mi and the second harmonic measurement microphone M2 are positioned at a distance of 50 cm from the reference microphone Mo.

[0141] In certain embodiments, the acquisition system 15 comprises a first projectivity measurement microphone M3i and a second projectivity measurement microphone M4b positioned in front of the individual along the x-axis. The two microphones can be positioned on vertical planes at different depths of the room relative to the individual. For example, a first projectivity microphone is positioned 2 m from the individual in the horizontal plane containing the reference microphone, and a second projectivity microphone is positioned 4 m from the individual in the horizontal plane containing the reference microphone.

[0142] By their positions, the first projectivity measurement microphone M3i and the second projectivity measurement microphone M4[ are configured to optimally measure the projectivity of the voice of individual I. In other words, the acquisition positions of the first projectivity measurement microphone M3i and the second projectivity measurement microphone Mu allow the acquisition of audio signals including information characterizing the degree of projectivity in the sequence of at least one sound S emitted by individual I.

[0143] By "projectivity" is meant a quantification of an individual's ability to carry their voice far into space and to be heard clearly at a distance. Projectivity is influenced by several factors, including vocal power, i.e., the force with which sounds are produced by the vocal cords; resonance, i.e., the effective use of resonating cavities such as the mouth, nose, or pharynx to amplify sound; and the appropriate use of breathing and support. diaphragmatic. Finally, clarity of diction, that is to say the ability to articulate sounds distinctly, can also influence projectivity.

[0144] The "directivity" of the voice refers to the way in which the sound of the voice propagates in a direction in space. It describes the spatial distribution of sound energy. It can be quantified with different measurement microphones. Thus, it is possible to measure lateral directivity, that is, the dispersion of a sound relative to the left or right of the individual. This angle can be measured from the individual's mouth and located relative to a vertical plane perpendicular to the plane of the individual's chest or the plane parallel to the orientation of their nose. It is also possible to measure vertical directivity, that is, the dispersion of the sound relative to the top or bottom of the mouth.

[0145] Voice directionality can be influenced by several factors such as the shape of the mouth during phonation, the orientation of the mouth, the orientation and shape of the lips, etc.

[0146] For example, the first directivity measurement microphone M3 and the second directivity measurement microphone M4 are positioned on a second vertical line, one of them being positioned above the other. In one example, the first projectivity measurement microphone M3 and the second projectivity measurement microphone M4 are located at a distance from plane P of between 1 cm and 10 m. The distance between the two planes can be adapted according to the dimensions of the recording room.

[0147] For example, a third directivity measurement microphone M3' and a fourth directivity measurement microphone M4' are positioned on a second vertical line, one of them positioned above the other. In one example, the first directivity measurement microphone M3' and the second directivity measurement microphone M4' are located at a distance from plane P of between 1 cm and 10 m. The distance between the two planes can be adjusted according to the dimensions of the recording room.

[0148] For example, a fifth directivity measurement microphone M3” and a sixth directivity measurement microphone M4” are positioned at the same height within a horizontal plane, for example, on the right side of the individual and on the left side of the individual. In one example, the first directivity measurement microphone M3” and the second directivity measurement microphone M4” are separated by a distance of between 1 cm and 10 m in a horizontal plane. The distance between the two planes can be adjusted according to the dimensions of the recording room.

[0149] Projectivity and directivity are two related measurements, insofar as the more directional an audio signal is, the greater its projection will be. Conversely, the more projective a signal is, the lower its directivity will be.

[0150] The dispersion of a signal can be heterogeneous in a solid projection cone. For example, the audio signal can have lateral directivity, that is to say, a characterization of the distribution of audio power according to azimuth or height.

[0151] Laterality measurement quantifies the dispersion of a signal in the horizontal plane on either side of the individual. This measurement is very similar to the measurement of lateral directivity. Laterality measures the amount of signal power distributed laterally to the individual, that is, the variation in intensity in the horizontal plane.

[0152] In certain embodiments, the acquisition system 15 comprises a first laterality measurement microphone M5 and a second laterality measurement microphone M6, positioned in the plane Ph plane of the reference microphone Mo, and at the same height along the vertical axis, i.e., the height of the mouth of individual I. By their positions, the first laterality measurement microphone M5 and the second laterality measurement microphone M6 are configured to optimally measure the degree of laterality of the voice of individual I. In other words, the acquisition positions of the first laterality measurement microphone M5 and the second laterality measurement microphone M6 make it possible to acquire audio signals comprising information characterizing the degree of laterality in the sequence of at least one sound S emitted by individual I.

[0153] According to one example, the first laterality measurement microphone M5 and the second laterality measurement microphone M6 are each located at a distance from the reference microphone Mo of between 5 cm and 2 m, and preferably between 40 cm and 60 cm. Preferably, the first laterality measurement microphone M5 and the second laterality measurement microphone M6 are each located at a distance of 50 cm from the reference microphone Mo.

[0154] In certain embodiments, the acquisition system 15 includes a reflectivity measurement microphone M7 positioned behind the individual I (along the x-axis). The reflectivity measurement microphone M7 is positioned along the x-axis at the same height as the reference microphone M0. By its position, the reflectivity measurement microphone M7 is configured to optimally measure the degree of reflectivity of the individual I's voice. In other words, the acquisition position of the reflectivity measurement microphone M7 allows for the acquisition of an audio signal containing information characterizing the degree of reflectivity in The sequence of at least one sound S emitted by individual I. Reflectivity is understood as the quantification of the power of the audio signal reflected from an obstacle, such as a wall, ceiling, or other object. For example, the reflectivity measurement microphone M7 is located at a distance of between 50 cm and 4 m from the plane Pi of the reference microphone Mo. The plane containing the reflectivity measurement microphone M7 is preferably located behind the individual, that is, in the portion of space on the side of the vertical plane containing the individual and positioned on the opposite side of their face.

[0155] In certain embodiments, the acquisition system 15 comprises a first ambient projectivity measurement microphone M8 and a second ambient projectivity measurement microphone M9, positioned in front of the individual I. By their positions, the first ambient projectivity measurement microphone M8 and the second ambient projectivity measurement microphone M9 are configured to optimally measure the degree of laterality of the voice of the individual I. In other words, the acquisition positions of the first ambient projectivity measurement microphone M8 and the second ambient projectivity measurement microphone M9 make it possible to acquire audio signals comprising information characterizing the degree of ambient projectivity in the sequence of at least one sound S emitted by the individual I."Ambient projectivity" refers to the ability of a sound to fill a space in such a way as to create a specific atmosphere or ambiance, taking into account how the acoustic characteristics of the environment affect the perception of the sound. Thus, ambient projectivity allows us to jointly quantify reflection, reverberation, and absorption by surrounding surfaces, as well as the uniformity of distribution—that is, the ability of the sound to reach all points in space in a balanced way.

[0156] According to one example, the first ambient projectivity measurement microphone M8 and the second ambient projectivity measurement microphone M9 are located at a distance from plane Pi (plane of the reference microphone Mo) of between 10 cm and 6 m, and preferably between 2 m and 5 m. However, the distance is adapted according to the dimensions of the room in which the recording takes place. Advantageously, the first ambient projectivity measurement microphone M8 and the second ambient projectivity measurement microphone M9 are positioned at a height higher than the height of the reference microphone Mo, preferably higher by a distance of between 10 cm and 2 m, preferably higher by 50 cm.

[0157] In certain embodiments, the acquisition system 15 comprises a first ambient reflectivity measurement microphone Mi0 and a second ambient reflectivity measurement microphone Mu, positioned behind the individual I. According to a For example, the ambient reflectivity microphones Mi0 and Mn are arranged in the same vertical plane, on either side of the longitudinal x-axis. By their positions, the first ambient reflectivity microphone Mi0 and the second ambient reflectivity microphone Mn are configured to optimally measure the degree of laterality of the reflected portion of the audio signal from the voice of individual I. In other words, the acquisition positions of the first ambient reflectivity microphone Mi0 and the second ambient reflectivity microphone Mn allow the acquisition of audio signals containing information characterizing the degree of ambient reflectivity in the sequence of at least one sound S emitted by individual I.

[0158] Ambient reflectivity means the quantification of the ability of a sound to reflect homogeneously and / or uniformly in space after being reflected off obstacles.

[0159] According to one example, the first ambient reflectivity measurement microphone Mi0 and the second ambient reflectivity measurement microphone Mn are located from the plane Pi (plane of the reference microphone Mo) by a distance of between 50 cm and 5 m. Advantageously, the first ambient reflectivity measurement microphone Mi0 and the second ambient reflectivity measurement microphone Mn are positioned at a height higher than the height of the reference microphone Mo, preferably higher by a distance of between 0 and 2 m, preferably higher by 50 cm.

[0160] In some embodiments, the acquisition system 15 includes a microphone Mi2 called a laryngeal pickup microphone, positioned near the throat of individual I. By its position, the laryngeal pickup microphone M12 is configured to optimally measure the vibrations of the larynx produced during the phonation of the voice of individual I. In other words, the acquisition position of the laryngeal pickup microphone Mi2 makes it possible to acquire an audio signal comprising information characterizing vocal vibrations, in particular the vibrations of the larynx, by providing precise data on the fundamental frequency of the voice in the sequence of at least one sound S emitted by individual I.

[0161] In certain embodiments, each microphone in the plurality of microphones of the acquisition system 15 has the same type, in other words, the same technical characteristics, such as bandwidth, response curve, sensitivity, and directivity. In this case, only the position (i.e., the location) of each microphone of the acquisition system 15 relative to the position of individual I distinguishes it from the other microphones of the acquisition system 15.

[0162] In operation, when positioned relative to an individual I standing and emitting a sequence of at least one sound S, the acquisition system 15 can then collect a plurality of input audio signals from each of the microphones of the acquisition system 15. Each input audio signal encodes the sequence of at least one sound S emitted by individual I. In other words, each audio signal is a representation of the sequence of at least one sound S. For example, if the audio signals are analog, each audio signal comprises an electrical voltage signal with a variable voltage level. In another example, if the audio signals are digital, each audio signal comprises a series of binary numbers.

[0163] Advantageously, since each microphone of the acquisition system 15 is positioned at a distinct location, the input audio signal from a microphone is specific to that location.

[0164] The acquisition system 15 described above can be used in a training phase to obtain audio signals constituting training data for a machine learning model generating a first learned function that will be used in the method 100 to determine a value of a voice descriptor of the voice of an individual I according to the invention described below. In this case, the acquisition system comprises at least two microphones positioned at two different locations relative to the position of an individual whose voice is recorded by the acquisition system 15.

[0165] The acquisition system 15 described above can be used in an operational phase to obtain audio signals constituting the input data for the learned function generated by the preceding machine learning model and used in the method 100 to determine a value of a voice descriptor of an individual's voice according to the invention described below. The acquisition system 15 comprises at least one microphone.

[0166] The method 100 for determining a value of a voice descriptor of an individual's voice according to the invention will now be described. Applications of the method 100 include, but are not limited to: establishing a vocal assessment of an individual for the purpose of vocal coaching, establishing a vocal assessment of an individual for the diagnosis of a vocal pathology or voice dysfunctions, and characterizing human voices for multimedia applications.

[0167] Establishing a vocal assessment according to the invention makes it possible to obtain a detailed representation of the characterization of an individual's voice. This characterization makes it possible to reduce diagnostic errors and / or to refine recommendations for preventing damage to the body caused by poor voice projection habits, etc.

[0168] The method of the invention makes it possible to improve the early detection of functional voice problems.

[0169] According to another aspect, voice characterization allows for better characterization of objective indicators that are precursors of pathological conditions, for example Parkinson's disease.

[0170] According to other applications, these descriptors can be used to refine lie detection when these descriptors are used in conjunction with audio signal processing, data characterizing an individual's breathing and data characterizing the individual's heartbeat.

[0171] Finally, the method of the invention makes it possible to classify characteristic descriptors of an individual's voice, which allow for the determination of a specific corrective action to be taken by the individual. For example, increasing lung air intake, taking better breaths, opening the mouth wider than a reference opening, or calibrating recording equipment according to the quantification of the measured and classified descriptors.

[0172] According to one embodiment, voice characterization allows for better corrections, for example, in the reconstruction of an individual's voice. Typically, during the recording of a phonogram, a disc, or a dubbing, corrective factors can be introduced to compensate for certain aspects of the voice based on the quantifications of the measured voice descriptors.

[0173] The method 100 can be implemented using a device 10 as illustrated in [Fig.2] and called a descriptor device, which represents a schematic diagram showing the components of an example of the device 10.

[0174] The device of descriptor 10 can be implemented as a single hardware device, for example in the form of a desktop personal computer (PC), a laptop computer, a personal digital assistant (PDA), a smartphone, a smartwatch, a server, a console, or can be implemented on separate interconnected hardware devices interconnected by one or more communication links, with wired and / or wireless segments. The device 20 can, for example, communicate with one or more cloud computing systems, one or more servers, or remote devices to implement the functions described in this description for the device concerned. The device 20 can also be implemented itself as a cloud computing system.

[0175] As shown in [Fig. 2], the device of descriptor 10 comprises a computer, this computer comprising a memory 11 for storing program instructions that can be loaded into a circuit 12 and adapted to cause a circuit to execute steps of the process 100 illustrated in [Fig. 3], described later, when the program information is executed by the circuit. The memory can also to store data and information useful for carrying out the steps of the present invention as described below.

[0176] Circuit 12 can be, for example: - a processor or processing unit adapted to interpret instructions in a computer language, the processor or processing unit being able to understand, be associated with, or be attached to a memory containing the instructions, or - the combination of a processor / processing unit and a memory, the processor or processing unit being adapted to interpret instructions in a computer language, the memory containing said instructions, or - an electronic circuit board in which the steps of the invention are described in silicon, or - a programmable electronic chip such as an FPGA chip (for "Field-Programmable Gate Array").

[0177] The memory 11 may include random access memory (RAM), cache memory, non-volatile memory, backup memory (e.g., programmable or flash memory), read-only memory (ROM), a hard disk drive (HDD), a solid-state drive (SSD), or any combination thereof. The ROM of the memory 11 may be configured to store, among other things, an operating system and / or one or more computer program codes for one or more software applications. The RAM of the memory 11 may be used by the circuit 12 for the temporary storage of data.

[0178] The computer may also include an input interface 13 for receiving input data and an output interface 14 for providing output data. Examples of input and output data will be provided later.

[0179] To facilitate interaction with the computer, a screen 15 and a keyboard 16 may be provided and connected to the computer circuit.

[0180] The descriptor device 10 is configured to implement the computer-implemented method 100 to determine a value D of a voice descriptor of an individual's voice I according to the invention.

[0181] The [Fig.3] is a flowchart representing an example of a set of steps that can be carried out to implement the process 100 to determine an intrinsic characteristic value D of a voice descriptor of an individual's voice I.

[0182] In this application, a voice descriptor is understood to mean a characteristic capable of describing an aspect of the sound of an individual's voice. The voice descriptor characterized by method 100 according to the invention can be chosen from a set of voice descriptors called reference voice descriptors, comprising: an attack of the sound, a pitch, a quantity of air over the voice, a sound termination, a harmonic mixture, a degree of laterality, a degree of projectivity, a degree of vibrato.

[0183] These reference vocal descriptors are classically known to human voice experts, who can evaluate them for a given individual by a physical examination in a given context, such as vocal coaching, or a clinical voice examination.

[0184] One advantage of the invention is that it allows for an objective characterization of these criteria through the implementation of a machine learning algorithm. Since the descriptor data is annotated—that is, labeled or categorized by domain experts—the classifier will provide a classification and quantification of the descriptors specific to the profession. Depending on the various implementations of the invention, it is possible to configure the probability threshold levels of the machine learning algorithm's predictions in order to personalize the classifier according to the specific needs of each professional.

[0185] In a first REC step, the descriptor device 10 receives and then stores in memory 11 at least one input audio signal S Ai, SA2, ... SAN.

[0186] In some embodiments, the at least one input audio signal SAb SA2, ... SAn is one or a plurality of signals from the acquisition system 15 previously described.

[0187] In other embodiments, the at least one input audio signal SAb SA2, ... SAn is one or a plurality of signals from any other acquisition system. These other signals encode a sequence of at least one sound that was emitted by individual I. In one example, the other signals each encode the sequence of at least one sound S, recorded from positions distinct relative to the individual's position when emitting the sequence of at least one sound S.

[0188] In certain embodiments, depending on the voice descriptor for which an intrinsic characteristic value is to be obtained with the method 100, the at least one input audio signal SAb SA2, ... SAN corresponds to one or a plurality of audio signals which have been acquired by a set of microphones configured to capture a set of audio signals including non-zero information relating to this voice descriptor.

[0189] The invention also relates to a method for first recording and calculating descriptors from a maximum number of microphones and then selecting the most representative descriptors for each descriptor. The selection of the most representative descriptors can be carried out by comparing the range of descriptor values ​​by varying the input signal according to an aspect of the voice, for example its power, harmonics, vibrato, etc.

[0190] Thus, according to one example, when the vocal descriptor is the attack of the sound, a preferred set of microphones configured to acquire the input audio signals S Ai, SA2, ... SAn includes: the reference microphone and the harmonic measurement microphones. Other microphones can also be selected to enrich the characterization of this descriptor.

[0191] According to one example, where the voice descriptor is the degree of projectivity, a preferred set of microphones configured to acquire the input audio signals S Ai, SA2, ... SAN includes, for example, projectivity, directionality, and laterality microphones. Other microphones can also be selected to enrich the characterization of this descriptor.

[0192] According to one example, where the voice descriptor is the degree of reflectivity, a preferred set of microphones configured to acquire the input audio signals S Ai, SA2, ... SAN includes reflectivity, ambient projectivity, or ambient reflectivity microphones. Other microphones may also be selected to enrich the characterization of this descriptor.

[0193] Advantageously, in certain embodiments, in a second REC2 reception step, the descriptor device 10 receives and then stores configuration information in memory 11. The configuration information includes, for each input audio signal among the at least one input audio signal (SAi, SA2, ... SAn), position information for an acquisition microphone associated with that input audio signal relative to a position of the individual I when it emitted the sequence of at least one sound S. More precisely, the position information associated with an input audio signal includes the position of the microphone with which the sequence of at least one sound S was acquired so as to produce the input audio signal.

[0194] After receiving the input audio signals SAb SA2, ... SAN and optionally configuration information, in a step called the first implementation step IMPI, the circuit 12 of the device 10 implements a first learned function generated from a first machine learning model previously trained from a first training domain comprising first training data. The first training data comprises at least two audio training signals encoding the same speech sequence called the training sequence. By speech sequence, we mean a sequence of at least one sound emitted by a human being, for example, an artist. The training of the first machine learning model from the first training domain will be described later in this description.

[0195] The first learned function receives as input at least one audio input signal SAi, SA2, ... SAN, so as to generate as output, for each of the input audio signals, a corresponding value of a descriptor characteristic of the individual's voice. In other words, a plurality of values ​​of the voice descriptor of individual I's voice is obtained at the end of the first implementation step IMPI, each corresponding to one of the at least one audio input signal SAi, SA2, ... SAN.

[0196] In one example, the values ​​of the voice descriptor are scores between 0 and 10. In another example, the values ​​of the voice descriptor are percentages between 0 and 100%. Advantageously, the values ​​of the voice descriptor characterize the intensity of that voice descriptor. Thus, a score of 8 out of 10 for the voice descriptor of the attack of the sound will be characteristic of a strong attack of the sound.

[0197] In some embodiments, the first machine learning model generating the first learned function is based on a convolutional neural network (CNN). Alternatively or in combination, other architectures can be used to build the first machine learning model. Thus, in some alternative embodiments or embodiments that can be combined with the preceding embodiments, the first machine learning model can be based on architectures using temporal correlations, such as recurrent neural networks (RNNs, LSTMs, U-Nets).

[0198] Advantageously, when the at least one input audio signal SAb SA2, ... SAN comprises at least two input audio signals corresponding to two different acquisition positions of the common sequence of at least one sound emitted S by the voice of individual I, the method 100 includes a step called the second implementation step IMP2, in which the circuit 12 of the device 10 implements a function called the global function. The global function receives as input the corresponding values ​​of the characteristic descriptor of the individual's voice for each of the input signals obtained during the first implementation step IMPI, so as to generate as output the intrinsic characteristic value D of the voice descriptor of the individual's voice.

[0199] In some embodiments, the global function is the maximum function. In other embodiments, the global function is the minimum function. In still other embodiments, the global function is an averaging function. In yet other embodiments, the global function is a weighted sum of the corresponding values ​​of the characteristic descriptor of each of the input audio signals SAb SA2, ... SAN.

[0200] When the at least one input audio signal comprises a single input signal, the intrinsic characteristic value D corresponds to the value obtained during the first implementation step IMP1. This is then the corresponding value of at least one first voice descriptor relating to the microphone that acquired said voice of said individual which was calculated by the first learned function.

[0201] According to the invention, a system 20 for determining an intrinsic characteristic value D of a voice descriptor of an individual's voice I, schematically represented in [Fig. 4], comprises the descriptor device 10 and the acquisition system 15 described above. In operation, the acquisition system 15 then implements a step called the ACQ acquisition step of a sequence of at least one sound emitted by individual I so as to obtain at least one input audio signal SAb SA2,... SAN. The acquisition system 15 and the descriptor device 10 are interconnected by one or more communication links, with wired and / or wireless segments. Thus, at least one audio input signal SAb SA2, ... SAN is then transmitted via the communication link(s) to the descriptor device 10. For example, the descriptor device 10 receives at least one audio signal SAb SA2, ... SAN via its input interface 13.Method 100 can be implemented after the ACQ acquisition step by the descriptor device 10. Different audio tracks generated by different sound recording and capture devices can be routed to the same acquisition equipment to use all the input audio signals.

[0202] In certain embodiments, the method 100 makes it possible to determine at least two characteristic values ​​of two different voice descriptors of an individual's voice, by implementing the first implementation step IMPI and the second implementation step IMP2 following the reception of at least one input audio signal SAb SA2, ... SAN. According to one embodiment, the two implementation steps are carried out simultaneously.

[0203] In certain embodiments, the method 100 makes it possible to determine a plurality of characteristic values, each of a voice descriptor of an individual's voice, by implementing the first implementation step IMPI and the second implementation step IMP2 following the reception of at least one input audio signal SAb SA2, ... SAN. According to one embodiment, the two implementation steps are carried out simultaneously.

[0204] Training the first machine learning model

[0205] In some embodiments, the first machine learning model is trained prior to the implementation of process 100.

[0206] It is recalled that training a machine learning model consists of determining a set of model parameters and / or teaching the model a corresponding learned function from training data, so that the model can then predict a label of new data received.

[0207] Advantageously, the training data are annotated (or labeled, or tagged). Advantageously, the training data are manually annotated by a panel of experts. By expert is meant a person reputed to have mastered the knowledge of a given field. For example, an expert might be a singing teacher, a sound engineer, an acoustician, or a doctor.

[0208] In some embodiments, the initial training data were collected and annotated during a research session involving a plurality of experts and one or more artists.

[0209] During the research session, the artists perform one or more vocal performances which are recorded by at least two microphones spatially arranged in an enclosed space, and which are observed by a plurality of experts. For example, the acquisition system 15 can be used to record the vocal performance(s).

[0210] Advantageously, each recording of the same vocal performance is associated, for example by a label, with information relating to the microphone that produced the recording. This information may be the position of the microphone relative to the position of the performer when they performed the vocal performance. The information may also be the type of microphone, for example: direct microphone, high-harmonic measurement microphone, low-harmonic measurement microphone, projectivity measurement microphone, lateral measurement microphone, reflectivity measurement microphone, ambient projectivity measurement microphone, ambient reflectivity measurement microphone, or laryngeal microphone.

[0211] For each recording of a given vocal performance, each expert assigns a corresponding value to a vocal descriptor of the voice of the artist performing the vocal performance. The vocal descriptor is a reference vocal descriptor such as: attack, pitch, air volume, ending, harmonic blend, degree of laterality, degree of projection, or degree of vibrato. Then, based on the values ​​assigned by each expert, a final value, called the characteristic value of the vocal descriptor, is assigned to the vocal performance for each recording.

[0212] At the end of the search session, a plurality of vocal performance recordings are obtained, in which each recording is associated with one or more vocal descriptor values ​​of the corresponding artist's voice.

[0213] Another aspect of the invention relates to a method 200 for characterizing at least one physiological descriptor of an individual's voice I.

[0214] In this application, a physiological descriptor of a voice means a characteristic describing an individual's voice from the point of view of the individual's physiology, that is to say, a characteristic related to one or more organs of the individual's body.

[0215] At least one physiological descriptor characterized by the method 200 according to the invention can be selected from a set of physiological descriptors referred to as reference physiological descriptors, including: vocal strain, muscular strain, and glottal supports. These reference physiological descriptors are commonly known to human voice experts, who can assess them for a given individual through a physical examination in a given context, such as vocal coaching, a clinical voice examination, or any other application utilizing certain descriptors in particular.

[0216] In some embodiments, the method 200 for characterizing at least one physiological descriptor of an individual's voice is implemented by a device 30 called pathology device 30 and similar to the descriptor device 10. The pathology device 30 is illustrated in [Fig. 5], which represents a schematic diagram showing the components of an example of the pathology device 30.

[0217] The Pathology Device 30 can be implemented as a single hardware device, for example in the form of a desktop personal computer (PC), a laptop computer, a personal digital assistant (PDA), a smartphone, a smartwatch, a server, a console, or it can be implemented on separate interconnected hardware devices linked by one or more communication links, with wired and / or wireless segments. The Pathology Device 30 can, for example, communicate with one or more cloud computing systems, one or more servers, or remote devices to implement the functions described in this description for the device concerned. The Pathology Device 30 can also be implemented itself as a cloud computing system.

[0218] As shown in [Fig. 5], the pathology device 30 comprises a computer, this computer comprising a memory 31 for storing program instructions that can be loaded into a circuit 32 and adapted to cause a circuit to execute steps of the process 200 illustrated in [Fig. 5], described below, when the program information is executed by the circuit 32. The memory 31 can also store data and information useful for the execution of the steps of the present invention as described below.

[0219] Circuit 32 can be, for example: - a processor or processing unit adapted to interpret instructions in a computer language, the processor or processing unit being able to understand, be associated with, or be attached to a memory containing the instructions, or - the combination of a processor / processing unit and a memory, the processor or processing unit being adapted to interpret instructions in a computer language, the memory containing said instructions, or - an electronic circuit board in which the steps of the invention are described in silicon, or - a programmable electronic chip such as an FPGA chip (for "Field-Programmable Gate Array").

[0220] The memory 31 may include random access memory (RAM), cache memory, non-volatile memory, backup memory (e.g., programmable or flash memory), read-only memory (ROM), a hard disk drive (HDD), a solid-state drive (SSD), or any combination thereof. The ROM of the memory 31 may be configured to store, among other things, an operating system and / or one or more computer program codes for one or more software applications. The RAM of the memory 31 may be used by the circuit 32 for the temporary storage of data.

[0221] The computer may also include an input interface 33 for receiving input data and an output interface 34 for providing output data. Examples of input and output data will be provided later.

[0222] To facilitate interaction with the computer, a screen 35 and a keyboard 36 may be provided and connected to the computer circuit.

[0223] The pathology device 30 is configured to implement the computer-implemented method 200 to characterize at least one physiological descriptor of an individual's voice according to the invention.

[0224] In some embodiments, the pathology device 30 is the descriptor device 10.

[0225] The [Fig.6] is a flowchart representing an example of a set of steps that can be carried out to implement the process 200 to characterize at least one physiological descriptor of an individual's voice and will be described below.

[0226] In a step called the REC3 descriptor reception step, the pathology device 30 receives and then stores in memory 31 at least one characteristic value of a voice descriptor of the individual's voice. This descriptor can be calculated by the global function or be directly derived from the execution of the first learned function, depending on the number of microphones used.

[0227] In some embodiments, the process 200 uses characteristic values ​​of voice descriptors obtained by the process 100 to determine a characteristic value of a voice descriptor of a voice of a previously described individual.

[0228] In other embodiments, at least one characteristic value of a voice descriptor of the individual's voice is obtained by another process or has another origin.

[0229] In a step called the third implementation step IMP3, the circuit 32 of the pathology device 30 implements a second learned function generated from a second machine learning model trained from a second training domain comprising second training data. The second learned function receives as input at least one characteristic value of the voice descriptor, so as to generate as output at least one score characterizing at least one physiological descriptor of the individual's voice.

[0230] In some embodiments, the second machine learning model is the first machine learning model that generated the first learned function.

[0231] In other embodiments, the second machine learning model is another model, different from the first machine learning model.

[0232] In some embodiments, the second machine learning model generating the second learned function is based on a convolutional neural network (CNN). Alternatively or in combination, other architectures can be used to build the second machine learning model. Thus, in some alternative embodiments or embodiments that can be combined with the preceding embodiments, the second machine learning model can be based on architectures using temporal correlations, such as recurrent neural networks (RNNs, LSTMs, U-Nets).

[0233] According to one embodiment, associations and / or combinations of characteristic values ​​of the analytical descriptors make it possible to associate said analytical descriptors with a physiological descriptor, for example by associating the corresponding score of the physiological descriptor with different conditions on the scores of several analytical descriptors.

[0234] In order to associate these values, a labeling of the data characterizing the pathological descriptors of the individual can be carried out so as to train a machine learning model.

[0235] According to an example, if the score of the analytical descriptor describing the amount of air on the voice is greater than 8 / 10 and the score of the analytical descriptor describing the pitch is greater than 6 / 10, then a given physiological descriptor can be associated with these two conditions.

[0236] In certain embodiments, the method 200 includes a step called the fourth implementation step IMP4, in which the circuit 32 of the pathology device 30 implements a function called the comparison function. The comparison function is configured to compare at least one characteristic value of the voice descriptor received during the REC3 descriptor reception step, so as to generate at least one output score characterizing at least one physiological descriptor of the individual's voice.

[0237] For example, when the voice descriptor is air over voice, the characteristic value can be compared to a threshold value of 80%. If the characteristic value is greater than this threshold value, the score obtained during the fourth stage of IMP4 implementation is 100% voice forcing.

[0238] In certain embodiments, the method 200 includes an additional step. In the additional step, the circuit 32 of the pathology device 20 implements a third learned function generated from a third machine learning model trained from a third training domain comprising third training data. The third learned function receives as input at least one score characterizing at least one physiological descriptor of the individual's voice and generated by the second learned function during the third implementation step IMP3 and generates as output at least one percentage value characterizing a contribution of at least one stage of the individual's vocal tract to the physiological descriptor.

[0239] At least one stage of the individual's vocal apparatus may be one of: the individual's lung, the individual's larynx, the individual's pharynx, an individual's articulator.

[0240] According to one embodiment, associations and / or combinations of characteristic values ​​of physiological descriptors make it possible to associate said physiological descriptors with the involvement of a level of the phonatory apparatus.

[0241] Training the second machine learning model and the third machine learning model

[0242] In some embodiments, the second machine learning model and the third machine learning model are trained prior to the implementation of method 200.

[0243] It is recalled that training a machine learning model consists of determining a set of model parameters and / or teaching the model a corresponding learned function from training data, so that the model can then predict a label of new data received.

[0244] Advantageously, the training data are annotated (or labeled, or tagged). Advantageously, the training data are annotated manually by a panel of experts. By expert, we mean a person renowned for mastering the knowledge of a given field. For example, an expert could be a singing teacher, a sound engineer, an acoustician, or a doctor.

[0245] The second training data allows for the training of a second machine learning model capable of identifying and quantifying a physiological descriptor of the voice. This training data comprises associations of voice audio files with at least one label defining at least one physiological descriptor. A panel of experts can label certain voices with a quantified physiological descriptor. This labeling can be repeated for different physiological descriptors and different individuals so that the training data can provide sufficient data for all physiological descriptors. The training data includes, for example, heterogeneous voice samples, i.e., from individuals of different ages, different genders, and / or with characteristics that allow for the characterization of a voice sample.

[0246] The physiological descriptor of an individual's voice may relate, for example and in a non-limiting manner, to vocal strain, muscular strain, glottal supports, etc.

[0247] The third training data allows for the training of a third machine learning model capable of identifying and quantifying at least one percentage value characterizing a contribution from at least one level of the vocal tract of said individual to said physiological descriptor. This training data comprises associations of voice audio files with at least one label defining at least one percentage value characterizing a contribution from at least one level of the vocal tract of said individual to said physiological descriptor. A panel of experts can label certain voices of individuals with a percentage value characterizing a contribution from at least one level of the vocal tract of said individual to said physiological descriptor.This labeling can be repeated for different physiological descriptors and different individuals so that the training data can provide sufficient data for all physiological descriptors. The training data includes, for example, heterogeneous voice samples, i.e., from individuals of different ages, different genders and / or with characteristics that allow for the qualification of a voice sample, as well as antecedents, for example.

[0248] The vocal tract of said individual may refer to a lung of said individual, to the larynx of said individual, to the pharynx of said individual, or to an articulator of said individual. According to one embodiment, the invention makes it possible to quantify the contribution of each articulator in the production of the voice, of which: the vocal cords; the tongue, including the apex, the dorsum, the radix; the lips; the teeth; the alveoli; the hard palate; the soft palate; the uvula; the cheeks; the lower jaw also called the mandible; the nasal cavities.

[0249] Another aspect of the invention relates to a method for predicting a pathology of an individual's voice: - the steps of process 200 previously described to characterize at least one physiological descriptor of the individual's voice; - the additional step of implementing the third learned function generated from the third machine learning model trained from the third training domain, said third learned function receiving as input at least one score characterizing at least one physiological descriptor of the individual's voice, so as to generate as output at least one percentage value characterizing a contribution of at least one stage of the individual's phonatory apparatus to the physiological descriptor; - a prediction stage, of the pathology of the individual's voice based on the contribution of at least one level of the individual's phonatory apparatus.

[0250] The training domain refers to a set of training data. Thus, it is possible to define domains according to age or gender, or for example domains for certain diseases or pathologies, or even according to certain individual profiles, or taking into account certain antecedents.

[0251] By way of example, a vocal strain descriptor may have a contribution distributed across different levels of an individual's vocal tract. For instance, a vocal strain of 1 / 10 is considered acceptable and may correspond to an engagement of less than 70% of the total laryngeal involvement. The remaining 30% results from the involvement of lower or higher levels, for example, the lungs, pharynx, or articulators. In this case, the vocal cords may have an involvement of less than 10% in the contribution of all the components of the vocal tract, which may be acceptable.

[0252] For a vocal strain of 9 / 10, the larynx can be involved up to 90%, and in particular the vocal cords can be involved between 80% and 90% in contributing to the The value of the physiological descriptor represents a guttural voice that is harmful to the vocal cords. Thus, the invention allows for the quantification of the proportion of each physiological descriptor using a trained model. This makes it possible to implement corrective actions to address this strain, such as appropriate vocal rehabilitation exercises, medication (under medical supervision) to reduce vocal cord inflammation, recommendations for speech therapy, vocal exercises, physical exercises, rest, articulator relaxation exercises, and so on.

Claims

1.

2. Demands A computer-implemented method (100) for determining an intrinsic characteristic value (D) of a first voice descriptor of an individual's voice (I), said first voice descriptor belonging to a set of reference voice descriptors, the method (100) comprising: • reception (REC) of at least one input audio signal (SAb SA2,.. SAN), each input audio signal (SAb SA2,.. SAN) encoding a common sequence of at least one sound emitted (S) by said voice of said individual, said at least one input audio signal (SAb SA2,.. SAN) being acquired by at least one microphone; • implementation (IMPI) of a first learned function generated from a first machine learning model previously trained from a first training domain comprising at least two training audio signals encoding the same sequence called the training sequence and acquired by two different microphones respectively, said first learned function receiving as input said at least one input audio signal (SAb SA2,.. SAN), so as to generate as output, for each of the input audio signals, a corresponding value of at least one first voice descriptor relating to the microphone which acquired said voice of said individual (I), • determination of an intrinsic characteristic value (D) of the first voice descriptor of an individual's voice (I) from the corresponding value(s) of the first voice descriptor relative to each microphone generated during the implementation step (IMPI). Method (100) according to claim 1 characterized in that when a single input audio signal is acquired by means of a single microphone, the corresponding value of the first generated descriptor corresponds to the intrinsic characteristic value of the first descriptor.

3. Method (100) according to claim 1, wherein the at least one input audio signal (SAb SA2, ... SAN) comprises at least two input audio signals corresponding to recordings acquired by two microphones arranged in two different acquisition positions, said recordings corresponding to audio signals from the common sequence of at least one sound emitted (S) by said voice of said individual (I), the method (100) further comprising: • implementation (IMP2) of a function called global function receiving as input each corresponding value of the first voice descriptor of said voice of said individual acquired by each microphone, so as to generate as output said intrinsic characteristic value (D) of the first voice descriptor of said voice of said individual (I).

4. Method (100) according to the preceding claim, comprising, prior to the implementation step of the first learned function: • receiving (REC2) configuration information comprising, for each input audio signal among the at least one input audio signal (SAb SA2,... SAN), position information from an acquisition microphone associated with said input audio signal relating to a position of the individual when he emitted the sequence of at least one sound (S), said first learned function also receiving as input all the position information of all the input audio signals.

5. Method (100) according to any one of claims 3 or 4, wherein the overall function is one of the following: maximum function, minimum function, averaging function, weighted sum of corresponding values ​​of the voice descriptor of each of the input audio signals.

6. A method (100) according to any one of the preceding claims, wherein the method (100) is a method for determining an intrinsic characteristic value of at least two different vocal descriptors of the individual's voice.

7. A method (100) according to any one of the preceding claims, wherein the reference voice descriptors are normalized on the same scale of values.

8. A method (100) according to any one of the preceding claims, wherein the first training domain comprises a set of training audio signals, each training audio signal being pre-labeled with at least one label, the at least one label comprising a value of at least one reference voice descriptor.

9. A method (100) according to any one of the preceding claims, wherein the first function comprises a neural network, such as a convolutional neural network.

10. A method (100) according to claim 9, wherein a training step comprises: • A plurality of recordings of a sequence of at least one sound by a plurality of microphones (Mi), each microphone (Mi) being associated with a recording configuration and being arranged in a predefined position relative to a reference recording position; • An estimation of the corresponding value of the voice descriptor of said voice of said individual (I) of each microphone; • An estimation of the intrinsic characteristic value of a voice descriptor of said voice; • The association of said estimated value of the intrinsic characteristic value of a voice descriptor of said voice with each corresponding value of the voice descriptor of said voice of said individual (I) of each microphone;• Learning the neural network by modifying the coefficients of said network based on associations of said values.

11. A method (100) according to any one of the preceding claims characterized in that at least one vocal descriptor is selected from the list of descriptors: {a quantification of the attack of the sound, a quantification of the pitch, a quantification of air on the voice, a quantification of the termination of the sound,

12.

13.

14.

15. a quantification of projectivity, a quantification of harmonic mixing, a quantification of a degree of laterality, a quantification of a degree of vibrato}. Method (100) according to any one of the preceding claims, comprising, before the reception step (REC), an acquisition step (ACQ), by an acquisition system (15), of said at least one input audio signal (SAb SA2, ... SAN). Programmable device configured to implement the method according to any one of claims 1 to 12. System (20) for determining an intrinsic characteristic value of a voice descriptor of an individual's voice and configured to implement the method according to claim 12, comprising: • an acquisition system (15) configured to implement the acquisition step (ACQ), comprising at least one microphone, the at least one microphone being configured to simultaneously acquire said common sequence of at least one sound emitted (D) by said voice of said individual (I), so as to generate at least one input audio signal (SAb SA2, ... SAn); • the programmable device according to claim 13. System (20) according to claim 14, wherein at least one microphone of the acquisition system (15) comprises at least one of: • a reference microphone (Mo) positioned, in operation, facing and at the height of the mouth of the individual (I) when the latter emits the sequence of at least one sound (S), at a distance between 1 and 40 cm from the mouth of the individual (I), the orientation of the body of the individual from his back towards his torso defining a direction of orientation and a line of orientation; • a first harmonic measurement microphone (MJ) and a second harmonic measurement microphone (M2), positioned in front of the individual along the same vertical line containing the reference microphone (M0) and located at a distance of between 1 cm and 5 m from a vertical plane containing the individual (I), the first harmonic measurement microphone (MJ) being located above the height of the reference microphone (Mo) and at a distance from it of between 5 cm and 60 cm, and the second harmonic measurement microphone (M2) being located below the height of the reference microphone (Mo) and at a distance from it of between 5 cm and 60 cm; a first projectivity measurement microphone (M3) and a second projectivity measurement microphone (M4), positioned in a vertical plane in front of the individual, the vertical plane being distant from the vertical plane containing the individual by a distance of between 5 cm and 10 m, the height of the first projectivity measurement microphone (M3) being distant from the height of the individual's mouth by a distance of between 10 cm and 5 m, the height of the second projectivity measurement microphone (M4) being distant from the height of the individual's mouth by a distance of between 10 cm and 5 m; a first laterality measurement microphone (M5) and a second laterality measurement microphone (M6), positioned in a vertical plane containing the reference microphone (Mo) and perpendicular to the orientation direction, and substantially arranged at the height of the individual's mouth, on either side of the orientation line; a reflectivity measurement microphone (M7) positioned behind the individual along the orientation line and at a distance from the vertical plane containing the individual of between 30 cm and 10 m; a first ambient projectivity measurement microphone (M8) and a second ambient projectivity measurement microphone (M9), positioned in a vertical plane in front of the individual, the vertical plane being at a distance of between 5 cm and 10 m from the vertical plane containing the individual, the height of the first ambient projectivity measurement microphone (M8) being at a distance of between 1 cm and 10 m from the height of the individual's mouth, the height

16. the second ambient projectivity measurement microphone (M9) being at a distance from the height of the individual's mouth of a distance between 1 cm and 10 m; • a first ambient reflectivity measurement microphone (Mi0) and a second ambient reflectivity measurement microphone (Mu), positioned in a vertical plane behind the individual, the vertical plane being distant from the vertical plane containing the individual by a distance of between 5 cm and 10 m, the height of the first ambient projectivity measurement microphone (M8) being distant from the height of the individual's mouth by a distance of between 1 cm and 5 m, the height of the second ambient projectivity measurement microphone (M9) being distant from the height of the individual's mouth by a distance of between 1 cm and 5 m; • a microphone (Mi2) called a laryngeal pickup microphone, positioned near the throat of individual I. System (20) according to any one of claims 14 to 15, wherein it is configured for driving the first function: • at least one reference microphone when the voice descriptor is a quantification of the sound attack; • at least one harmonic microphone and possibly also the reference microphone when the vocal descriptor is a quantification of the pitch; • at least one reference microphone and one projection microphone when the voice descriptor is a quantity of air over the voice; • at least one reference microphone and one projection and / or harmonics microphone when the speech descriptor is the sound termination, • at least two harmonic microphones or, for example, a reference microphone and a microphone of harmonics when the vocal descriptor is a quantification of the harmonic mixture, at least one microphone of directionality or laterality and possibly also the reference microphone when the vocal descriptor is a quantification of a degree of laterality, at least one projectivity microphone and possibly also the reference microphone when the speech descriptor is a quantification of a degree of projectivity, at least one laryngeal microphone, called a "stethomicrophone" or laryngophone, when the vocal descriptor is a quantification of a degree of vibrato.

Citation Information

Patent Citations

  • Acoustic Based Speech Analysis Using Deep Learning Models

    US20210118426A1