Method for determining a value of a voice descriptor of a voice of an individual

A machine learning-based method addresses the variability in voice characterization by determining intrinsic descriptors from audio signals, enhancing vocal health and performance through standardized and reliable assessments.

WO2025248062A1PCT designated stage Publication Date: 2025-12-04INDICE DATA RECORDING
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/064941
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-30
Filing Date
2025-05-28
Publication Date
2025-12-04

AI Technical Summary

Technical Problem

Existing voice characterization methods are subjective, poorly standardized, and difficult to reproduce, leading to unreliable and non-interoperable assessments due to variability in equipment and recording conditions, which hinders effective vocal health protection and performance enhancement.

Method used

A computer-implemented method using machine learning models to determine intrinsic voice descriptors from audio signals, leveraging a predefined microphone configuration to accurately characterize voice characteristics independently of recording equipment and arrangement, allowing for objective assessment without expert intervention.

Benefits of technology

Enables accurate and reproducible characterization of voice descriptors, facilitating improved vocal health protection and performance through efficient and standardized vocal assessments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025064941_04122025_PF_FP_ABST
    Figure EP2025064941_04122025_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a computer-implemented method (100) for determining an intrinsic characteristic value of a voice descriptor of a voice of an individual, said voice descriptor belonging to a set of reference voice descriptors, the method (100) comprising: • - receiving (REC) at least one input audio signal encoding a common sequence of at least one sound emitted by the voice of said individual; • - implementing (IMP1) a first learned function generated from a first machine learning model previously trained using a first training domain.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] METHOD FOR DETERMINING A VALUE OF A VOICE DESCRIPTOR OF AN INDIVIDUAL'S VOICE

[0002] Scope of the invention

[0003] The invention relates to the field of computer-assisted voice analysis. More specifically, the invention relates to a method and / or a device for determining a value of a vocal descriptor of an individual's voice and / or for characterizing at least one physiological descriptor of the individual's voice.

[0004] State of the art

[0005] In many fields, a thorough analysis and rich understanding of the human voice are invaluable. For example, in the vocal arts, performers require vocal preparation to optimize their performance or ensure it while preserving their health, such as by combating or preventing vocal strain and, more generally, protecting their vocal instrument. In the healthcare field, a vocal assessment conducted by a healthcare professional can identify damaged areas of an individual's vocal tract and diagnose vocal pathologies—in other words, pathologies related to the vocal tract. Such vocal assessments are currently performed by professionals or groups of professionals with expertise in human voice analysis.Characterizing the human voice is, in general, a non-trivial problem and depends in particular on the person whose voice we want to characterize, the equipment recording the voice, the recording configuration, the recording location, etc.

[0006] However, these characterizations are generally subjective, poorly standardized, and / or difficult to reproduce and / or compare between different professionals, due in particular to the variability of the equipment used (types of sensors, spatial configurations, ambient acoustics, etc.), which can compromise the reliability, traceability, and / or interoperability of voice assessments. Finally, various functional elements constituting an aspect of the vocal tract can influence the voice, including the vocal cords, lungs, thorax, trachea, larynx, soft palate, hard palate, nasal cavity, pharynx, teeth, lips, tongue, and oral cavity.

[0007] Characterizing an individual's voice and identifying the elements responsible for a pathology on the phonatory apparatus would allow the application of the most appropriate treatments or would allow the choice of equipment configurations and arrangement of said equipment, for example, during audio recordings to be made.

[0008] Characterizing the voices would allow us to:

[0009] - to achieve better protection of the vocal cords in order to prevent damage due to intensive use without preparation.

[0010] - improve performance through appropriate vocal warm-ups to enhance the flexibility and capacity of the vocal cords;

[0011] - reduce tension which can affect voice quality and cause vocal fatigue.

[0012] Today, there is a need to be able to perform such vocal assessments more efficiently and quickly, by more accurately characterizing an individual's voice. This invention improves the situation by making it possible to objectively assess the characteristics of the voice.

[0013] Summary of the invention

[0014] A first aspect of the invention relates to a computer-implemented method for determining an intrinsic characteristic value of a first voice descriptor of an individual's voice, said voice descriptor belonging to a set of reference voice descriptors, the method comprising:

[0015] - reception of at least one input audio signal, each input audio signal encoding a common sequence of at least one sound emitted by the individual's voice;

[0016] - implementation of a first learned function generated from a first machine learning model trained from a first training domain comprising at least two training audio signals encoding the same sequence called the training sequence and acquired by two different microphones respectively, the first learned function receiving as input at least one input audio signal, so as to generate as output, for each of the input audio signals, a corresponding value of at least one first voice descriptor relative to the microphone that acquired the individual's voice,

[0017] - determination of an intrinsic characteristic value of the first voice descriptor of an individual's voice from the corresponding value(s) of the first voice descriptor relative to each microphone generated during the implementation step.

[0018] Thus, the invention makes it possible to obtain information on one or more vocal descriptors of an individual's voice, from audio signals constituting a representation of that individual's voice, without the intervention of a voice expert such as a vocal coach, a voice practitioner, or a voice osteopath. Such information can be advantageously used in various applications, such as for the evaluation and coaching of artists, therapeutic applications for voice pathologies, or multimedia applications where voice samples are used.

[0019] One advantage of the invention is that it leverages a hardware configuration and microphone arrangement to qualify and characterize voice descriptors. This allows for the accurate reproduction of the descriptor value from a single microphone or a few microphones when the voice is recorded using a single reference microphone. This is made possible by a machine learning algorithm trained on a predefined microphone configuration, which associates the actual descriptor value with the specific reference microphone used. The benefit lies in using only one or a few microphones to accurately characterize the voice.

[0020] One advantage is to determine an intrinsic value of a voice descriptor independently of the material used, its arrangement, for example its position or its type.

[0021] The characteristic value of such a descriptor, made independent of the recording microphone, is called an intrinsic characteristic value. In one embodiment, when a single input audio signal is acquired by means of a single microphone, the corresponding value of the first generated descriptor corresponds to the intrinsic characteristic value of the first descriptor.

[0022] The equipment used to implement the process can be an electronic terminal such as a personal computer, a smartphone, a tablet, or any other device containing a computer. In one example, the process is executed on a remote server.

[0023] According to one embodiment and for all aspects of the invention, at least one vocal descriptor is selected from the list of descriptors: {a quantification of the attack of the sound, a quantification of the pitch, a quantification of air on the voice, a quantification of the termination of the sound, a quantification of projectivity, a quantification of harmonic mixing, a quantification of a degree of laterality, a quantification of a degree of vibrato}.

[0024] In certain embodiments, the at least one input audio signal comprises at least two input audio signals corresponding to recordings acquired by two microphones arranged in two different acquisition positions, said recordings corresponding to audio signals from the common sequence of at least one sound emitted (S) by the individual's voice, the method further comprising:

[0025] ■ implementation of a function called global function receiving as input each corresponding value of the voice descriptor of said voice of said individual acquired by each microphone, so as to generate as output said intrinsic characteristic value of the voice descriptor of the individual's voice.

[0026] Thus, the global function makes it possible to determine the intrinsic characteristic value of the voice descriptor on the basis of at least two corresponding values ​​of input audio signals and therefore, with increased accuracy with the number of input audio signals.

[0027] In some embodiments, the process includes, prior to the implementation step of the first learned function:

[0028] - receiving configuration information including, for each input audio signal among the at least one input audio signal, position information from an acquisition microphone associated with said input audio signal relating to a position of the individual when he emitted the sequence of at least one sound, the first learned function also receiving as input all the position information from all the input audio signals.

[0029] Thus, the corresponding value(s) generated by the first learned function are obtained using information related to the context of the input audio signal acquisition. More specifically, the corresponding value generated for a given input audio signal depends on the acquisition parameters of that input audio signal, such as the position of the microphone that acquired the input audio signal.

[0030] In some embodiments, the global function is one of the following: the maximum function, the minimum function, an averaging function, a weighted sum of the corresponding values ​​of the voice descriptor of each of the input audio signals.

[0031] Thus, the global function allows the contribution of each microphone, based on the corresponding values ​​obtained during the implementation step of the first learned function, to be adjusted to the intrinsic characteristic value. In this way, when configuration information has also been received, the global function allows the contribution of different microphones that acquired the different input audio signals to be adjusted to the characteristic value of the voice descriptor.

[0032] In some embodiments, the method is a method for determining an intrinsic characteristic value of at least two different vocal descriptors of the individual's voice.

[0033] In some embodiments, the reference voice descriptors are normalized on the same scale of values.

[0034] For example, the intrinsic or non-intrinsic characteristic value of a voice descriptor is a score out of a predetermined number of points. In another example, the intrinsic or non-intrinsic characteristic value of a voice descriptor is a percentage.

[0035] Thus, thanks to the normalization of vocal descriptors, it is possible to compare the characteristic values ​​of different vocal descriptors. In some embodiments, the reference vocal descriptors include: the attack of the sound, the pitch, the amount of air over the voice, the ending of the sound, the mixture of harmonics, the degree of laterality, the degree of projectivity, and the degree of vibrato. Furthermore, the reference vocal descriptors may include an indication of the location or proportion of the sound produced in the larynx or pharynx, for example, at the level of the hypopharynx, oropharynx, or nasopharynx.

[0036] In some embodiments, the first training domain comprises a set of training audio signals, each training audio signal being pre-labeled with at least one label, the at least one label comprising a value of at least one voice reference descriptor.

[0037] In one example, each training audio signal was previously labeled with at least one label by a panel of experts.

[0038] Thus, by the nature of the training of the first machine learning model, the result provided by the combination of the first learned function and the global function makes it possible to reproduce values ​​of voice descriptors as close as possible to the estimation of these descriptors by experts.

[0039] According to one embodiment, the process includes training the model comprising:

[0040] ■ A plurality of recordings of a sequence of at least one sound by a plurality of microphones, each microphone being associated with a recording configuration and being arranged in a predefined position relative to a reference recording position;

[0041] ■ An estimate of the corresponding value of the voice descriptor of said voice of said individual from each microphone;

[0042] ■ An estimate of the intrinsic characteristic value of a vocal descriptor of said voice;

[0043] ■ The association of said estimated value of the intrinsic characteristic value of a voice descriptor of said voice with each corresponding value of the voice descriptor of said voice of said individual of each microphone; ■ Learning of the neural network by modifying the coefficients of said network from the associations of said values.

[0044] One advantage is having a model trained with numerous microphone configurations. This offers two key benefits. First, it provides a model capable of performing well with an arbitrary recording setup, such as one with a given number of microphones, for example, a studio with seven microphones. Thus, all tracks can be used by the first learned function to produce corresponding values ​​for each microphone. The global function then generates an intrinsic characteristic value. The second advantage is that even with a single recording of a track, for example, using a reference microphone, the first learned function can generate a corresponding descriptor value that is intrinsic to the voice and relatively independent of the recording device.

[0045] In some embodiments, the at least one input audio signal consists of a single input audio signal.

[0046] Thus, advantageously, the invention makes it possible to determine one or more characteristic values ​​of vocal descriptors of an individual's voice from a single recording of it.

[0047] In some embodiments, at least one input audio signal is / are acquired by means of several tracks, generally between 1 and 13 tracks. Each track allows the acquisition of one input audio signal. Therefore, there can be as many input audio signals as there are tracks provided for this purpose. According to one embodiment, between 3 and 4 tracks are configured to acquire input audio signals from 3 or 4 microphones. The invention relates to the case where a single microphone can record one track, that is to say, a single input audio signal.

[0048] In some embodiments, at least one input audio signal consists of thirteen input audio signals.

[0049] In some embodiments, the method includes, prior to the reception step, an acquisition step, by an acquisition system, of at least one input audio signal. A second aspect of the invention relates to a programmable device configured to implement the method described above.

[0050] A third aspect of the invention relates to a system for determining an intrinsic characteristic value of a voice descriptor of an individual's voice and configured to implement the method described above, comprising:

[0051] - an acquisition system configured to implement the acquisition step, comprising at least one microphone, the at least one microphone being configured to simultaneously acquire the common sequence of at least one sound emitted by said voice of said individual, so as to generate at least one input audio signal;

[0052] - the programmable device described previously.

[0053] In some embodiments, each microphone among the at least one microphone is identical to the others.

[0054] In some embodiments, at least one microphone of the acquisition system includes at least one of the following:

[0055] - A reference microphone is positioned, when in operation, facing and at the height of the individual's mouth when they emit a sequence of at least one sound, at a distance of between 1 cm and 40 cm from the individual's mouth. The individual's body orientation from their back towards their torso defines a direction of orientation and a line of orientation. Such an arrangement corresponds to a typical microphone setup in a studio, concert hall, or auditorium. A distance of between 5 cm and 30 cm between the microphone and the mouth is optimal for recording a track using the reference microphone. The reference microphone is preferably oriented towards the individual's mouth.

[0056] - A first harmonic measurement microphone and a second harmonic measurement microphone are positioned in front of the individual along the same vertical line containing the reference microphone and located at a distance of between 1 cm and 5 m (meters) from a vertical plane containing the individual, or even between 1 cm and 2.15 m. The first harmonic measurement microphone is located above the reference microphone. The vertical offset between the first harmonic measurement microphone and the reference microphone allows for the detection of variations between the amplitude of one or more harmonic frequencies and the fundamental frequency, depending on the position of the mouth. The first harmonic measurement microphone is located above the reference microphone, for example, at a height greater than the height of the reference microphone.Increasing the height of the first harmonic measurement microphone relative to the reference microphone allows for the capture of sound transformed by body parts above the mouth, such as the nasal cavity and / or skull, and / or acoustic effects from environmental elements above the mouth, such as the ceiling. For example, this could involve increasing the height of the reference microphone from 5 cm to 60 cm. This is also referred to as an "overhead" microphone. Optionally, the first harmonic measurement microphone can be directed towards the mouth.

[0057] The second harmonic measurement microphone is positioned below the height of the reference microphone, for example, at a height less than the reference microphone's height (e.g., 5 cm to 60 cm). This microphone is advantageously placed at a similar distance to the overhead microphone, but angled upwards towards the mouth. The lower height of the second harmonic measurement microphone relative to the reference microphone allows it to capture sound transformed by parts of the individual's body located below the mouth, such as the torso, and / or acoustic effects from parts of the environment located below the mouth, such as the floor.

[0058] The distance between the first and second harmonic microphones is preferably between 5 cm and 1.4 m (for example, 1 m). These distances allow the harmonic measurement microphones to be spaced apart in a way that is favorable for detecting the harmonic mixture while respecting the typical morphology of a human being. Optionally, the separation can be between 30 cm and 1.4 m. From a distance of 30 cm, the signals captured by the two harmonic measurement microphones can enable a directional characterization of the spectral components of the voice.

[0059] - A first projectivity or directivity measurement microphone and a second projectivity or directivity measurement microphone are positioned in a vertical plane in front of the individual. Advantageously, this vertical plane is located at a distance from the vertical plane containing the individual, ranging from a few centimeters to several meters, depending on the room configuration in which the individual's voice is recorded. The height of the first projectivity or directivity measurement microphone is a distance from the height of the individual's mouth between a minimum and a maximum limit, determined by the orientation line. For example, the minimum limit could be 10 cm, 30 cm, or 2 m. Alternatively, the maximum limit could be 30 cm, 2 m, or 5 m.The projectivity measurement microphone(s) can be positioned further from the mouth than the reference microphone. This distance, when compared to the reference microphone, allows for the determination of sound evolution as a function of distance from the mouth. A distance between 10 cm and 30 cm allows for measuring sound evolution over relatively short distances. A distance between 30 cm and 2 m allows for measuring voice diffusion and / or energy loss. A distance between 2 m and 5 m allows for measuring the voice's ability to maintain its intelligibility and clarity. The height of the second projectivity or directivity measurement microphone is positioned between 10 cm and 5 m from the height of the individual's mouth, and preferably between 30 cm and 2 m, depending on the line of orientation, to avoid excessive or muddy reverberation.

[0060] - A first and second laterality measurement microphones are positioned approximately in a vertical plane containing the reference microphone and perpendicular to the orientation direction, at the height of the individual's mouth, on either side of the orientation line. By laterally moving the laterality measurement microphone(s) away from the orientation line, it may be possible, by comparison with the reference microphone, to detect a change in the sound in a lateral direction relative to the individual's mouth. "Approximately" means that a shift of a few centimeters can be considered as being in the same plane. In another embodiment, the laterality measurement can be performed in a plane parallel to the plane containing the reference microphone and located in the portion of space in front of the individual's face.The arrangement of two laterality measurement microphones positioned on either side of the orientation line can detect any asymmetry in vocal projection. Such asymmetry can result, for example, from a postural and / or articulatory disorder.

[0061] - A reflectivity measurement microphone positioned behind the individual along the orientation line. Such a microphone can be positioned at different distances from the vertical plane containing the individual, depending on the dimensions of the room in which the recording is made. For example, it can be located at a distance of between 30 cm and 10 m.

[0062] - A first ambient projectivity microphone and a second ambient projectivity microphone are positioned in a vertical plane in front of the individual. This vertical plane is located between 5 cm and 10 m from the vertical plane containing the individual. The height of the first ambient projectivity microphone is between 1 m and 10 m from the height of the individual's mouth, and the height of the second ambient projectivity microphone is between 1 cm and 10 m from the height of the individual's mouth. The ambient projectivity or ambient directionality microphones are preferably placed at a height of approximately 2 to 3 meters above the floor. This height allows for a good balance between direct and indirect reflections from the room, avoiding the capture of excessive reflections from the floor or ceiling.If the room is particularly tall with a high ceiling, the microphones can be placed even higher to capture more of the space's natural reverberation.

[0063] - A first ambient reflectivity measurement microphone and a second ambient reflectivity measurement microphone are positioned in a vertical plane behind the individual. The vertical plane is located between 5 cm and 10 m from the vertical plane containing the individual, depending on the dimensions of the recording room. The height of the first ambient reflectivity measurement microphone M8 is between 1 cm and 5 m from the height of the individual's mouth, depending on the dimensions of the recording room. The height of the second ambient reflectivity measurement microphone is also between 1 cm and 5 m from the height of the individual's mouth, depending on the dimensions of the recording room.

[0064] - a microphone called a laryngeal microphone, for example a stethoscope microphone or a laryngophone, positioned near the individual's throat. Such a microphone allows the direct vibrations of the skin on the throat, where the larynx is located, to be captured without picking up ambient sounds from the mouth or other sources.

[0065] In some embodiments:

[0066] - The vocal descriptor is the attack of the sound; this descriptor allows us to characterize how a note is initiated and produced by the voice.

[0067] More specific descriptors can also be defined, such as "soft attack," which is characterized by a gradual onset and generally a slight breath preceding the note. One advantage of the invention is its ability to characterize the gradation and amount of breath present, particularly through training different types of soft attacks found in a set of audio recordings of voices. Thus, the machine learning model makes it possible to deduce a characteristic output from a new input defining an individual's voice.

[0068] Similarly, a hard attack can be characterized by a sound that begins with a well-defined consonant and a rapid closure of the vocal cords, producing a clear and precise sound from the very beginning of the note. Quantifying the sound at the start, identifying the presence of a consonant, and measuring the duration of vocal cord closure allows for the characterization and / or labeling of recordings in order to train a machine learning model.

[0069] Similarly, the glottal attack is characteristic of a glottal stop, where the vocal cords close abruptly before sound production, creating subglottal tension. Consequently, this descriptor can be characterized using a machine learning model trained with labeled data that specifically characterizes the vocal cord closure time before sound production.

[0070] Vocal cord closure measurements can be performed using an imaging device positioned near the mouth or nose, such as a laryngoscope. Here are some preferred associations between specific microphones and certain descriptors to train the model during the initial learning of the first learned function or to best characterize the descriptors during the exploitation phase of the first learned function. These associations are advantageously made during training to select the microphones of interest for learning the first function. During the exploitation of the learned function, a single microphone may be used to acquire an individual's voice; however, some recording facilities may use different microphones to refine the estimation of one or more vocal descriptors of an individual.

[0071] Preferably, to characterize the descriptor related to the attack of the sound, at least one microphone in the acquisition system includes, for example, the reference microphone. The arrangement of the reference microphone allows for the capture not only of the individual's initiation of phonation, but also of the phonation itself, with high temporal fidelity compared to other arrangements. Optionally, to characterize the attack, the acquisition system includes another microphone such as a harmonic microphone and / or a projection microphone. While this other microphone (or microphones) is / are, due to its / their distance from the mouth, less temporally accurate than the reference microphone, it / these other microphones are, due to this distance, less exposed to mechanical, non-phonatory noises that may emanate from the mouth in connection with the initiation of phonation than the reference microphone.Thus, this / these other microphone(s) can limit / avoid the influence of these mechanical noises on the capture of the attack of the sound.

[0072] In some embodiments:

[0073] - The voice descriptor is a quantification of the pitch of the note,

[0074] - At least one microphone in the acquisition system is, for example, a harmonics microphone. Because it is positioned in front of the individual but away from the line of sight, the first harmonics measurement microphone can be protected from mechanical and / or plosive noises that could distort a sound frequency measurement. Regardless of the presence of at least a second harmonics measurement microphone, the acquisition system optionally includes, in addition to the first harmonics measurement microphone, a reference microphone. Since the harmonics measurement microphone's position differs from the reference microphone and / or the other harmonics measurement microphone, using them together can compensate for and / or cross-reference the vocal spectral components received by each microphone.

[0075] In some embodiments:

[0076] - The vocal descriptor is a quantity of air over the voice,

[0077] - at least one microphone in the acquisition system includes, for example, a reference microphone and a projection microphone. Since a projection measurement microphone is positioned further from the mouth than the reference microphone, the combination of a reference microphone and a projection measurement microphone makes it possible to differentiate between breath noise, a vocal component characteristic of air on the voice, and vocalization, a vocal component that projects better in space than breath noise.

[0078] In some embodiments:

[0079] - The voice descriptor is a quantification of the ending of a sound,

[0080] - At least one microphone in the acquisition system includes a reference microphone and a projection and / or harmonics microphone. Once the individual ceases vocalizing, a given microphone may continue to detect noise, for example, due to the individual's breathing and / or reverberations in the room caused by the previous vocalization. Since breathing is primarily a breath sound, it does not travel as well through space as vocalization. As reverberations propagate in directions other than the orientation axis, these sounds arrive at the microphones with a different time delay compared to the initial vocalization. Thus, the differences in the arrangement of these microphones can allow us to distinguish between noise due to vocalization and incidental noise that may be present after vocalization has ceased.

[0081] In some embodiments:

[0082] - The voice descriptor is a quantization of a harmonic mixture. - The acquisition system must include, for example, at least two harmonic microphones, or at least one reference microphone and at least one harmonic measurement microphone. Using the reference microphone in combination with any of the harmonic measurement microphones can enhance fundamental frequency capture compared to using a single harmonic measurement microphone, while using the first and second harmonic measurement microphones in combination can enhance harmonic capture compared to using a single harmonic measurement microphone.

[0083] In some embodiments:

[0084] - The vocal descriptor is a quantification of a degree of laterality,

[0085] - At least one microphone in the acquisition system includes at least one directional or lateral microphone. Thanks to its perpendicular position relative to the line of direction, the primary direction of voice propagation, this microphone allows for the measurement of voice propagation in other directions. Optionally, the acquisition system also includes a reference microphone. Due to the different orientations of these microphones, it is possible to differentiate between voice propagation along the line of direction and in a lateral direction. When the acquisition system includes at least the first and second lateral measurement microphones, it may be possible to measure asymmetry in voice propagation.

[0086] In some embodiments:

[0087] - The voice descriptor is a quantification of a degree of projectivity,

[0088] - At least one microphone in the acquisition system includes at least one projection microphone. Optionally, the acquisition system also includes a reference microphone. Due to the differences in distance between the mouth and each of these microphones, it is possible to measure how the sound changes with distance from the mouth.

[0089] In some embodiments:

[0090] - The vocal descriptor is a quantification of a degree of vibrato,

[0091] - At least one microphone in the acquisition system must include at least one laryngeal microphone, also known as a "stethomicrophone" or laryngophone. Vibrato is a periodic and intentional modulation of the fundamental frequency of glottal oscillations around a principal frequency that corresponds to the pitch. A laryngeal microphone, positioned on the individual's neck, primarily captures phonation by skin conduction, which can allow for more precise measurement of the fundamental frequency than by air capture. Due to its proximity to the vocal tract, such a microphone is less susceptible to other phenomena that can be mistaken for vibrato, even though these phenomena are not caused by modulations of glottal oscillations.

[0092] A fourth aspect of the invention relates to a device for characterizing at least one physiological descriptor of an individual's voice, comprising:

[0093] - at least one input interface configured to receive at least one characteristic value of a voice descriptor of the individual's voice;

[0094] - at least one configured processing unit implemented by computer:

[0095] - an implementation step of a second learned function generated from a second machine learning model trained from a second training domain, the second learned function receiving as input at least one characteristic value of the voice descriptor; the implementation step of the second learned function allowing to generate as output at least one score characterizing at least one physiological descriptor of the individual's voice.

[0096] - at least one output interface configured to return at least one score characterizing at least one physiological descriptor of said voice of said individual.

[0097] In this case, the characteristic value of a voice descriptor for an individual's voice can be an intrinsic characteristic value or one measured by a microphone. Recall that the intrinsic characteristic value of a descriptor is calculated from a function learned from several microphones in order to produce a characteristic value of the descriptor that is independent, or as independent as possible, of the microphone(s) that acquired the voice. In one embodiment, the voice descriptor used as input to the second learned function is one or more audio sequences of an individual directly measured by at least one microphone and possibly processed. This case corresponds to a specific instance in which the voice descriptor is directly an audio signal.

[0098] According to this fourth aspect of the invention, the characteristic value of a descriptor may be intrinsic or not. It is more generally called the characteristic value of the descriptor.

[0099] Thus, thanks to this device, it is possible, from characteristic values ​​of vocal descriptors, such as the attack of the sound, the air on the voice, or even the pitch of the note, to characterize a physiological descriptor of an individual's voice, without the intervention of an expert voice practitioner, such as a voice osteopath.

[0100] In some embodiments, at least one processing unit is further configured to implement the following on a computer:

[0101] - an implementation step of a comparison function configured to compare at least one intrinsic or non-intrinsic characteristic value of the voice descriptor to a threshold value, the implementation step of the comparison function allowing to generate as output at least one score characterizing at least one physiological descriptor of the individual's voice.

[0102] Thus, thanks to the invention, it is possible to characterize a physiological descriptor of an individual's voice from a single characteristic value of a vocal descriptor, and this without the intervention of an expert voice practitioner, such as a voice osteopath.

[0103] In some embodiments, at least one physiological descriptor is included among: vocal strain, muscular strain, and glottal stops. The physiological descriptor corresponds to a quantification of vocal strain, muscular strain, or one or more glottal stops.

[0104] In some embodiments, the vocal descriptor is included among a set of reference vocal descriptors including: a quantification of the attack of the sound, a quantification of a pitch, a quantity of air over the voice, a quantification of a sound termination, a quantification of a harmonic mixture, a degree of laterality, a degree of projectivity, a degree of vibrato.

[0105] In some embodiments, at least one processing unit is further configured to implement by computer, following the implementation step:

[0106] - an implementation step of a third learned function generated from a third machine learning model trained from a third training domain, the third learned function receiving as input at least one score characterizing at least one physiological descriptor of the individual's voice, so as to generate as output at least one percentage value characterizing a contribution of at least one stage of the phonatory apparatus of said individual to said physiological descriptor.

[0107] In some embodiments, the at least stage is one of: the individual's lung, the individual's larynx, the individual's pharynx, an individual's articulator.

[0108] A fifth aspect of the invention relates to a method for characterizing at least one physiological descriptor of an individual's voice implemented by the device described above, comprising:

[0109] - a stage of receiving at least one characteristic value of a voice descriptor of said individual's voice;

[0110] - an implementation step of a second learned function generated from a second machine learning model trained from a second training domain, said second learned function receiving as input said at least one characteristic value of the voice descriptor; the implementation step of the second learned function enabling the generation as output of at least one score characterizing said at least one physiological descriptor of said voice of said individual.

[0111] According to this aspect of the invention, the process can be carried out with an intrinsic or non-intrinsic characteristic value of at least one descriptor.

[0112] A sixth aspect of the invention relates to a method for characterizing a contribution of at least one stage of the phonatory apparatus to a physiological descriptor of an individual's voice implemented by the device previously described, comprising: - the steps of the method for characterizing at least one physiological descriptor of the voice of said individual previously described;

[0113] - a step called additional implementation step of a third learned function generated from a third machine learning model trained from a third training domain, said third learned function receiving as input said at least one score characterizing said at least one physiological descriptor of said voice of said individual, so as to generate as output at least one percentage value characterizing a contribution of at least one stage of the phonatory apparatus of said individual to said physiological descriptor.

[0114] A seventh aspect of the invention relates to a method for predicting a voice pathology in an individual, the method comprising:

[0115] - the steps of the process to characterize at least one physiological descriptor of the voice of the individual previously described,

[0116] - the steps of the process to characterize a contribution from at least one level of the phonatory apparatus of said individual to a previously described physiological descriptor,

[0117] - a step of predicting the pathology of the individual's voice based on the contribution of at least one level of the individual's phonatory apparatus.

[0118] Different types of measurement microphones can be used and configured for recording an individual's voice. They can be used in combination when different types of measurement microphones are used.

[0119] Other types of microphones are presented here in addition to those already described.

[0120] One type of measurement microphone provides a flat frequency response. These measurement microphones have a very flat frequency response. This means they are designed to respond equally to all audible frequencies, without coloration or emphasis of any particular frequency. This characteristic allows for accurate measurements of the sound as it is, without distortion introduced by the microphone itself. A second type of measurement microphone provides a measurement of directivity. These allow for an estimation of sound dispersion.

[0121] A third type of measurement microphone offers high sensitivity and good accuracy. These microphones allow for the measurement of audio signals with low sound levels. This characteristic enables precise acoustic measurements, including noise level measurements.

[0122] Simple measurement microphones can also be used, such as studio microphones.

[0123] Brief description of the figures

[0124] Other features and advantages of the invention will become apparent from the detailed description that follows, with reference to the attached figures, which illustrate:

[0125] Fig. 1: an example of an acquisition system configured to acquire a set of audio signals that can be used in a method to determine an intrinsic characteristic value of a voice descriptor of an individual's voice according to the invention;

[0126] Fig. 2: an example of a device called a descriptor device configured to implement the method for determining an intrinsic characteristic value of a voice descriptor of an individual's voice according to the invention;

[0127] Fig. 3: an example of a set of steps that can be carried out to implement the method for determining an intrinsic characteristic value of a voice descriptor of an individual's voice according to the invention;

[0128] Fig. 4: an example of a system configured to implement the method for determining an intrinsic characteristic value of a voice descriptor of an individual's voice according to the invention;

[0129] Fig. 5: an example of a device called a pathology device configured to implement a method for characterizing at least one physiological descriptor of an individual's voice according to the invention;

[0130] Fig. 6: an example of a set of steps that can be carried out to implement the method for characterizing at least one physiological descriptor of an individual's voice according to the invention.

[0131] Description of the Invention In this document, the word "micro" (and variants) means "microphone" (and variants), and the expression "harmonic / projectivity / lateral / ambient projectivity / ambient reflectivity microphone" (and variants) is interchangeable with the expression "harmonic / projectivity / lateral / ambient projectivity / ambient reflectivity microphone" (and variants). For the sake of simplicity, the terms "micro" and "microphone" denote both the sensor that transforms an audio signal into an electrical signal and the electrical signal itself. In this document, no particular structure is implied for any microphone arranged at a distance from the individual—in particular, such a microphone may be omnidirectional or not.As a non-limiting example, an omnidirectional microphone is "pointed" towards a sound source when its sensor is closer than the sensor support to the sound source.

[0132] One aspect of the invention relates to a method 100, a device called a descriptor device 10, and a system 20 for determining a value D of a vocal descriptor of an individual's voice. Another aspect of the invention relates to a method 200 and a device called a pathology device 30 for characterizing at least one physiological descriptor of an individual's voice. A further aspect of the invention relates to a method for characterizing the contribution of at least one level of the vocal tract to a physiological descriptor of an individual's voice. Finally, another aspect of the invention relates to a method for predicting a pathology of an individual's voice.

[0133] Figure 1 represents an example of an acquisition system 15 configured to acquire and record a sequence of at least one sound S emitted by the voice of an individual I, so as to obtain at least one audio signal SAi, SA2,.. SAN encoding the sequence of at least one sound, with N an integer greater than or equal to 1.

[0134] In operation, the acquisition system 15 is placed in an indoor space, such as a recording room, a studio, an examination room.

[0135] The acquisition system 15 includes at least one microphone positioned to acquire the sequence of at least one sound S, the microphone being positioned at a predetermined acquisition position (or location). When the acquisition system 15 includes at least two microphones, these are positioned at separate acquisition positions.

[0136] Some embodiments of the acquisition system 15 will be described below, which can be combined together.

[0137] In some embodiments, the acquisition system 15 comprises between 1 and 13 microphones.

[0138] In some embodiments such as that illustrated in Figure 1, the acquisition system 15 includes 13 microphones (Mo, Mi, M2... M12).

[0139] In some embodiments, the acquisition system 15 includes a single microphone.

[0140] When individual I emits the sequence of at least one sound S for a recording by the acquisition system 15, the orientation of his body from back to front, defining a direction called orientation direction x and a line called orientation line passing through the center of gravity of the individual and parallel to the orientation direction x, and the body of individual I defining approximately a plane P perpendicular to the direction x.

[0141] Advantageously, a microphone Mo, serving as a geometric reference and hereafter referred to as the reference microphone Mo, is positioned, while in operation, in front of the mouth of individual I, whose sequence of at least one sound S is acquired and recorded. The individual is standing when emitting the sequence of at least one sound S. The reference microphone Mo is located in a plane Pi, a distance from plane P between 1 cm and 40 cm, and preferably between 5 cm and 30 cm. At this distance (or these distances), the reference microphone Mo is positioned at the mouth so as to capture the vocalization itself, as well as any incidental non-vocalized noises, such as articulation noises (e.g., mechanical mouth noises) and breath sounds. The signal from the reference microphone can then highlight the sequence of actions undertaken by the individual during phonation, the onset of phonation, and / or the cessation of phonation.

[0142] In some embodiments, the acquisition system 15 comprises a first harmonic measurement microphone M1 and a second harmonic measurement microphone M2, positioned at acquisition locations on the same vertical line positioned in front of the individual when he or she emits the sequence of at least one sound S. "Vertical line" means a line oriented along the vertical axis, i.e., the axis of gravity. For example, the first harmonic measurement microphone M1 is positioned on the vertical line below the reference microphone M0, and the second harmonic measurement microphone M2 is positioned on the vertical line above the reference microphone M0.

[0143] In this document, the concept of "harmonic measurement" denotes the measurement of wave amplitude whose frequency is an integer multiple (plus or minus 0.1) of a fundamental frequency of a sound.

[0144] The concepts of "below" and "above" are interpreted in terms of a vertical line extending from the floor to the ceiling of a room and passing through the reference microphone. The position on the z-axis gives the altitude and allows any object to be positioned along this axis relative to the reference microphone. Since the altitude of the reference microphone is generally adjusted according to the height of the individual and the position of their mouth, any object in the room can also be positioned on the vertical axis relative to an individual's mouth.

[0145] Advantageously, the acquisition position of the first harmonic measurement microphone M1, placed above the mouth, allows for the acquisition of an audio signal containing more information about the high harmonic components contained in the sequence of at least one sound emitted by individual I.

[0146] Advantageously, the acquisition position of the second harmonic measurement microphone M2, placed below the mouth, allows the acquisition of an audio signal including more information on low harmonic components contained in the sequence of at least one sound S emitted by individual I.

[0147] As an example, the first harmonic measurement microphone M1 and the second harmonic measurement microphone M2 are positioned at a distance of between 40 and 60 cm from the reference microphone Mo. Preferably, the first harmonic measurement microphone M1 and the second harmonic measurement microphone M2 are positioned at a distance of 50 cm from the reference microphone Mo. In general, in this document, separating two microphones from each other in space allows for the capture of spatial phenomena of the voice, such as its propagation ability and localized and / or directional variations, and / or spectral phenomena such as its harmonic mixing, and / or phenomena of interaction with space such as reverberation, reflection, and / or spatial uniformity.

[0148] In some embodiments, the acquisition system 15 includes a first projectivity measurement microphone M31 and a second projectivity measurement microphone M41, positioned in front of the individual along the x-axis. The two microphones can be positioned on vertical planes at different depths of the room relative to the individual. For example, a first projectivity microphone is positioned 2 m from the individual in the horizontal plane containing the reference microphone, and a second projectivity microphone is positioned 4 m from the individual in the horizontal plane containing the reference microphone.

[0149] By their positions, the first projectivity measurement microphone M31 and the second projectivity measurement microphone M41 are configured to optimally measure the projectivity of the voice of individual I. In other words, the acquisition positions of the first projectivity measurement microphone M31 and the second projectivity measurement microphone M41 allow the acquisition of audio signals including information characterizing the degree of projectivity in the sequence of at least one sound S emitted by individual I.

[0150] Projectivity refers to a quantification of an individual's ability to project their voice far into space and be heard clearly at a distance. Projectivity is influenced by several factors, including vocal power (the force with which sounds are produced by the vocal cords), resonance (the effective use of resonating cavities such as the mouth, nose, and pharynx to amplify sound), and the proper use of breath control and diaphragmatic support. Finally, clarity of diction (the ability to articulate sounds distinctly) can also affect projectivity.

[0151] Voice "directivity" refers to how the sound of the voice propagates in a given direction in space. It describes the spatial distribution of sound energy. It can be quantified using various measurement microphones. For example, it is possible to measure lateral directivity, that is, the dispersion of sound relative to the left or right of the individual. This angle can be measured from the individual's mouth and located relative to a vertical plane perpendicular to the plane of the individual's chest or a plane parallel to the orientation of their nose. It is also possible to measure vertical directivity, that is, the dispersion of sound relative to the top or bottom of the mouth.

[0152] Voice directionality can be influenced by several factors such as the shape of the mouth during phonation, the orientation of the mouth, the orientation and shape of the lips, etc.

[0153] For example, the first directivity measurement microphone M3 and the second directivity measurement microphone M4 are positioned on a second vertical line, one above the other. In another example, the first projectivity measurement microphone M3 and the second projectivity measurement microphone M4 are positioned at a distance from plane P ranging from 1 cm to 10 m. The distance between the two planes can be adjusted according to the dimensions of the recording room.

[0154] For example, a third directivity measurement microphone M3' and a fourth directivity measurement microphone M4' are positioned on a second vertical line, one above the other. In one example, the first directivity measurement microphone M3' and the second directivity measurement microphone M4' are positioned at a distance from plane P ranging from 1 cm to 10 m. The distance between the two planes can be adjusted according to the size of the recording room.

[0155] For example, a fifth directivity measurement microphone (M3) and a sixth directivity measurement microphone (M4) are positioned at the same height within a horizontal plane, for example, one on the right side of the individual and the other on the left. In another example, the first directivity measurement microphone (M3) and the second directivity measurement microphone (M4) are separated by a distance between 1 cm and 10 m in a horizontal plane. The distance between the two planes can be adjusted according to the size of the recording room. Projectivity and directivity are two related measurements: the more directional an audio signal, the greater its projection. Conversely, the more projective a signal, the lower its directivity.

[0156] The dispersion of a signal can be heterogeneous within a solid projection cone. For example, an audio signal can have lateral directivity, that is, a characterization of the distribution of audio power according to azimuth or height.

[0157] Laterality measurement quantifies the dispersion of a signal in the horizontal plane on either side of the individual. This measurement is very similar to the measurement of lateral directivity. Laterality measures the amount of signal power distributed laterally to the individual, that is, the variation in intensity in the horizontal plane.

[0158] In some embodiments, the acquisition system 15 comprises a first laterality measurement microphone Ms and a second laterality measurement microphone Me, positioned in the plane Pi, the plane of the reference microphone Mo, and at the same height along the vertical axis, i.e., the height of the mouth of individual I. By their positions, the first laterality measurement microphone Ms and the second laterality measurement microphone Me are configured to optimally measure the degree of laterality of the voice of individual I. In other words, the acquisition positions of the first laterality measurement microphone Ms and the second laterality measurement microphone Me make it possible to acquire audio signals comprising information characterizing the degree of laterality in the sequence of at least one sound S emitted by individual I.

[0159] In one example, the first laterality measurement microphone Ms and the second laterality measurement microphone Me are each located between 5 cm and 2 m from the reference microphone Mo, and preferably between 40 cm and 60 cm. Preferably, the first laterality measurement microphone Ms and the second laterality measurement microphone Me are each located 50 cm from the reference microphone Mo.

[0160] In some embodiments, the acquisition system 15 includes a reflectivity measurement microphone M? positioned behind the individual I (along the x-axis). The reflectivity measurement microphone M? is positioned along the x-axis at the same height as the reference microphone Mo. By its position, the reflectivity measurement microphone M? is configured to optimally measure the degree of reflectivity of the individual I's voice. In other words, the acquisition position of the reflectivity measurement microphone M? allows for the acquisition of an audio signal containing information characterizing the degree of reflectivity in the sequence of at least one sound S emitted by the individual I. Reflectivity is understood to be the quantification of the power of the audio signal reflected from an obstacle, such as a wall, ceiling, or other object.According to an example, the reflectivity measurement microphone M? is distant from the plane Pi of the reference microphone Mo by a distance of between 50 cm and 4 m. The plane containing the reflectivity measurement microphone M? is preferably located behind the individual, that is to say in the part of the space located on the side of the vertical plane containing the individual and arranged on the other side of his face.

[0161] In some embodiments, the acquisition system 15 includes a first ambient projectivity measurement microphone Ms and a second ambient projectivity measurement microphone Mg, positioned in front of individual I. By their positions, the first ambient projectivity measurement microphone M8 and the second ambient projectivity measurement microphone M9 are configured to optimally measure the degree of laterality of individual I's voice. In other words, the acquisition positions of the first ambient projectivity measurement microphone M8 and the second ambient projectivity measurement microphone M9 allow the acquisition of audio signals comprising information characterizing the degree of ambient projectivity in the sequence of at least one sound S emitted by individual I."Ambient projectivity" refers to the ability of a sound to fill a space in such a way as to create a specific atmosphere or ambiance, taking into account how the acoustic characteristics of the environment affect the perception of the sound. Ambient projectivity thus allows for the joint quantification of reflection, reverberation, and absorption by surrounding surfaces, as well as the uniformity of distribution—that is, the ability of the sound to reach all points in space in a balanced manner.

[0162] In one example, the first ambient projectivity measurement microphone Ms and the second ambient projectivity measurement microphone Mg are positioned at a distance from plane Pi (the plane of the reference microphone Mo) of between 10 cm and 6 m, and preferably between 2 m and 5 m. However, the distance is adjusted according to the dimensions of the room in which the recording takes place. Advantageously, the first ambient projectivity measurement microphone Ms and the second ambient projectivity measurement microphone Mg are positioned at a height higher than the height of the reference microphone Mo, preferably higher by a distance of between 10 cm and 2 m, and preferably higher by 50 cm.

[0163] In some embodiments, the acquisition system 15 includes a first ambient reflectivity measurement microphone Mw and a second ambient reflectivity measurement microphone Mu, positioned behind individual I. In one example, the ambient reflectivity microphones Mw and Mu are arranged in the same vertical plane, on either side of the longitudinal x-axis. By their positions, the first ambient reflectivity measurement microphone Mw and the second ambient reflectivity measurement microphone Mu are configured to optimally measure the degree of laterality of the reflected portion of the audio signal from individual I's voice.In other words, the acquisition positions of the first ambient reflectivity measurement microphone Mw and the second ambient reflectivity measurement microphone Mu allow the acquisition of audio signals including information characterizing the degree of ambient reflectivity in the sequence of at least one sound S emitted by individual I.

[0164] Ambient reflectivity refers to the quantification of a sound's ability to reflect homogeneously and / or uniformly in space after reflecting off obstacles.

[0165] According to an example, the first ambient reflectivity measurement microphone M and the second ambient reflectivity measurement microphone Mu are located at a distance from plane Pi (plane of reference microphone Mo) of a distance between 50 cm and 5 m. Advantageously, the first ambient reflectivity measurement microphone Mw and the second ambient reflectivity measurement microphone Mu are positioned at a height higher than the height of the reference microphone Mo, preferably higher by a distance between 0 and 2 m, preferably higher by 50 cm.

[0166] In some embodiments, the acquisition system 15 includes a microphone M12 called a laryngeal pickup microphone, positioned near the throat of individual I. By its position, the laryngeal pickup microphone M12 is configured to optimally measure the vibrations of the larynx produced during the phonation of individual I's voice. In other words, the acquisition position of the laryngeal pickup microphone M12 makes it possible to acquire an audio signal comprising information characterizing vocal vibrations, in particular laryngeal vibrations, by providing precise data on the fundamental frequency of the voice in the sequence of at least one sound S emitted by individual I.

[0167] In some embodiments, each microphone in the plurality of microphones of the acquisition system 15 has the same type, in other words, the same technical characteristics, such as bandwidth, frequency response, sensitivity, and directivity. In this case, only the position (i.e., the location) of each microphone in the acquisition system 15 relative to the position of individual I distinguishes it from the other microphones in the acquisition system 15.

[0168] In operation, when positioned relative to an individual I standing and emitting a sequence of at least one sound S, the acquisition system 15 collects a plurality of input audio signals from each of the system's microphones. Each input audio signal encodes the sequence of at least one sound S emitted by individual I. In other words, each audio signal is a representation of the sequence of at least one sound S. For example, if the audio signals are analog, each audio signal comprises an electrical voltage signal with a variable voltage level. In another example, if the audio signals are digital, each audio signal comprises a series of binary numbers.

[0169] Advantageously, since each microphone in the acquisition system 15 is positioned at a distinct location, the input audio signal from a microphone is specific to that location.

[0170] The acquisition system 15 described above can be used in a training phase to obtain audio signals constituting training data for a machine learning model, generating a first learned function that will be used in the method 100 to determine a value for a voice descriptor of an individual's voice according to the invention described below. In this case, the acquisition system comprises at least two microphones positioned at two different locations relative to the position of an individual whose voice is recorded by the acquisition system 15.

[0171] The acquisition system 15 described above can be used in an operational phase to obtain audio signals constituting the input data for the learned function generated by the preceding machine learning model and used in the method 100 to determine a value for a voice descriptor of an individual's voice according to the invention described below. The acquisition system 15 comprises at least one microphone.

[0172] The method 100 for determining a value of a vocal descriptor of an individual's voice according to the invention will now be described. Applications of the method 100 include, but are not limited to: establishing a vocal assessment of an individual for the purpose of vocal coaching, establishing a vocal assessment of an individual for the diagnosis of a vocal pathology or voice dysfunctions, and characterizing human voices for multimedia applications.

[0173] The vocal assessment performed according to the invention provides a detailed representation of an individual's voice characteristics. This characterization helps reduce diagnostic errors and / or refine recommendations for preventing damage to the body caused by poor voice projection habits, etc.

[0174] The method of the invention makes it possible to improve the early detection of functional voice problems.

[0175] From another perspective, voice characterization allows for better characterization of objective indicators that are precursors to pathological conditions, for example, Parkinson's disease.

[0176] According to other applications, these descriptors can be used to refine lie detection when these descriptors are used in conjunction with audio signal processing, data characterizing an individual's breathing, and data characterizing said individual's heartbeat.

[0177] Finally, the method of the invention makes it possible to classify characteristic descriptors of an individual's voice, which allow for the determination of a specific corrective action to be taken by the individual. For example, increasing lung air intake, taking deeper breaths, opening the mouth wider than a reference opening, or calibrating recording equipment based on the quantification of the measured and classified descriptors.

[0178] In one embodiment, voice characterization allows for better corrections, for example, in the reconstruction of an individual's voice. Typically, during the recording of a phonogram, disc, or dubbing, corrective factors can be introduced to address certain aspects of the voice based on the quantification of the measured voice descriptors.

[0179] The process 100 can be implemented using a device 10 as illustrated in Figure 2 and named the descriptor device, which represents a schematic diagram showing the components of an example of the device 10.

[0180] The device described in descriptor 10 can be implemented as a single hardware device, for example, as a desktop personal computer (PC), laptop computer, personal digital assistant (PDA), smartphone, smartwatch, server, or console, or it can be implemented on separate, interconnected hardware devices linked by one or more communication links, with wired and / or wireless segments. The device 20 can, for example, communicate with one or more cloud computing systems, one or more servers, or remote devices to implement the functions described in this description for the device in question. The device 20 can also be implemented itself as a cloud computing system.

[0181] As shown in Figure 2, the device of descriptor 10 comprises a computer, this computer including a memory 11 for storing program instructions that can be loaded into a circuit 12 and adapted to cause a circuit to execute steps of the process 100 illustrated in Figure 3, described below, when the program information is executed by the circuit. The memory can also store data and information useful for the execution of the steps of the present invention as described below.

[0182] Circuit 12 could be, for example:

[0183] - a processor or processing unit adapted to interpret instructions in a computer language, the processor or processing unit being able to understand, be associated with, or be attached to a memory containing the instructions, or

[0184] - the combination of a processor / processing unit and a memory, the processor or processing unit being adapted to interpret instructions in a computer language, the memory containing said instructions, or

[0185] - an electronic circuit board in which the steps of the invention are described in silicon, or

[0186] - a programmable electronic chip such as an FPGA chip (for "Field-Programmable Gate Array").

[0187] Memory 11 may include random access memory (RAM), cache memory, non-volatile memory, backup memory (e.g., programmable or flash memory), read-only memory (ROM), a hard disk drive (HDD), a solid-state drive (SSD), or any combination thereof. The ROM of memory 11 may be configured to store, among other things, an operating system and / or one or more computer program codes for one or more software applications. The RAM of memory 11 may be used by circuit 12 for temporary data storage.

[0188] The computer may also include an input interface 13 for receiving input data and an output interface 14 for providing output data. Examples of input and output data will be provided later.

[0189] To facilitate interaction with the computer, a 15-inch screen and a 16-inch keyboard can be provided and connected to the computer circuit.

[0190] The descriptor device 10 is configured to implement the computer-implemented method 100 for determining a value D of a voice descriptor of an individual's voice I according to the invention. Figure 3 is a flowchart representing an example of the set of steps that can be performed to implement the method 100 for determining an intrinsic characteristic value D of a voice descriptor of an individual's voice I.

[0191] In the present application, a vocal descriptor of a voice is understood to mean a characteristic capable of describing an aspect of the sound of an individual's voice. The vocal descriptor characterized by the method 100 according to the invention can be chosen from a set of vocal descriptors called reference vocal descriptors, comprising: an attack of the sound, a pitch, a quantity of air over the voice, a termination of the sound, a harmonic mixture, a degree of laterality, a degree of projectivity, a degree of vibrato.

[0192] These reference vocal descriptors are classically known to human voice experts, who can evaluate them for a given individual through a physical examination in a given context, such as vocal coaching, or a clinical voice examination.

[0193] One advantage of the invention is that it allows for the objective characterization of these criteria through the implementation of a machine learning algorithm. Since the descriptor data is annotated—that is, labeled or categorized by domain experts—the classifier will provide a classification and quantification of the descriptors specific to the industry. Depending on the various implementations of the invention, it is possible to configure the probability threshold levels of the machine learning algorithm's predictions, thus customizing the classifier to meet the specific needs of each professional.

[0194] In a first REC step, the descriptor device 10 receives and then stores in memory 11 at least one input audio signal SAi, SA2, ... SAN.

[0195] In some embodiments, the at least one input audio signal SA1, SA2, ... SAN is one or a plurality of signals from the previously described acquisition system 15.

[0196] In other embodiments, the at least one input audio signal SA1, SA2, ... SAN is one or a plurality of signals from any other acquisition system. These other signals encode a sequence of at least one sound that was emitted by individual I. In one example, the other signals each encode the sequence of at least one sound S, recorded from positions distinct relative to the individual's position when the sequence of at least one sound S was emitted.

[0197] In some embodiments, depending on the voice descriptor from which an intrinsic characteristic value is to be obtained with process 100, the at least one input audio signal SAi, SA2, ... SAN corresponds to one or a plurality of audio signals that have been acquired by a set of microphones configured to capture a set of audio signals including non-zero information relating to that voice descriptor.

[0198] The invention also relates to a method for first recording and calculating descriptors from a maximum number of microphones and then selecting the most representative descriptors for each descriptor. The selection of the most representative descriptors can be carried out by comparing the range of values ​​of the descriptor by varying the input signal according to an aspect of the voice, for example, its power, harmonics, vibrato, etc.

[0199] Thus, as an example, when the vocal descriptor is the attack of the sound, a preferred set of microphones configured to acquire the input audio signals SA1, SA2, ... SAN includes: the reference microphone and the harmonic measurement microphones. Other microphones can also be selected to enrich the characterization of this descriptor.

[0200] For example, when the vocal descriptor is the degree of projectivity, a preferred set of microphones configured to acquire the input audio signals SA1, SA2, ... SAN might include projectivity, directionality, and laterality microphones. Other microphones could also be selected to further enhance the characterization of this descriptor.

[0201] For example, when the voice descriptor is the degree of reflectivity, a preferred set of microphones configured to acquire the input audio signals SA1, SA2, ... SAN includes reflectivity, ambient projectivity, or ambient reflectivity microphones. Other microphones can also be selected to further enhance the characterization of this descriptor.

[0202] Advantageously, in certain embodiments, in a second reception step REC2, the descriptor device 10 receives and then stores configuration information in memory 11. The configuration information includes, for each input audio signal among the at least one input audio signal (SAi, SA2, ... SAN), position information for an acquisition microphone associated with that input audio signal relative to a position of the individual I when they emitted the sequence of at least one sound S. More precisely, the position information associated with an input audio signal includes the position of the microphone with which the sequence of at least one sound S was acquired so as to produce the input audio signal.

[0203] After receiving the input audio signals SA1, SA2, ... SAN and possibly configuration information, in a step called the first implementation step IMP1, circuit 12 of device 10 implements a first learned function generated from a first machine learning model previously trained from a first training domain containing initial training data. The initial training data includes at least two audio training signals encoding the same speech sequence called the training sequence. By speech sequence, we mean a sequence of at least one sound emitted by a human being, for example, a performer. The training of the first machine learning model from the first training domain will be described later in this document.

[0204] The first learned function receives at least one input audio signal SA1, SA2, ... SAN as input, so as to generate as output, for each of the input audio signals, a corresponding value of a descriptor characteristic of the individual's voice. In other words, a plurality of values ​​of the voice descriptor of individual I's voice is obtained at the end of the first implementation step IMP1, each corresponding to one of the at least one input audio signal SA1, SA2, ... SAN.

[0205] In one example, the values ​​of the voice descriptor are scores between 0 and 10. In another example, the values ​​of the voice descriptor are percentages between 0 and 100%. Advantageously, the values ​​of the voice descriptor characterize the intensity of that voice descriptor. Thus, a score of 8 out of 10 for the voice attack descriptor will be characteristic of a strong attack.

[0206] In some embodiments, the first machine learning model generating the first learned function is based on a convolutional neural network (CNN). Alternatively, or in combination, other architectures can be used to build the first machine learning model. Thus, in some alternative embodiments, or those that can be combined with the previous embodiments, the first machine learning model can be based on architectures using temporal correlations, such as recurrent neural networks (RNNs, LSTMs, U-Nets).

[0207] Advantageously, when the at least one input audio signal SA1, SA2, ... SAN comprises at least two input audio signals corresponding to two different acquisition positions of the common sequence of at least one sound emitted S by the voice of individual I, the method 100 includes a step called the second implementation step IMP2, in which the circuit 12 of the device 10 implements a function called the global function. The global function receives as input the corresponding values ​​of the characteristic descriptor of the individual's voice for each of the input signals obtained during the first implementation step IMP1, so as to generate as output the intrinsic characteristic value D of the voice descriptor of the individual's voice.

[0208] In some embodiments, the global function is the maximum function. In other embodiments, the global function is the minimum function. In still other embodiments, the global function is an averaging function. In yet other embodiments, the global function is a weighted sum of the corresponding values ​​of the characteristic descriptor of each of the input audio signals SAi, SA2, ... SAN.

[0209] When at least one input audio signal comprises a single input signal, the intrinsic characteristic value D corresponds to the value obtained during the first implementation step IMP1. This is then the corresponding value of at least one first voice descriptor relating to the microphone that acquired said voice of said individual which was calculated by the first learned function.

[0210] According to the invention, a system 20 for determining an intrinsic characteristic value D of a voice descriptor of an individual's voice, schematically represented in Figure 4, comprises the descriptor device 10 and the acquisition system 15 described above. In operation, the acquisition system 15 then implements a step called the ACQ acquisition step of a sequence of at least one sound emitted by individual I in order to obtain at least one input audio signal SA1, SA2, ... SAN. The acquisition system 15 and the descriptor device 10 are interconnected by one or more communication links, with wired and / or wireless segments. Thus, at least one audio input signal SA1, SA2, ... SAN is then transmitted via the communication link(s) to the descriptor device 10. For example, the descriptor device 10 receives at least one audio signal SA1, SA2, ... SAN via its input interface 13.The process 100 can be implemented after the ACQ acquisition step by the descriptor device 10. Different audio tracks generated by different sound recording and capture devices can be routed to the same acquisition equipment to use all the input audio signals.

[0211] In some embodiments, the method 100 makes it possible to determine at least two characteristic values ​​of two different voice descriptors of an individual's voice, by implementing the first implementation step IMP1 and the second implementation step IMP2 following the reception of at least one input audio signal SA1, SA2, ... SAN. According to one embodiment, the two implementation steps are carried out simultaneously.

[0212] In some embodiments, process 100 allows for the determination of a plurality of characteristic values, each representing a voice descriptor of an individual's voice, by implementing the first implementation step IMP1 and the second implementation step IMP2 following the reception of at least one input audio signal SA1, SA2, ... SAN. According to one embodiment, the two implementation steps are performed simultaneously.

[0213] Training the first machine learning model

[0214] In some embodiments, the first machine learning model is trained prior to the implementation of process 100.

[0215] It is recalled that training a machine learning model consists of determining a set of model parameters and / or teaching the model a corresponding learned function from training data, so that the model can then predict a label of new data received.

[0216] Advantageously, the training data is annotated (or labeled, or tagged). Advantageously, the training data is annotated manually by a panel of experts. By expert, we mean a person recognized as having mastered the knowledge of a given field. For example, an expert might be a singing teacher, a sound engineer, an acoustician, or a doctor.

[0217] In some embodiments, the initial training data was collected and annotated during a research session involving a plurality of experts and one or more artists.

[0218] During the research session, the artists perform one or more vocal performances which are recorded by at least two microphones spatially arranged in an enclosed space, and which are observed by a panel of experts. For example, the 15 acquisition system can be used to record the vocal performance(s).

[0219] Advantageously, each recording of the same vocal performance is associated, for example by a label, with information about the microphone that produced the recording. This information could be the microphone's position relative to the artist's position when performing the vocal performance. It could also be the type of microphone, for example: direct microphone, high-harmonic measurement microphone, low-harmonic measurement microphone, projectivity measurement microphone, lateral measurement microphone, reflectivity measurement microphone, ambient projectivity measurement microphone, ambient reflectivity measurement microphone, or laryngeal microphone.

[0220] For each recording of a given vocal performance, each expert assigns a corresponding value to a vocal descriptor of the artist's voice. The vocal descriptor is a reference characteristic such as: attack, pitch, breath support, ending, harmonic blend, degree of laterality, degree of projection, or degree of vibrato. Then, based on the values ​​assigned by each expert, a final value, known as the characteristic value of the vocal descriptor, is assigned to the vocal performance for each recording. In acoustics, "harmonic blend" (sometimes called "harmonic frequency blending / distribution") refers to the distribution of sound between several harmonic frequencies relative to the fundamental frequency of a sound.

[0221] In acoustics, the term "pitch" is traditionally understood to refer to the fundamental frequency of a sound. In an audio sequence emitted by an individual, the frequency with the highest spectral density or the frequency with the highest spectral power in the emitted spectrum is generally considered. In other cases, the frequency with the greatest amplitude may be used.

[0222] Pitch is generally used to assess an individual's spectral range in the low and high registers. As such, the invention allows for the deduction, from among the generated descriptors, of the pitch corresponding to all or part of an acquired sequence or a plurality of acquired sequences. For the purpose of determining the pitch value, an average can be calculated over a sample.

[0223] In acoustics, vibrato is classically defined as an oscillation of a frequency component within an acquired audio sequence. In one embodiment, vibrato is estimated from the fundamental frequency of the audio sequence. In another example, vibrato is estimated from the oscillations of several characteristic frequencies within the spectrum of the sample under consideration. The characteristic frequency may be the one with the greatest amplitude, the one with the highest spectral power, or the one with the highest spectral density.

[0224] In acoustics, the term "air on the voice" is typically understood to refer to vocalization preceded and / or accompanied by a breathy noise. This breathy noise is usually caused by the incomplete closure of the vocal folds during phonation. When produced before vocalization, it is considered a non-vocalized and non-harmonic noise; when produced during vocalization, it is a non-harmonic component complementing the harmonic component of vocalization.

[0225] Breath noise is detectable at close range, but barely detectable at a distance from the subject's mouth. As a non-limiting example, breath noise can be detected by a microphone placed preferably less than 40 cm from an individual's mouth. Such an arrangement allows, for example, the detection of a sound pressure level (SPL) of 15 dB or more and 30 dB or less for a typical adult human airflow. In English-language literature, "SPL" refers to "Sound Pressure Level," meaning a level of acoustic pressure. This contrasts with a sound pressure level of 1 dB or more and 5 dB or less at a distance of 10 meters from the subject's mouth, which can render breath noise inaudible to the human ear at that distance.

[0226] "Projectivity" refers to a subject's ability to carry their voice far in space and clearly over a distance. As a non-limiting example, a voice is projective when it loses less than 6 dB SPL for every doubling of distance from the mouth, over a distance ranging from 10 centimeters to 5 meters from the subject's mouth, measured along the line of orientation, in a controlled environment such as a low-reverberation room or a (semi-)anechoic chamber (for example, a space conforming to ISO 3382-1 for measurement rooms, with a reverberation time TReo of 0.3 seconds or less in a frequency band ranging from 500 Hz to 3000 Hz).

[0227] Projectivity can also be defined as the ability of a voice to maintain its perceived intensity at a distance, that is, to maintain a high sound level despite the distance from the receiver (microphone or ear). It generally depends on the vocal power emitted (initial sound pressure), the timbre (presence of harmonics), and the directionality of the emission (how the voice is projected in space).

[0228] As an example, voice projection can be measured using microphones positioned at varying distances along a frontal axis from the speaker's mouth. By positioning the microphones at increasingly longer distances and repeating the transmission of a stable vocal sequence, or by recording sequences simultaneously with different microphones, it is possible to calibrate an individual's voice projection.

[0229] The invention enables the training of a machine learning model to produce a projectivity indicator from a single microphone. To this end, the machine learning model is trained to generate descriptor values ​​based on the acquired audio sequence. Here are some theoretically possible values:

[0230] Measured SPL Distance Variation

[0231] 10 cm 85 dB —

[0232] 20 cm 79 dB -6 dB

[0233] 40 cm 73 dB -6 dB

[0234] 80 cm 67 dB -6 dB

[0235] 160 cm 61 dB -6 dB

[0236] 320 cm 58 dB -6 dB

[0237] 640 cm 55 dB -6 dB

[0238] We consider a loss of 6 dB SPL at each doubling for a voice whose projectivity is considered to be a standard.

[0239] When a speaker loses less than 6 dB SPL with each doubling, the voice can be considered projective. If the speaker loses more than 6 dB with each doubling, the voice can be considered less projective.

[0240] Here is an example illustrating the difference in projective voice, with variations in IB SPL in one case and 7 to 8 dB in the other, at each doubling of distance. Distance Projective Voice Less Projective Voice 10 cm 85 dB SPL 78 dB SPL 20 cm 80 dB SPL 71 dB SPL 40 cm 75 dB SPL 63 dB SPL 80 cm 70 dB SPL 55 dB SPL

[0241] An indi of projectivity pc ut then be generated by considering these variations and in normalizing or associating them on a scale of variation.

[0242] During training, the invention allows, from the recording of an audio sequence using several microphones arranged in different positions, the evaluation of the projectivity index of the acquired audio sequence. This index allows the projectivity of the audio sequence to be classified. The audio sequence is then labeled with the projectivity class and the characteristics of the microphones that acquired all the audio sequences.

[0243] One advantage is that it allows for the direct classification of audio sequences, taking into account all the intrinsic characteristics of the individual's voice, and not just the recorded volume. The invention not only reduces the number of microphones needed to estimate projectivity but also allows for the exploitation of all the characteristics of the recorded audio sequence, such as pitch, timbre, etc.

[0244] When the model is learned, it can be used in a recording setup with at least one microphone whose layout configuration is known and / or whose microphone type is known.

[0245] In a first example, an acquisition setup consists of a single microphone whose type is known, and therefore its position relative to the speaker's mouth can be estimated. For example, it could be a reference microphone. The trained model then takes as input the audio sequence acquired by the microphone and is trained to estimate a class of the projectivity classification. This class can be a value on a scale of projectivity index values.

[0246] In a second example, an acquisition setup includes two microphones whose respective types are known, and therefore their respective positions relative to the speaker's mouth can be estimated. The type of microphones can be determined based on a reference microphone, a projection microphone, a harmonic measurement microphone, a laterality microphone, etc.

[0247] The trained model then considers as input the audio sequences acquired by each microphone and is trained to estimate a projectivity class from the two acquisitions. This example yields a better result and a higher confidence score for the classification. Indeed, projectivity depends primarily on the variation in the spatial volume of the sound emitted by an individual. Consequently, the presence of two microphones in two different positions allows for a more accurate estimation.

[0248] However, it should be noted that the use of the trained learning model can be effective even with the acquisition of the input signal(s) by a single microphone because:

[0249] - on the one hand, for each recording corresponding to the training data, the configuration of each microphone used during training is associated with the acquired audio sequences and;

[0250] - on the other hand, the invention makes it possible to exploit all the spectral characteristics of a recorded audio sequence and not only the variation of the sound level as is the case in current solutions.

[0251] "Sound projection" refers to the ability of a voice to fill a space in front of an individual in a balanced manner. As a non-limiting example, a voice exhibits sufficient sound projection when its amplitude at a given distance from the mouth varies by less than 10 dB SPL per steradian in an area in front of the vertical plane containing the individual, within a space conforming to ISO 3382-1 for measurement rooms. This vocal descriptor corresponds to the voice's capacity to interact with a space, and can be expressed in several values ​​depending on the characteristics of different spaces.Thus, the ambient projectivity value determined from the recording can correspond to the voice in combination with the recording space, and can be modified to correspond to the voice in combination with another space, by adopting a weighting that represents a relationship between the recording space and that other space. For example, the value determined from a studio recording corresponds to the voice in combination with the studio, but by replacing the weighting associated with the studio with a weighting associated with a concert hall (or by adopting a weighting that represents the relationship between the studio and the concert hall), it is possible to determine an ambient projectivity value for that same voice in the concert hall.

[0252] Laterality refers to the ability of the voice to disperse in a horizontal plane relative to the subject's mouth.

[0253] The term "attack" refers to the beginning of phonation. The attack can be "soft" when the phonation begins gradually, possibly accompanied by a hissing sound; it can be "hard" when the phonation begins rapidly, possibly accompanied by a consonant; it can be "glottal" when the phonation is preceded by a subglottal tension.

[0254] The term "termination" refers to the end of a phonation, which can be assessed in the same way as the attack.

[0255] At the end of the search session, a plurality of vocal performance recordings are obtained, in which each recording is associated with one or more vocal descriptor values ​​of the corresponding artist's voice.

[0256] Another aspect of the invention relates to a method 200 for characterizing at least one physiological descriptor of an individual's voice I.

[0257] In this application, a physiological descriptor of a voice is defined as a characteristic describing an individual's voice from the perspective of that individual's physiology, that is, a characteristic related to one or more organs of the individual's body.

[0258] At least one physiological descriptor characterized by method 200 according to the invention can be selected from a set of physiological descriptors referred to as reference physiological descriptors, including: vocal strain, muscular strain, and glottal stops. These reference physiological descriptors are commonly known to human voice experts, who can assess them for a given individual through a physical examination in a specific context, such as vocal coaching, a clinical voice examination, or any other application utilizing certain descriptors in particular.

[0259] Optionally, the system includes a module for detecting muscular and / or vocal strain, and / or a module for analyzing glottal attacks.

[0260] Glottal pressure corresponds to an increase in subglottic pressure followed by a rapid release of the glottis causing an initial vibration of the vocal cords presenting a sudden variation in the amplitude and / or spectrum of the voice.

[0261] As a non-limiting example, a glottal stop can be detected non-invasively from the speech signal by identifying the onset of phonation in a recorded speech signal. The onset of phonation can correspond to a portion of the signal where the amplitude exceeds a predetermined threshold. For example, such a threshold could be 5 dB SPL. A window of predetermined duration can be defined around the identified onset of phonation, sliding around it. As a non-limiting example, the window's start is 10 milliseconds or more before the onset of phonation and lasts for 50 milliseconds or less, or 30 milliseconds or more. The speech signal within this window can be segmented into segments of one or more predetermined durations. Optionally, a given segment can have a duration of 10 milliseconds or more, or 20 milliseconds or less.The RMS (root-mean-square) energy, fundamental frequency, and / or spectral burst of the signal can be calculated for each segment. By comparing the RMS energy of the segments, it is possible to determine the rate of change in the signal's RMS energy at the beginning of phonation. By comparing the fundamental frequency of the segments, it is possible to detect a change in the fundamental frequency at the beginning of phonation. By comparing the spectral burst of the segments, it is possible to detect a change in the voice's bandwidth.

[0262] Glottal support is present when:

[0263] - the calculated rate of change for the RMS energy is equal to or greater than 1 dB SPL per second;

[0264] - the calculated variation for the fundamental frequency between successive segments is equal to or greater than 10 Hz per second; and / or

[0265] - the spectral burst is wider than a predetermined threshold.

[0266] As an example, the width threshold for the spectral burst can be 100 Hz between two consecutive segments, or a bandwidth of 200 Hz for a given segment.

[0267] As a non-limiting example, muscle strain is detected in a 60-minute recording of a vocal signal, including 50 minutes of vocalization. Muscle strain is present when:

[0268] - a first part of the recording, comprising, starting from an initial vocal attack by the subject during the recording and for a predetermined duration, presents: o a first average shimmer over the predetermined duration, o a first average spectral inclination over the predetermined duration, and / or o a first average variation in the ratio between harmonics and noise over the predetermined duration; and

[0269] - a second part of the recording, ending with a final vocal utterance from the subject during the recording and lasting the same duration as the first part of the recording, exhibits: o a second average shimmer over the predetermined duration, greater than the first average shimmer, o a second average spectral inclination over the predetermined duration, lower than the first average spectral inclination, and / or o a second average variation in the ratio between harmonics and noise over the predetermined duration, greater than the first average variation in the ratio between harmonics and noise; and possibly when

[0270] - a third part of the recording, starting after the start of the first part of the recording and ending before the end of the second part of the recording and lasting the same duration as the first part of the recording, presents: o a third average shimmer over the predetermined duration, greater than the first average shimmer and less than the second average shimmer, o a second average spectral inclination over the predetermined duration, less than the first average spectral inclination and greater than the second average spectral inclination, and / or o a second average variation of the ratio between harmonics and noise over the predetermined duration, greater than the first average variation of the ratio between harmonics and noise and less than the first average variation of the ratio between harmonics and noise.

[0271] In some embodiments, the method 200 for characterizing at least one physiological descriptor of an individual's voice is implemented by a device 30 called pathology device 30 and similar to the descriptor device 10. The pathology device 30 is illustrated in Figure 5, which represents a schematic diagram showing the components of an example of the pathology device 30. The pathology device 30 can be implemented as a single hardware device, for example in the form of a desktop personal computer (PC), a laptop computer, a personal digital assistant (PDA), a smartphone, a smartwatch, a server, a console, or can be implemented on separate interconnected hardware devices interconnected by one or more communication links, with wired and / or wireless segments.The Pathology 30 device can, for example, communicate with one or more cloud computing systems, one or more servers, or remote devices to implement the functions described in this description for the device in question. The Pathology 30 device can also be implemented itself as a cloud computing system.

[0272] As shown in Figure 5, the pathology device 30 comprises a computer, this computer including a memory 31 for storing program instructions that can be loaded into a circuit 32 and adapted to cause a circuit to execute steps of the process 200 illustrated in Figure 5, described below, when the program information is executed by the circuit 32. The memory 31 can also store data and information useful for the execution of the steps of the present invention as described below.

[0273] Circuit 32 could be, for example:

[0274] - a processor or processing unit adapted to interpret instructions in a computer language, the processor or processing unit being able to understand, be associated with, or be attached to a memory containing the instructions, or

[0275] - the combination of a processor / processing unit and a memory, the processor or processing unit being adapted to interpret instructions in a computer language, the memory containing said instructions, or

[0276] - an electronic circuit board in which the steps of the invention are described in silicon, or

[0277] - a programmable electronic chip such as an FPGA (Field-Programmable Gate Array) chip. The memory 31 may include random access memory (RAM), cache memory, non-volatile memory, backup memory (e.g., programmable or flash memory), read-only memory (ROM), a hard disk drive (HDD), a solid-state drive (SSD), or any combination thereof. The ROM of the memory 31 may be configured to store, among other things, an operating system and / or one or more computer program codes for one or more software applications. The RAM of the memory 31 may be used by the circuit 32 for temporary data storage.

[0278] The computer may also include an input interface 33 for receiving input data and an output interface 34 for providing output data. Examples of input and output data will be provided later.

[0279] To facilitate interaction with the computer, a 35-inch screen and a 36-inch keyboard can be provided and connected to the computer circuit.

[0280] The pathology device 30 is configured to implement the computer-implemented method 200 to characterize at least one physiological descriptor of an individual's voice according to the invention.

[0281] In some embodiments, the pathology device 30 is the descriptor device 10.

[0282] Figure 6 is a flowchart representing an example of a set of steps that can be carried out to implement process 200 to characterize at least one physiological descriptor of an individual's voice and will be described below.

[0283] In a step called the REC3 descriptor reception step, the pathology device 30 receives and then stores in memory 31 at least one characteristic value of a voice descriptor of the individual's voice. This descriptor can be calculated by the global function or be directly derived from the execution of the first learned function, depending on the number of microphones used.

[0284] In some embodiments, process 200 uses characteristic values ​​of voice descriptors obtained by process 100 to determine a characteristic value of a voice descriptor for a previously described individual's voice. In other embodiments, at least one characteristic value of a voice descriptor for the individual's voice is obtained by another process or has a different origin.

[0285] In a step called the third implementation step IMP3, circuit 32 of the pathology device 30 implements a second learned function generated from a second machine learning model trained on a second training domain comprising second training data. The second learned function receives as input at least one characteristic value of the voice descriptor, so as to generate as output at least one score characterizing at least one physiological descriptor of the individual's voice.

[0286] In some embodiments, the second machine learning model is the first machine learning model that generated the first learned function.

[0287] In other embodiments, the second machine learning model is another model, different from the first machine learning model.

[0288] In some embodiments, the second machine learning model generating the second learned function is based on a convolutional neural network (CNN). Alternatively, or in combination, other architectures can be used to build the second machine learning model. Thus, in some alternative embodiments, or those that can be combined with the previous embodiments, the second machine learning model can be based on architectures using temporal correlations, such as recurrent neural networks (RNNs, LSTMs, ll-Nets).

[0289] According to one embodiment, associations and / or combinations of characteristic values ​​of analytical descriptors make it possible to associate said analytical descriptors with a physiological descriptor, for example by associating the corresponding score of the physiological descriptor with different conditions on the scores of several analytical descriptors.

[0290] To link these values, data characterizing an individual's pathological descriptors can be labeled to train a machine learning model. For example, if the score of the analytical descriptor describing the amount of air in the voice is greater than 8 / 10 and the score of the analytical descriptor describing pitch is greater than 6 / 10, then a given physiological descriptor can be associated with these two conditions.

[0291] In some embodiments, the process 200 includes a step called the fourth implementation step IMP4, in which the circuit 32 of the pathology device 30 implements a function called the comparison function. The comparison function is configured to compare at least one characteristic value of the voice descriptor received during the REC3 descriptor reception step, so as to generate at least one output score characterizing at least one physiological descriptor of the individual's voice.

[0292] For example, when the voice descriptor is air over voice, the characteristic value can be compared to a threshold value of 80%. If the characteristic value is greater than this threshold value, the score obtained during the fourth stage of IMP4 implementation is 100% voice forcing.

[0293] As a non-limiting example, voice forcing is detected in a recorded voice signal by dividing the signal (or a portion of the signal) into segments of one or more predetermined durations. The duration of a given segment may be equal to or greater than 10 milliseconds and equal to or less than 25 milliseconds. Two consecutive segments may partially overlap, for example, by 5 milliseconds or more, down to 15 milliseconds or less.

[0294] The shimmer can be calculated for each segment by calculating the amplitude variations between successive vibration cycles. As a non-limiting example, this involves calculating variations in the signal envelope.

[0295] The ratio between harmonics and noise can be calculated for each segment—this is the ratio between the periodic and aperiodic components of the speech signal. Optionally, segments can be compared to each other to determine the variation in this ratio.

[0296] The spectral energy of the voice signal can be calculated within a predetermined spectral band of interest for each segment, and within a subband of interest within that band. As a non-limiting example, the spectral band of interest corresponds to a range from 2 kilohertz to 4 kilohertz, inclusive. As a further non-limiting example, a subband of interest for such a band of interest corresponds to a range from 2,500 Hertz to 4,000 Hertz, inclusive.

[0297] The RMS amplitude can be calculated for each segment. Optionally, the average RMS amplitude can be calculated for all segments.

[0298] When the signal portion includes the onset of phonation, the vocal attack can be calculated by measuring the average amplitude of a given segment relative to the segment containing the onset of phonation, relative to the duration between the ends of these two segments.

[0299] Spectral slant can be calculated for each segment - this is a calculation of dB per octave, relative to the fundamental frequency, typically over a range from the fundamental frequency up to 5 kHz (e.g.).

[0300] Vocal straining is present when:

[0301] - the "shimmer" of a given segment exceeds a predetermined minimum threshold;

[0302] - the ratio between the harmonics and the noise of a given segment falls below a predetermined minimum threshold;

[0303] - the variation in the ratio between harmonics and noise for at least two given consecutive segments exceeds a predetermined maximum threshold;

[0304] - the ratio between the spectral energy in the sub-band of interest of a given segment and the spectral energy in the band of interest of the same given segment exceeds a predetermined minimum threshold;

[0305] - the RMS amplitude for a given segment exceeds a predetermined minimum threshold;

[0306] - the average RMS amplitude for at least 5 consecutive given segments exceeds a predetermined minimum threshold;

[0307] - the vocal attack for a given segment exceeds a predetermined maximum threshold; and / or

[0308] - the average spectral inclination for a given segment falls below a predetermined minimum threshold on a predetermined spectral band, for example a spectral band ranging from 20 hertz up to 5 kilohertz, inclusive.

[0309] Optionally, the threshold(s) for shimmer, the ratio between harmonics and noise, the variation of the ratio between harmonics and noise, the spectral energy in the band of interest, the RMS amplitude, the average RMS amplitude, the vocal attack and / or the spectral slant can be derived from:

[0310] - a personal model of the subject, obtained from recordings of the subject that do not include vocal strain, a normative model of comfortable phonation, obtained from recordings that do not include vocal strain, and / or

[0311] - of a model learned through machine learning.

[0312] As a complement or alternative:

[0313] - the minimum threshold for shimmer is 3% or more, 5% or more, or even 15% relative to the minimum amplitude of the segment;

[0314] - the minimum threshold for the ratio between harmonics and noise is 10 dB;

[0315] - the maximum threshold for variation in the ratio between harmonics and noise is 15% over a period of 50 milliseconds or less;

[0316] - the minimum threshold for the ratio between the spectral energy of the sub-band of interest of a given segment and the spectral energy in the band of interest of the same given segment is 50%;

[0317] - the minimum threshold for the RMS amplitude of a given segment is 60 to 65 dB SPL inclusive;

[0318] - the minimum threshold for the average RMS amplitude is 50 to 55 dB SPL over at least 5 consecutive segments and / or over a duration of 50 ms or more and 100 ms or less;

[0319] - the maximum threshold for the vocal attack is 0.2 to 0.4 dB SPL per millisecond inclusive (e.g., 0.3 dB SPL / ms); and / or

[0320] - the minimum threshold for spectral inclination is -10 dB / octave relative to the fundamental frequency.

[0321] In some embodiments, the process 200 includes an additional step. In this additional step, the circuit 32 of the pathology device 20 implements a third learned function generated from a third machine learning model trained on a third training domain comprising third training data. The third learned function receives as input at least one score characterizing at least one physiological descriptor of the individual's voice, generated by the second learned function during the third implementation step IMP3, and outputs at least one percentage value characterizing the contribution of at least one stage of the individual's vocal tract to the physiological descriptor.

[0322] At least one level of the individual's vocal apparatus may be one of: the individual's lung, the individual's larynx, the individual's pharynx, or an individual's articulator.

[0323] According to one embodiment, associations and / or combinations of characteristic values ​​of physiological descriptors make it possible to associate said physiological descriptors with the involvement of a level of the phonatory apparatus.

[0324] Training the second machine learning model and the third machine learning model

[0325] In some embodiments, the second machine learning model and the third machine learning model are trained prior to the implementation of process 200.

[0326] It is recalled that training a machine learning model consists of determining a set of model parameters and / or teaching the model a corresponding learned function from training data, so that the model can then predict a label of new data received.

[0327] Advantageously, the training data is annotated (or labeled, or tagged). Advantageously, the training data is annotated manually by a panel of experts. By expert, we mean a person recognized as having mastered the knowledge of a given field. For example, an expert might be a singing teacher, a sound engineer, an acoustician, or a doctor.

[0328] The second set of training data allows for the training of a second machine learning model to identify and quantify a physiological descriptor of the voice. This training data comprises associations of voice audio files with at least one label defining at least one physiological descriptor. A panel of experts can label certain voices with a quantified physiological descriptor. This labeling can be repeated for different physiological descriptors and different individuals so that the training data can provide sufficient data for all physiological descriptors. The training data includes, for example, heterogeneous voice samples, that is, samples from individuals of different ages, genders, and / or with characteristics that allow for the characterization of a voice sample.

[0329] As a non-limiting example, the second set of training data includes several groups of descriptors, each group comprising one or more vocal descriptors and one or more physiological descriptors. Optionally, a given group of descriptors can be associated with a voice audio recording such that one or more vocal descriptors and one or more physiological descriptors are annotations on the audio recording. For example, an expert panel can annotate the audio recording with one or more physiological descriptors, and this same audio recording can also be associated with one or more vocal descriptors. For example, an expert panel—possibly the same panel of experts for generating one or more physiological descriptors—can annotate the audio recording with at least one of the one or more vocal descriptors.In addition or alternative, at least one or more of the voice descriptors can be generated by the first function learned during the processing of this audio recording.

[0330] The physiological descriptor of an individual's voice can relate, for example and in a non-limiting way, to vocal strain, muscular strain, glottal supports, etc.

[0331] The third training data allows for the training of a third machine learning model capable of identifying and quantifying at least one percentage value characterizing the contribution of at least one level of the vocal tract of the individual in question to the physiological descriptor. This training data comprises associations of voice audio files with at least one label defining at least one percentage value characterizing the contribution of at least one level of the vocal tract of the individual in question to the physiological descriptor. A panel of experts can label certain voices of individuals with a percentage value characterizing the contribution of at least one level of the vocal tract of the individual in question to the physiological descriptor.This labeling process can be repeated for different physiological descriptors and different individuals so that the training data can provide sufficient data for all physiological descriptors. The training data includes, for example, heterogeneous voice samples, i.e., from individuals of different ages, genders, and / or with characteristics that allow for the characterization of a voice sample, as well as their medical history, for example.

[0332] The vocal tract of said individual may refer to a lung of said individual, to the larynx of said individual, to the pharynx of said individual, or to an articulator of said individual. According to one embodiment, the invention makes it possible to quantify the contribution of each articulator to voice production, including:

[0333] ■ the vocal cords;

[0334] ■ the tongue, including the apex, the dorsum, the radix;

[0335] ■ the lips;

[0336] ■ the teeth;

[0337] ■ the alveoli;

[0338] ■ the hard palate;

[0339] ■ the soft palate;

[0340] ■ the uvula;

[0341] ■ the cheeks;

[0342] ■ the lower jaw also called the mandible;

[0343] ■ the nasal cavities.

[0344] Another aspect of the invention relates to a method for predicting a pathology of an individual's voice:

[0345] - the steps of process 200 previously described to characterize at least one physiological descriptor of the individual's voice;

[0346] - the additional step of implementing the third learned function generated from the third machine learning model trained from the third training domain, said third learned function receiving as input at least one score characterizing at least one physiological descriptor of the individual's voice, so as to generate as output at least one percentage value characterizing a contribution of at least one stage of the individual's phonatory apparatus to the physiological descriptor;

[0347] - a prediction stage, of the pathology of the individual's voice based on the contribution of at least one level of the individual's phonatory apparatus.

[0348] The training domain refers to a set of training data. Thus, it is possible to define domains according to age or gender, or for example domains for certain diseases or pathologies, or even according to certain individual profiles, or taking into account certain medical histories.

[0349] As an example, a vocal strain descriptor can have a contribution distributed across different levels of an individual's vocal tract. For instance, a vocal strain of 1 / 10 is considered acceptable and may correspond to less than 70% involvement of the entire laryngeal system. The remaining 30% results from the involvement of lower or higher levels, such as the lungs, pharynx, or articulators. In this case, the vocal cords may contribute less than 10% to the overall contribution of the vocal tract components, which may be acceptable.

[0350] For a vocal strain of 9 / 10, the larynx can be involved up to 90%, and in particular the vocal cords can contribute between 80% and 90% to the physiological descriptor value, resulting in a guttural voice that is harmful to the vocal cords. Thus, the invention allows for the quantification of the proportion of each physiological descriptor using a trained model. This makes it possible to implement corrective actions for this strain, such as appropriate vocal rehabilitation exercises, medication (under medical supervision) to reduce vocal cord inflammation, recommendations for speech therapy, vocal exercises, physical exercises, or even rest, articulator relaxation exercises, etc.

[0351] In this document, the various examples and / or thresholds presented are given for illustrative purposes only and are not exhaustive.

Claims

DEMANDS 1. A computer-implemented method (100) for determining an intrinsic characteristic value (D) of at least a first voice descriptor of an individual's voice (I), said at least a first voice descriptor belonging to a set of reference voice descriptors, the method (100) comprising: ■ reception (REC) of at least one input audio signal (SA1, SA2,.. SAN), each input audio signal (SA1, SA2,. SAN) encoding a common sequence of at least one sound emitted (S) by said voice of said individual, said at least one input audio signal (SA1, SA2,. SAN) being acquired by at least one microphone; ■ implementation (IMP1) of a first learned function generated from a first machine learning model previously trained from a first training domain comprising at least two training audio signals encoding the same sequence called the training sequence and acquired by two different microphones respectively, said first learned function receiving as input said at least one input audio signal (SA1, SA2, SAN), so as to generate as output, for each of the input audio signals, a corresponding value of at least one first voice descriptor relating to the microphone which acquired said voice of said individual (I), ■ determination of an intrinsic characteristic value (D) of the first voice descriptor of an individual's voice (I) from the corresponding value(s) of the first voice descriptor relative to each microphone generated during the implementation step (IMP1).

2. Method (100) according to claim 1 characterized in that when a single input audio signal is acquired by means of a single microphone, the corresponding value of at least one first descriptor generated corresponds to the intrinsic characteristic value of the first descriptor.

3. A method (100) according to claim 1, wherein the at least one input audio signal (SA1, SA2, ... SAN) comprises at least two input audio signals corresponding to recordings acquired by two microphones arranged in two different acquisition positions, said recordings corresponding to audio signals from the common sequence of at least one sound emitted (S) by said voice of said individual (I), the method (100) further comprising: ■ implementation (IMP2) of a function called global function receiving as input each corresponding value of the first voice descriptor of said voice of said individual acquired by each microphone, so as to generate as output, said intrinsic characteristic value (D) of at least one first voice descriptor of said voice of said individual (I).

4. Method (100) according to the preceding claim, comprising, prior to the step of implementing the first learned function: ■ reception (REC2) of configuration information including, for each input audio signal among the at least one input audio signal (SA1, SA2,... SAN), a position information of an acquisition microphone associated with said input audio signal relating to a position of the individual when he emitted the sequence of at least one sound (S), said first learned function also receiving as input all the position information of all the input audio signals.

5. Method (100) according to any one of claims 3 or 4, wherein the overall function is one of the following: the maximum function, the minimum function, an averaging function, a weighted sum of the corresponding values ​​of at least one voice descriptor of each of the input audio signals.

6. A method (100) according to any one of the preceding claims, wherein the method (100) is a method for determining a value intrinsic characteristic of at least two different vocal descriptors of the individual's voice.

7. Method (100) according to any one of the preceding claims, wherein the reference voice descriptors are normalized on the same scale of values.

8. Method (100) according to any one of the preceding claims, wherein the first training domain comprises a set of training audio signals, each training audio signal being pre-labeled with at least one label, the at least one label comprising a value of at least one reference voice descriptor.

9. Method (100) according to any one of the preceding claims, wherein the first function comprises a neural network, such as a convolutional neural network.

10. Method (100) according to claim 9, wherein a training step comprises: ■ A plurality of recordings of a sequence of at least one sound by a plurality of microphones ({Mi}i), each microphone (Mi) being associated with a recording configuration and being arranged in a predefined position relative to a reference recording position; ■ An estimate of the corresponding value of the voice descriptor of said voice of said individual (I) of each microphone; ■ An estimate of the intrinsic characteristic value of a vocal descriptor of said voice; ■ The association of said estimated value of the intrinsic characteristic value of a voice descriptor of said voice with each corresponding value of the voice descriptor of said voice of said individual (I) of each microphone; Learning the neural network by modifying the coefficients of said network based on associations of said values.

11. A method (100) according to any one of the preceding claims, characterized in that at least a first voice descriptor comprises: - a quantification of the sound attack, - a quantification of the pitch of the note, - a quantification of air in the voice, - a quantification of the ending of sound, - a quantification of projectivity, - a quantification of harmonic mixing, - a quantification of a degree of laterality, and / or - a quantification of a degree of vibrato.

12. Method (100) according to any one of the preceding claims, comprising, before the reception step (REC), an acquisition step (ACQ), by an acquisition system (15), of said at least one input audio signal (SAi , SA2, ... SAN).

13. Programmable device configured to implement the method according to any one of claims 1 to 12.

14. System (20) for determining an intrinsic characteristic value of at least a first voice descriptor of an individual's voice and configured to implement the method according to claim 12, comprising: ■ an acquisition system (15) configured to implement the acquisition step (ACQ), comprising at least one microphone, the at least one microphone being configured to simultaneously acquire said common sequence of at least one sound emitted (D) by said voice of said individual (I), so as to generate at least one input audio signal (SA1, SA2, ... SAN); ■ the programmable device according to claim 13.

15. System (20) according to claim 14, wherein at least one microphone of the acquisition system (15) comprises at least one of: ■ a reference microphone (Mo) positioned, in operation, facing and at the height of the mouth of the individual (I) when the latter emits the sequence of at least one sound (S), at a distance between 1 and 40 cm from the mouth of the individual (I), the orientation of the body of the individual from his back towards his torso defining a direction of orientation and a line of orientation; ■ a first harmonic measurement microphone (Mi) and a second harmonic measurement microphone (M2), positioned in front of the individual along the same vertical line containing the reference microphone (Mo) and located at a distance of between 1 cm and 5 m from a vertical plane containing the individual (I), the first harmonic measurement microphone (M1) being located above the height of the reference microphone (Mo) and at a distance of between 5 cm and 60 cm from it, and the second harmonic measurement microphone (M2) being located below the height of the reference microphone (Mo) and at a distance of between 5 cm and 60 cm from it; ■ a first projectivity measurement microphone (M3) and a second projectivity measurement microphone (M4), positioned in a vertical plane in front of the individual, the vertical plane being distant from the vertical plane containing the individual by a distance of between 5 cm and 10 m, the height of the first projectivity measurement microphone (M3) being distant from the height of the individual's mouth by a distance of between 10 cm and 5 m, the height of the second projectivity measurement microphone (M4) being distant from the height of the individual's mouth by a distance of between 10 cm and 5 m; ■ a first laterality measurement microphone (Ms) and a second laterality measurement microphone (Me), positioned in a vertical plane containing the reference microphone (Mo) and perpendicular to the orientation direction, and substantially arranged at the height of the individual's mouth, on either side of the orientation line; ■ a reflectivity measurement microphone (M?) positioned behind the individual along the orientation line and at a distance from the vertical plane containing the individual of between 30 cm and 10 m; ■ a first ambient projectivity measurement microphone (Ms) and a second ambient projectivity measurement microphone (Mg), positioned in a vertical plane in front of the individual, the vertical plane being distant from the vertical plane containing the individual by a distance of between 5 cm and 10 m, the height of the first ambient projectivity measurement microphone (Ms) being distant from the height of the individual's mouth by a distance of between 1 cm and 10 m, the height of the second ambient projectivity measurement microphone (Mg) being distant from the height of the individual's mouth by a distance of between 1 cm and 10 m; ■ a first ambient reflectivity measurement microphone (Mw) and a second ambient reflectivity measurement microphone (Mu), positioned in a vertical plane behind the individual, the vertical plane being distant from the vertical plane containing the individual by a distance of between 5 cm and 10 m, the height of the first ambient projectivity measurement microphone (Ms) being distant from the height of the individual's mouth by a distance of between 1 cm and 5 m, the height of the second ambient projectivity measurement microphone (Mg) being distant from the height of the individual's mouth by a distance of between 1 cm and 5 m cm; ■ a microphone (M12) called a laryngeal pickup microphone, positioned near the throat of individual I.

16. System (20) according to claim 15, in which is configured for driving the first function: ■ at least one reference microphone when at least one first voice descriptor includes a quantification of the sound attack; ■ at least one harmonic measurement microphone and possibly also the reference microphone when at least one first voice descriptor includes a quantification of the note pitch; ■ at least one reference microphone and one projectivity measurement microphone when at least one first voice descriptor includes air quantification on the voice; ■ at least one reference microphone, accompanied by at least one projectivity measurement microphone and / or at least one harmonic measurement microphone when at least one first speech descriptor includes the sound termination, ■ at least two harmonic measurement microphones or, for example, a reference microphone and a harmonic measurement microphone when at least one first speech descriptor includes a harmonic mixture quantization, ■ at least one laterality measurement microphone and possibly also the reference microphone when at least one first speech descriptor includes a quantification of a degree of laterality, ■ at least one projectivity measurement microphone and possibly also the reference microphone when at least one first speech descriptor includes a quantification of a degree of projectivity, ■ at least one laryngeal recording microphone, called a "stethomicrophone" or laryngophone, when at least one first vocal descriptor includes a quantification of a degree of vibrato.

Citation Information

Patent Citations

  • Dynamic text-to-speech provisioning

    US20180122361A1

  • Acoustic Based Speech Analysis Using Deep Learning Models

    US20210118426A1