Synthesis of patient-specific voice models

By synthesizing subject-specific discriminants using reference models and available voice samples, the system addresses the challenge of limited training data for state discrimination, enabling accurate physiological state assessment.

JP7847856B2Active Publication Date: 2026-04-20CORDIO MEDICAL LTD
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
CORDIO MEDICAL LTD
Filing Date
2021-06-07
Publication Date
2026-04-20

AI Technical Summary

Technical Problem

Existing technologies face challenges in generating discriminators to accurately determine the physiological state of a subject, particularly when sufficient training samples for both stable and unstable states are not readily available.

Method used

The system synthesizes subject-specific discriminants using a combination of reference discriminants and voice samples from the subject in a known state, allowing for the generation of a unique discriminator without requiring samples from the other state.

Benefits of technology

This approach enables effective discrimination between stable and unstable physiological states, even with limited data availability, by leveraging existing voice samples to create a personalized model for accurate state assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007847856000001
    Figure 0007847856000001
  • Figure 0007847856000002
    Figure 0007847856000002
  • Figure 0007847856000003
    Figure 0007847856000003
Patent Text Reader

Abstract

The device (40) has a communications interface (26) and a processor (28). The processor is configured to: receive, via the communications interface, a plurality of speech samples (Equation 1) uttered by the subject while in a first state with respect to the disease; and synthesize a subject-specific discriminator using (Equation 1) and at least one reference discriminator that is not subject-specific; the discriminator is subject-specific and is configured to generate an output indicative of a likelihood that the subject is in a second state with respect to the disease in response to one or more test utterances uttered by the subject.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This invention relates particularly to the field of audio signal processing for diagnostic purposes. [Background technology]

[0002] The article "Optimization of Dynamic Programming Algorithms for Speech Recognition" by Sakoe and Chiba, IEEE Transaction 26.2(1978):43-49 (Non-Patent Literature 1) on Acoustics, Speech and Signal Processing, which is incorporated herein by reference, describes optimal dynamic programming (DP) based on time-normalization algorithms for spoken language recognition. First, the general principle of time normalization is given using a time-warping function. Next, two definitions of time-normalized distances, called symmetric and asymmetric forms, are derived from this principle. These two forms are compared with each other through theoretical discussion and experimental studies. The superiority of the symmetric algorithm is established. A technique called slope constraint is introduced to restrict the slope of the warping function, improving the identification between words of different categories.

[0003] Rabiner, Lawrence R., “A Tutorial on Hidden Markov Models and Selected Applications in Speech Recognition,” Proceedings of the IEEE 77.2 (1989): 257-286 (Non-Patent Literature 2), which is incorporated herein by reference, examines the theoretical aspects of types of statistical modeling and describes how they have been applied to selected problems in machine recognition of speech.

[0004] U.S. Patent No. 5,864,810 (Patent Document 1) describes a method and apparatus for automatic speech recognition that uses adaptive data to adapt to a specific speaker and develops a transformation that converts a speaker-independent model into a speaker-adapted model. The speaker-adapted model is then used for speaker recognition, achieving better recognition accuracy than an unadapted model. In further embodiments, the transformation-based adaptive technique is combined with known Bayesian adaptive techniques.

[0005] U.S. Patent No. 9,922,641 (Patent Document 2) describes a method that includes receiving input speech data from a speaker in a first language and estimating a speaker transformation that represents speaker characteristics associated with the input speech data based on a universal speech model. The method also includes accessing a speaker-independent speech model to generate speech data in a second language different from the first language. The method further includes modifying the speaker-independent speech model using the speaker transformation to obtain a speaker-specific speech model and generating speech data in the second language using the speaker-specific speech model. [Prior art documents] [Patent Documents]

[0006] [Patent Document 1] U.S. Patent No. 5,864,810 [Patent Document 2] U.S. Patent No. 9,922,641 [Non-patent literature]

[0007] [Non-Patent Document 1] Sakoe and Chiba, "Optimization of Dynamic Programming Algorithms for Speech Recognition," IEEE Transactions 26.2(1978):43-49 on Acoustics, Speech, and Signal Processing. [Non-Patent Document 2] Rabiner and Lawrence R., "A Tutorial on Hidden Markov Models in Speech Recognition and Selected Applications," Proceedings of the IEEE 77.2 (1989): 257-286. [Overview of the Initiative]

[0008] In some embodiments of the present invention, the device has a communication interface and a processor. The processor receives, via the communication interface, a plurality of voice samples {u m 0}, receive m=1…M; and {u m 0 The system is configured to synthesize a subject-specific discriminant using at least one reference discriminant that is not subject-specific, the discriminant being subject-specific and configured to produce an output indicating the likelihood that the subject is in a second state with respect to the disease in response to one or more test utterances uttered by the subject.

[0009] In some embodiments, the first state is a stable state and the second state is an unstable state.

[0010] In some embodiments, the disease is selected from a group of diseases consisting of congestive heart failure (CHF), coronary heart disease, arrhythmia, chronic obstructive pulmonary disease (COPD), asthma, interstitial lung disease, edema, pleural effusion, Parkinson's disease, and depression.

[0011] In some embodiments, the processor returns a subject-specific speech model θ of a first state for any given speech sample s, which returns a first distance measure indicating a first similarity between s and the speech of the subject in a first state. 0 The steps are to generate a subject-specific speech model θ of the second state and return a second distance measure indicating a second similarity between s and the speech of the subject in the second state. 1 The system is configured to generate a discriminant specific to the subject by synthesizing a discriminant specific to the subject.

[0012] In some embodiments, at least one reference discriminator has K reference discriminators {φk}, k = 1…K, where {φk} is: each first reference speech model that returns each first distance {D k 0 (s)}, and {D k 0 (s)} is the first reference speech model of the first state that indicates the first similarity between s and the reference speech of the first state uttered by one or more other subjects in K groups; each second reference speech model that returns each second distance {D k 1 (s)}, and {D k 1 (s)} is the second reference speech model of the second state that indicates the second similarity between s and the reference speech of the second state uttered by one or more other subjects in K groups; and has, where θ 0 applies a function to {D k 0 (s)} to return a first distance measure, where θ 1 applies a function to {D k 1 (s)} to return a second distance measure.

[0013] In some embodiments, the function, when applied to {D k 0 (s)}, returns a weighted average of {D´ k 0 (s)}, where {D´ k 0 (s)} is a non-decreasing function of D k 0 (s).

[0014] In some embodiments, the weighted average is Σ k for K weights {w k=1 K w k D´ k 0(s) is the case, and it is a voice sample {u m 0 Minimize the sum of each distance measure for}, {u m 0 Each audio sample belonging to} m The distance measure for Σ is k=1 K w k D' k 0 (u m ) based on.

[0015] In some embodiments, at least one reference discriminant is a first distance D that indicates a first similarity between s and a reference sound of a first state. 0 The first state reference speech model returns (s), and the second distance D indicates the second similarity between s and the second state reference speech. 1 The second state has a reference speech model that returns (s).

[0016] In some embodiments, the reference speech model of the first state is obtained by applying a first function to a set of feature vectors V(s) extracted from s. 0 The reference speech model of the second state is obtained by applying the second function to V(s) D 1 (s) is returned, where θ 0 and θ 1 The step of generating {V(u m 0 Use the normalization transformation T that best transforms )} to θ 0 and θ 1 The process includes the step of generating [something].

[0017] In some embodiments, the normalization transformation T is such that, with respect to one constraint, Σ u∈{um0} Minimize Δ(T(V(u)),V(u0)), Here, Δ is a third distance measure between two sets of features, u0 is u∈{u m 0This is a regular utterance for the content of}.

[0018] In some embodiments, Δ is a non-decreasing function of the dynamic time warping (DTW) distance.

[0019] In some embodiments, the normalization transformation T is such that, with respect to one constraint, Σ u∈{um0} Minimize f'0(T(V(u))), where f'0 is a non-decreasing function of the first function.

[0020] In some embodiments, θ 0 This applies the first function to T(V(s)) to return the first distance measure, and θ 1 This applies the second function to T(V(s)) and returns the second distance measure.

[0021] In some embodiments, θ 0 The step of generating the first parameter involves applying a denormalized transformation T' to the first parameter, where the denormalized transformation T' optimally transforms the first parameter of the reference speech model of the first state under one or more predefined constraints, and θ 1 The step of generating is to apply the denormalized transformation T' to the second parameter of the reference speech model of the second state θ 1 The process includes the step of generating [something].

[0022] In some embodiments, the denormalization transformation T' is, under constraints, Σ u∈{um0} T'(D 0 Minimize )(u), T'(D 0 )(s) is the first distance returned by the reference speech model of the first state under the transformation.

[0023] In some embodiments, the reference speech model for a first state includes a first hidden Markov model (HMM) comprising a plurality of first kernels, where the first parameter includes the first kernel parameter of the first kernel; the reference speech model for a second state includes a second HMM comprising a plurality of second kernels, where the second parameter includes the second kernel parameter of the second kernel.

[0024] In some embodiments, the first and second kernels are Gaussian kernels, and T' has: an affine transformation acting on the mean vectors of any one or more Gaussian kernels; and a quadratic transformation acting on the covariance matrices of any one or more Gaussian kernels.

[0025] In some embodiments, a reference speech model for a first state includes a plurality of first reference frames, and a first parameter includes a first reference frame feature of the first reference frames; a reference speech model for a second state includes a plurality of second reference frames, and a second parameter includes a second reference frame feature of the second reference frames.

[0026] In some embodiments, the reference voice of the first state includes a plurality of reference voice samples of the first state emitted by a first subset of other subjects of R people. The reference voice for the second state includes multiple reference voice samples for the second state emitted by a second subset of other subjects, where the processor further: converts each of the other subjects {T r A step in identifying {Φ}, r=1…R, where Tr is a subject for each r-th subject of other subjects under one or more predefined constraints. r This is a normalization transformation that optimally transforms {Φ}, r} is the union of (i) a reference voice sample of the first state uttered by another subject, and (ii) a reference voice sample of the second state uttered by another subject, step; for each r-th subject, T r to {V(Φ rBy applying this to ), it is configured to perform the steps of: calculating the set of modified features; and generating a reference discriminator from the set of modified features.

[0027] In some embodiments, the reference speech model for the first state and the reference speech model for the second state are identical with respect to a first set of parameters, but differ from each other with respect to a second set of parameters, and the processor is θ 0 θ such that it is identical to the reference speech model of the first state with respect to the second set of parameters 0 The processor is configured to generate θ 1 with respect to the first set of parameters θ 0 It is identical to the reference speech model of the second state with respect to the second set of parameters, θ 1 It is configured to generate.

[0028] In some embodiments, the reference speech model for the first state and the reference speech model for the second state each include different Hidden Markov Models (HMMs), each HMM including multiple kernels, each having its own kernel weights, with a first set of parameters including the kernel weights and a second set of parameters including the kernel parameters of the kernels.

[0029] In some embodiments, at least one reference discriminator includes a reference neural network associated with a plurality of parameters, the neural network returns another output for any one or more speech samples indicating the likelihood that the speech sample was emitted in a second state, and the processor processes the speech sample {u m 0 It is configured to synthesize subject-specific discriminators by synthesizing subject-specific neural networks by tuning a subset of parameters to minimize the error of other outputs for a set of input speech samples containing}.

[0030] In some embodiments, the parameters include the weights of multiple neurons, and a subset of the parameters includes a subset of the weights.

[0031] In some embodiments, the reference neural network includes multiple layers, and the subset of weights includes at least some of the weights associated with one of the layers, but not the weights associated with another layer.

[0032] In some embodiments, the layer comprises (i) one or more acoustic layers of the nerve that generate an acoustic layer output in response to an input based on a speech sample, (ii) one or more speech layers of the nerve that generate a speech layer output in response to the output of the acoustic layer, and (iii) one or more discrimination layers of the nerve that generate other outputs in response to the speech layer output, wherein the subset of weights includes at least some of the weights associated with the acoustic and discrimination layers, but does not include the weights associated with the speech layers.

[0033] In some embodiments, a subset of parameters includes speaker identification parameters that identify the speaker of the audio sample.

[0034] In some embodiments, the set of input audio samples further includes one or more audio samples of a second state.

[0035] In some embodiments of the present invention, a plurality of voice samples {u m 0 A method is provided that includes the step of receiving m=1…M. The method further includes {u m 0 The method comprises the step of synthesizing a subject-specific discriminant using a reference discriminant that is not subject-specific and at least one reference discriminant that is not subject-specific, wherein the discriminant is subject-specific and configured to produce an output indicating the likelihood that the subject is in a second state with respect to the disease in response to one or more test utterances uttered by the subject.

[0036] In some embodiments of the present invention, a computer software product is provided which includes a tangible, non-transient, computer-readable medium storing program instructions. When the instructions are read by the processor, the processor receives: a plurality of voice samples {u m 0}, step of receiving m=1…M; and {u m 0 The system is configured to perform the steps of synthesizing a subject-specific discriminant using at least one reference discriminant that is not subject-specific, the discriminant being subject-specific and configured to produce an output indicating the likelihood that the subject is in a second state with respect to the disease in response to one or more test utterances uttered by the subject. [Brief explanation of the drawing]

[0037] The present invention will be better understood from a detailed description of embodiments with reference to the accompanying drawings: [Figure 1] This is a schematic diagram of a system for evaluating the physiological state of a subject according to several embodiments of the present invention. [Figure 2] This is a flowchart of a method for generating a subject-specific voice model according to several embodiments of the present invention. [Figure 3] This is a flowchart of a method for generating a subject-specific voice model according to several embodiments of the present invention. [Figure 4] This is a flowchart of a method for generating a subject-specific voice model according to several embodiments of the present invention. [Figure 5] This is a schematic diagram of a neural network discriminator according to several embodiments of the present invention. [Modes for carrying out the invention]

[0038] (term) In the context of this application, including the claims, a subject is said to be in an “unstable state” with respect to his or her physiological condition (or “disease”) if he or she is suffering from a rapid deterioration of his or her condition. Otherwise, a subject is said to be in a “stable state” with respect to his or her physiological condition.

[0039] In the context of this application, including the claims, “speech model” refers to a computer implementation function configured to map a speech sample to an output that characterizes the sample. For example, given a speech sample spoken by a subject, the speech model may return a distance measure D(s) indicating the similarity between s and a reference speech of the subject or another subject.

[0040] In the context of this application, including the claims, “discriminator” refers to a group of one or more models, typically machine learning models, configured to distinguish between different states. For example, given a set of states such as “stable” and “unstable” for a particular physiological state, the discriminator could, based on a voice sample of a subject, generate an output indicating that the subject is likely to be unstable.

[0041] (overview) In subjects suffering from physiological conditions, it may be desirable to train a discriminator configured to determine whether the subject is stable or unstable based on their voice. However, the challenge is that obtaining a sufficient number of training samples for each condition can be difficult. For example, for a subject who is generally stable, a sufficient number of voice samples uttered in a stable state may be available, but obtaining a sufficient number of voice samples uttered in an unstable state may be difficult. For other subjects, it may be easy to collect a sufficient number of samples in an unstable state (e.g., after the subject is hospitalized), but it may not be possible to collect a sufficient number of samples in a stable state.

[0042] To address this challenge, embodiments of the present invention generate a subject-specific (i.e., configured to discriminate a subject) discriminator from a subject-nonspecific reference discriminator. to To generate a unique discriminator, the processor, While in one state The audio samples were recorded by the subjects. do , change or adapt the reference discriminator do This process involves the subject to This is called the "synthesis" of unique discriminants. Advantageously, the voice samples produced by the subject when they were in a different state are not needed. stomach .

[0043] Using the techniques described herein, discriminants for any appropriate physiological condition such as congestive heart failure (CHF), coronary heart disease, atrial fibrillation, or any other type of arrhythmia, chronic obstructive pulmonary disease (COPD), asthma, interstitial disease, lung disease, pulmonary edema, pleural effusion, Parkinson's disease, or depression can be synthesized.

[0044] (System description) First, refer to Figure 1, a schematic diagram of a system 20 for evaluating the physiological state of a subject 22 according to several embodiments of the present invention.

[0045] System 20 includes an acoustic receiver 32, such as a mobile phone, tablet computer, laptop computer, desktop computer, voice-controlled personal assistant (such as an Amazon Echo® or Google Home® device), or smart device. The acoustic receiver 32 includes circuitry including an acoustic sensor 38 (e.g., a microphone) that converts sound waves into analog electrical signals, and an analog-to-digital (A / D) converter. Typically, the acoustic receiver 32 further includes a storage device such as a solid-state drive or a screen (e.g., a touchscreen), and / or other user interfaces such as a keyboard or speakers. In some embodiments, the acoustic sensor 38 (and optionally the A / D converter 42) belongs to a unit located outside the acoustic receiver 32. For example, the acoustic sensor 38 may belong to a headset that connects to the acoustic receiver 32 via a wired or wireless connection such as Bluetooth.

[0046] System 20 further comprises a server 40, which includes a processor 28, a storage device 30 such as a hard drive or flash drive, and circuitry including a network interface such as a network interface controller (NIC) 26. Server 40 further comprises a screen, keyboard, and / or other appropriate user interface elements. Typically, server 40 is located away from acoustic receiver 32, for example, in a control center, and server 40 and acoustic receiver 32 communicate with each other over a network 24, including a cellular network and / or the internet, via their respective network interfaces.

[0047] System 20 is configured to assess the physiological state of a subject by processing one or more audio signals (also referred to herein as “audio samples”) received from the subject. Typically, the processor 36 of the acoustic receiver 32 and the processor 28 of the server 40 coordinately perform the reception and processing of at least some of the audio samples. For example, when a subject speaks into the receiver 32, the sound waves of the subject’s voice are converted into analog signals by the acoustic sensor 38 and sampled and digitized by the A / D converter 42. The subject’s voice can be sampled at any appropriate rate, such as between 8kHz and 45kHz. The resulting digital audio signal can be received by the processor 36. The processor 36 can then communicate the audio signal to the server 40 via the NIC 34. The processor 28 can then process the audio signal.

[0048] To process the subject's voice signals, the processor 28 uses subject-specific discriminators 44, which are unique to the subject 22 and stored in the memory device 30. Based on each voice signal, the subject-specific discriminators 44 generate an output indicating the likelihood that the subject is in a particular physiological state. For example, the output may indicate, with respect to physiological state, the likelihood that the subject is in a stable state and / or an unstable state. Alternatively or additionally, the output may include a score indicating the degree of likelihood that the subject's state is unstable. The processor 28 is further configured to synthesize the subject-specific discriminators 44 before using them, as will be described in detail below with reference to the following figures.

[0049] In response to the output from the subject-specific discriminator, the processor can generate any appropriate audible or visual output to the subject and / or another person, such as the subject's physician. For example, processor 28 can transmit an output to processor 36, which can then transmit the output to the subject, for example, by displaying a message on the screen of the receiving device 32. If the output indicates a relatively high probability that the subject's condition is unstable, the processor can generate an alert indicating that the subject needs to take medication or see a doctor. Such a warning can be communicated by calling or sending a message (e.g., a text message) to the subject, the subject's physician, and / or a monitoring center. Alternatively or additionally, in response to the output from the discriminator, the processor can control a drug delivery device to adjust the amount of medication administered to the subject.

[0050] In other embodiments, following the synthesis of subject-specific discriminants, processor 28 communicates the subject-specific discriminants to processor 36, which then stores the subject-specific discriminants in a memory device belonging to the receiving device 32. Processor 36 can then use the subject-specific discriminants. As yet another alternative, even the synthesis of subject-specific discriminants may be performed by processor 36. (Notwithstanding the foregoing, the following description in this specification generally assumes that processor 28—hereinafter simply referred to as “processor”—performs the synthesis.)

[0051] In some embodiments, the receiving device 32 includes an analog telephone that does not include an A / D converter or processor. In such embodiments, the receiving device 32 transmits an analog voice signal from a voice sensor 38 to a server 40 via a telephone network. Typically, in a telephone network, voice signals are digitized, communicated digitally, and converted to analog before reaching the server 40. Thus, the server 40 has an A / D converter that converts the input acoustic analog signal received via a suitable telephone-network interface into a digital voice signal. The processor 28 receives the digital voice signal from the A / D converter and then processes the signal as described above. Alternatively, the server 40 can receive the signal from the telephone network before it is converted to analog, and the server does not necessarily need to have an A / D converter.

[0052] As will be further explained below with reference to the following figures, the processor 28 synthesizes a subject-specific discriminator 44 using training voice samples emitted by subject 22 while in a known physiological state. Each of these samples may be received over the network via the interface, or any other suitable communication interface such as a flash drive interface. Similarly, at least one reference discriminator not specific to subject 22, or training samples from other subjects that may be used to generate a reference discriminator, may also be received by the processor 28 via a suitable communication interface.

[0053] The processor 28 may be embodied as a single processor or as a set of processors that are collaboratively networked or clustered. For example, a control center may include a number of interconnected servers, each containing a processor, that collaboratively perform the technologies described herein. In some embodiments, the processor 28 belongs to a virtual machine.

[0054] In some embodiments, the functionality of processor 28 and / or processor 36 may be implemented only in hardware, as described herein, using, for example, one or more fixed function or general purpose integrated circuits, application specific integrated circuits (ASICs) and / or field programmable gate arrays (FPGAs). Alternatively, this functionality may be implemented at least in part in software. For example, processor 28 and / or processor 36 may be embodied as a programmed processor including, for example, a central processing unit (CPU) and / or a graphics processing unit (GPU). Program code and / or data including a software program may be loaded for execution and processing by the CPU and / or GPU. The program code and / or data may be downloaded to the processor in electronic form, for example, via a network. Alternatively or additionally, the program code and / or data may be provided and / or stored on a non-transitory tangible medium such as a magnetic, optical, or electronic memory. Such program code and / or data, when provided to the processor, generate a machine or a dedicated computer configured to perform the tasks described herein.

[0055] (Synthesis of Subject-Specific Identifiers) As described in the above overview, in conventional techniques for generating an identifier for discriminating between two states, usually, a sufficient number of training samples are required for each state. However, depending on the situation, there may be cases where the processor has sufficient training samples for only one of the states. To address such situations, the processor synthesizes subject-specific identifiers.

[0056] To perform this synthesis, the processor first ,disease receives a plurality of voice samples {u 、} emitted by the subject while in a first state (e.g., a stable state) regarding qi, where m = 1…M. Next, {u m 0} m 0} and using at least one reference identifier that is not unique to the subject, the processor synthesizes a subject-specific identifier. Advantageously, the processor synthesizes a subject-specific identifier even though the processor has little or no voice samples uttered by the subject while the subject is in a second state (e.g., an unstable state) with respect to the disease. Unique The identifier can generate an output indicating the likelihood that the subject is in a second state in response to one or more test utterances uttered by the subject.

[0057] (Multi-model identifier) In some embodiments, the subject-specific identifier is a subject-specific voice model θ of the first state 0 and a subject-specific voice model θ of the second state 1 including. For any voice sample s, θ 0 returns a first distance measure indicating the similarity between s and the voice of the subject in the first state, and θ 1 returns a second distance measure indicating the similarity between s and the voice of the subject in the second state. In such embodiments, the subject-specific identifier can generate an output based on a mutual comparison of the two distance measures. For example, assuming the convention that the greater the distance, the lower the similarity, the subject-specific identifier can generate an output indicating that the subject is likely to be in the first state in response to the ratio between the first distance measure and the second distance measure being less than a threshold. Alternatively, the subject-specific identifier can output the likelihood of each of the two states based on the distance measure or simply output the two distance measures.

[0058] Various techniques can be used to synthesize such multi-model identifiers. Examples of such techniques are described herein with reference to FIGS. 1-4.

[0059] (i) First technique Here, refer to FIG. 2, which is a flowchart of a first technique 46 for generating θ 0 and θ 1 according to some embodiments of the present invention.

[0060] Technique 46 begins with a first receiving or generating step 48, where the processor uses a reference discriminant {φ} for K≧1. k}, receive or generate k=1…K. (Note that the processor can receive some of the discriminants while generating others.) See discriminant {φ k {D} includes each first-state reference speech model and each second-state reference speech model unique to the same K groups of one or more other subjects, referred to herein as “reference subjects”. In other words, for any speech sample s, the first-state reference speech model is each first distance {D} that indicates the similarity between s and each first-state reference speech uttered by the K groups. k 0 (s)}, k=1…K returns. On the other hand, the reference speech model of the second state is the second distance {D k 1 (s) returns k=1…K. In some embodiments, each of the reference speech models includes a parametric statistical speech model such as a Hidden Markov Model (HMM).

[0061] Next, in the voice sample reception step 50, the processor receives one or more first state voice samples {u m 0 The processor receives {D}. Next, in the first model generation step 52 of the first state, the processor receives the set of distances {D}. k 0 (s)} is a single transformed distance f({D k 0 Calculate the function "f" to convert to (s)}) and the audio sample {u m 0 Another function of the transformed distance of} is minimized with respect to one or more suitable constraints. Thus, for any speech sample s, the processor creates a subject-specific speech model θ of the first state. 0 The distance measure returned by the function "f" is {D k 0The subject-specific voice model θ of the first state is calculated by applying it to (s)}. 0 Generates.

[0062] For example, a processor has a total Σ with respect to constraints. m=1 M |f({D k 0 (u m )})| q Identify the function "f" that minimizes q ≥ 0. Alternatively, the function "f" is a weighted sum Σ with respect to the constraints. m=1 M β m f({D k 0 (u m )})| q This can be minimized. In such an embodiment, the weight β of each audio sample m This can be a function of sample quality, with higher quality samples being assigned a greater weight. Alternatively or additionally, speech samples whose transformed distance is greater than a given threshold (such as a specific percentile of the transformed distance) are excluded. It is considered a value and therefore assigned a weight of zero.

[0063] Next, in the first model generation step 54 of the second state, the processor performs the same function {D k 1 By applying (s)}, the subject-specific voice model θ of the second state is obtained. 1 In other words, for any speech sample s, the processor generates a subject-specific speech model θ of the second state. 1 The distance measure returned by is f({D k 1 (s)) equals θ 1 Generates.

[0064] In effect, technique 46 uses a speech sample of the subject in a first state to learn how the subject's voice in the first state best approximates a function of the voices of K groups of reference subjects in the first state. The processor then assumes that the same approximation applies to the second state, and θ 0 The function used is θ 1 Make it usable for that purpose as well.

[0065] As a specific example, the function computed in the first model generation step 52 of the first state is {D k 0 When applied to (s), {D' k 0 In some cases, the weighted average of (s) is returned. Here, {D' k 0 (s)} is |D for p≧1 k 0 (s) p {D k 0 It is a non-decreasing function of (s). In other words, it is the subject-specific speech model θ of the first state for any speech sample s. 0 The distance measure returned by is the weights of K {w k}, for k = 1...K, Σ k=1 K w k D' k 0 (s) may be equal to the subject-specific voice model θ of the second state. Similarly, in such embodiments, the subject-specific voice model θ of the second state may be equal to the subject-specific voice model θ of the second state. 1 The distance measure returned by is Σ k=1 K w k D' k 1 (s) may be equal to D'. k 1 (s) is D k 1 It is the same non-decreasing function as (s). In effect, such a function approximates the subject's speech as a weighted average of the speeches of K groups of reference subjects.

[0066] In such an embodiment, in order to compute K weights in the first model generation step 52 of the first state, the processor uses constraints (e.g., Σ k=1 K w k Regarding =1), audio sample {u m 0 We can also minimize the sum of the distance measures for each of the}, where the audio sample {u m 0 Each audio sample belonging to} m The distance scale is the transformed distance: Σ k=1 K w k D' k 0 (u m ) is based on. For example, the processor has a validity constraint, Σ when q ≥ 0 m=1 M |Σ k=1 K w k D' k 0 (u m )| q In some cases, it is necessary to minimize it. ({D' k 0 (s) = |D k 0 (s) p In this embodiment, q is generally equal to 1 / p. As described above, the transformed distance can be weighted, for example, according to the different qualities of the sample.

[0067] In some embodiments, to simplify the subject-specific model, the processor uses weights {w k Relatively low weights, such as weights below a certain percentile and / or a predefined threshold of}, are set to zero. The processor can then rescale the remaining non-zero weights so that their sum equals 1. For example, the processor can set θ 0 The distance measure returned by is D' kmax 0 The maximum weight w maxIn some cases, all weights except for k are set to zero. max is the maximum weight w max This is the index. Therefore, in effect, the voice of a subject can be approximated by the voice of one of the K groups of the reference subject, ignoring the other K-1 groups.

[0068] (ii) Second technique Here, according to several embodiments of the present invention, θ 0 and θ 1 Refer to Figure 3, which is a flowchart of the second technique 56 for generating.

[0069] Technique 56 begins with a second receiving or generating step 58, in which the processor receives or generates a first-state reference speech model and a second-state reference speech model (neither of which is specific to the subject). Similar to each of the first-state reference models in Technique 46 (Figure 2), the first-state reference speech models in Technique 56 represent a first distance D indicating the similarity between any given speech samples. 0 It returns (s), which indicates the similarity between any speech sample s and the reference speech of the first state. Similarly, the reference speech model of the second state in Technique 56, as with each of the reference speech models of the second state in Technique 46, indicates the similarity between s and the reference speech of the second state, the second distance D. 1 It returns (s).

[0070] For example, a reference speech model of the first state is obtained by applying the first function f0 to a set of feature vectors V(s) extracted from speech samples s, D 0 (s) can be returned (that is, D 0 (s) may be equal to f0(V(s)). On the other hand, the reference speech model of the second state is obtained by applying the second function f1 to V(s) D 1 (s) can return (that is, D 1(s) is equal to f1(V(s)). Each of the reference speech models can include a parametric statistical speech model such as a Hidden Markov Model (HMM).

[0071] However, in contrast to technique 46, the two reference models are not necessarily generated from the reference voices of subjects in the same group. For example, the reference voice model for the first state may be generated from the reference voice of the first state of one or more subjects in one group, and the reference voice model for the second state may be generated from the reference voice of the second state of another group of one or more subjects. Alternatively, one or both models may be generated from artificial voices produced by a speech synthesizer. Thus, technique 56 differs from technique 46, as will be explained in detail below.

[0072] Following the execution of the second receiving or generating step 58, the processor receives the audio sample {u m 0 The processor receives}. Next, in some embodiments, in the transformation calculation step 60, the processor receives the feature vector {V(u m 0 We compute a transformation T that optimally transforms the )}. T can be called a “feature normalization” transformation in that it transforms the features of the subject’s voice sample to neutralize the specificity of the subject’s vocal tract, i.e., T renders the voice sample more general or standard.

[0073] For example, T is related to the constraint: Σ uΕ{um0} f'0(T(V(u))) We can minimize |f0(*)| where f'0 is a non-decreasing function of f0. (For example, f'0(*) is |f0(*)| when p≧1.) p (This may be equal to:) Or, T is under one or more predefined effectiveness constraints: Σ uΕ{um0} Δ(T(V(u)),V(u0)) We can minimize this, where Δ is the distance measure between any two sets of feature vectors, and u0 is {u m 0 For each sample u belonging to}, this is a canonical utterance of the content of u, such as a synthesized speech of the content. In some embodiments, Δ is a non-decreasing function of the dynamic time warping (DTW) distance, which can be calculated as described in the reference by Sakoe and Chiba cited in the background art (Non-Patent Literature 1), incorporated herein by reference. For example, Δ(T(V(u)),V(u0)) is: |DTW(T(V(u)),V(u0))| p This may be equal to the following: Here, DTW(V1,V2) is the DTW distance between two sets of feature vectors V1 and V2, and p≧1.

[0074] (Typically, the DTW distance between two sets of feature vectors is calculated by mapping each feature vector from one set to its corresponding feature vector from the other set, minimizing the sum of the local distances between each pair of features. The local distance between each pair of vectors can be calculated by summing the squared differences between the corresponding elements of the vectors, or by using another suitable function.)

[0075] Typically, the processor extracts N duplicate or non-duplicate frames from each received audio sample s, where N is a function of the predefined length of each frame. Thus, V(s) contains N feature vectors {v}, with one feature vector for each frame. n}, n=1…N, is included. (Each feature vector can include, for example, a set of cepstrum coefficients and / or a set of linear predictor coefficients for the frame.) Typically, T includes a transformation that acts independently on each feature vector, i.e., T(V(s)) = {T(v n )}, n=1…N. For example, T can include affine transformations that operate independently on each feature vector. That is, T(V(s)) is {Av nThis is equivalent to { + b}, n=1...N, where A is an LxL matrix, b is an Lx1 vector, and L is each vector v n It is the length.

[0076] Following the calculation of T, the processor, in the second model generation step 62 of the first state, generates a subject-specific speech model θ of the first state for any speech sample s. 0 This returns f0(T(V(s))). Similarly, in the second model generation step 64 of the second state, the processor generates the subject-specific speech model θ of the second state. 1 θ such that f1(T(V(s))) returns 1 Generates.

[0077] In other embodiments, instead of calculating T, the processor calculates an alternative transformation T' in the alternative transformation calculation step 66 that optimally transforms the parameters of the first state reference speech model under one or more predefined constraints. For example, under the constraints: Σ uΕ{um0} T'(D 0 )(u) Minimize the expression, and here T'(D 0 )(s) is the distance returned by the reference speech model of the first state under the transformation. Alternatively, following the calculation of T, the processor derives T' from T, and applying T' to the model parameters has the same effect as applying T to the features of the subject's speech sample. T' is sometimes called a “parameter denormalization” transformation, because T' transforms the parameters of the reference model to better match the specificity of the subject's vocal tract. That is, T' makes the reference model more specific to the subject.

[0078] In such an embodiment, after calculating T', the processor applies T' to the parameters of the reference speech model of the first state in the third model generation step 68 of the first state to generate the subject-specific speech model θ of the first state. 0Similarly, in the third model generation step 70 of the second state, the processor generates the subject-specific speech model θ of the second state by applying T' to the parameters of the reference speech model of the second state. 1 It generates θ for any audio sample s. In other words, the processor generates θ for any audio sample s. 0 but: T'(D 0 )(s) = f'0(V(s)) θ to return 0 This generates f'0, which is different from f'0 because we used the parameters of the reference speech model of the first state modified by T'. Similarly, the processor generates θ 1 but: T'(D 1 )(s) = f'1(V(s)) θ to return 1 This generates f'1 is different from f1 because we used the parameters of the reference speech model of the second state modified by T'. (In embodiments where T' is derived from T as described above, f'0(V(s))=f0(T(V(s)) and f'1(V(s)) = f1(T(V(s))).

[0079] For example, if each of the reference speech models includes an HMM containing multiple kernels, each subject-specific model can input T(V(s)) into the kernel of the corresponding reference speech model, according to the previous embodiment. Alternatively, in the latter embodiment, the kernel parameters can be transformed using T', and V(s) can then be input into the transformed kernel.

[0080] As a specific example, each reference HMM can contain multiple Gaussian kernels for each state, and each kernel is: g(v;μ,σ)=(1 / √(2π|σ|))*e -(v-μ)T σ-1 (v-μ) The form is as follows, where v is an arbitrary feature vector belonging to V(s), μ is the mean vector, and σ is the covariance matrix with determinant |σ|. For example, assuming a state x with J kernels, the local distance between v and x is: L(Σ j=1 J w x,j g(v;μ x,j ,σ x,j ) It is calculated as follows, where g(v;μ x,j ,σ x,j ) is the j-th Gaussian kernel belonging to state x in j=1...J, and w x,j μ is the weight of this kernel, and L is a suitable scalar function such as the identity function or a negative logarithm function. In this case, T' can include affine transformations acting on one or more mean vectors of the kernel and quadratic transformations acting on one or more covariance matrices of the kernel. In other words, T' is μ' = A -1 Let σ be (μ+b) and σ' = A -1 σA T The Gaussian kernel can be transformed by replacing it with: For example, each local distance is: L(Σ j=1 J w x,j g(v;μ´ x,j ,σ´ x,j ) It is calculated as follows: (In the embodiment in which T' is derived from T as described above: g(v;μ´ x,j ,σ´ x,j ) is g(T(v);μ x,j ,σ x,j ) is equal to , where T(v) = Av + ​​b.

[0081] Alternatively, each of the reference speech models may contain multiple reference frames. In such an embodiment, for each speech sample s, the distance returned by each reference speech model is equal to each feature vector v nThis can be computed by mapping it to one of the reference frames (for example, using dynamic time warping (DTW)). The sum of the local distances between each feature vector and the reference frame to which the feature vector is mapped is minimized. In this case, according to the previous embodiment, each subject-specific model is mapped to the reference frame of the corresponding reference model {T(v) for n=1…N such that the sum of local distances is minimized. n )} can be mapped. Alternatively, according to the latter embodiment, the features of the reference frame can be transformed using T', then {v n} can be mapped to a transformed reference frame for n=1…N.

[0082] Regardless of whether T is applied to the subject's voice sample or T' is applied to the reference model, it is generally advantageous for the reference model to be as standard or subject-independent as possible. Therefore, in some embodiments, particularly when the reference voices used to generate the reference model are from a relatively small number of other subjects, the processor normalizes the reference voices in the receiving or generating step 58 before generating the reference model.

[0083] For example, the processor can first receive a reference voice sample of the first state, emitted by a first subset of the other R subjects, along with a reference voice sample of the second state, emitted by a second subset of the other subjects. (The subsets may overlap; that is, at least one of the other subjects may provide both the reference voice sample of the first state and the reference voice sample of the second state.) The processor then receives {Φ} for each of the r-th other subjects. r Identify {T}, which is the union of (i) a reference voice sample of the first state uttered by the r-th other subject, and (ii) a reference voice sample of the second state uttered by the r-th other subject. The processor then performs the respective transformations {T} for each of the other subjects. r}, r=1...R can be identified. rUnder the above constraints, {Φ r This is another normalization transformation that optimally transforms}. For example, T r Under the defined validity constraints: Σ uΕ{Φr} Δ(T(V(Φ)),V(Φ0)) This can be minimized. Φ0 is the regular (e.g., synthesized) speech of the content of Φ. The processor then processes T for each r-th subject of the other subjects. r to {V(Φ r By applying this to ), the modified set of features can be computed. Finally, the processor can generate a reference discriminator from the modified set of features that includes both reference models.

[0084] (ii) The third technique Here, according to some embodiments of the present invention, the subject-specific voice model θ of the first state is presented. 0 and subject-specific voice model θ of the second state 1 Refer to Figure 4, which is a flowchart of the third technique 72 for generating.

[0085] Similar to technique 56 (Figure 3), technique 72 can handle the case where the reference voice for the first state and the reference voice for the second state originate from different subject groups. Technique 72 requires that the two reference models are identical with respect to a first set of parameters, but different with respect to a second set of parameters that are assumed to represent the effect of the subject's health status on the reference voice. Since this effect is assumed to be the same for subject 22 (Figure 1), technique 72 sets θ such that the two reference models are identical with respect to their corresponding reference models with respect to the second set of parameters. 0 and θ 1 It generates a result, but with respect to the first set of parameters, it is different.

[0086] Technique 72 begins with a third receiving or generating step 74. In this step, the processor receives or generates a reference speech model for a first state and a reference speech model for a second state, the two models being identical with respect to a first set of parameters but different with respect to a second set.

[0087] For example, the processor may first receive or generate a reference model of the first state. The processor may then adapt the reference speech model of the second state to the reference speech model of the first state by changing a second set of parameters (without changing the first set of parameters), so that the sum of each distance returned by the second state model for the reference speech samples of the second state is minimized with respect to appropriate validity constraints. (Any appropriate non-decreasing function, such as the absolute value raised to the power of q ≥ 1, can be applied to each of the distances in this sum.) Alternatively, the processor may first receive or generate a reference model of the second state, and then adapt the reference model of the first state to the reference model of the second state.

[0088] In some embodiments, the reference models include different HMMs, each containing multiple kernels, each with its own kernel weights. In such embodiments, the first set of parameters may include the kernel weights. In other words, two reference models may include identical states, and in each state, they may contain the same number of kernels with the same kernel weights. The first set of parameters may further include state transition distances or probabilities. If the reference models differ in this respect, the second set of parameters may include kernel parameters (e.g., mean and covariance).

[0089] For example, in the first state reference model, the local distance between any state x and any feature vector v is: L(Σ j=1 J w x,j g(v;μ 0 x,j ,σ 0x,j ) is possible. The reference model of the second state may contain the same states as the reference model of the first state, and the local distance to any state x is: L(Σ j=1 J w x,j g(v;μ 1 x,j ,σ 1 x,j ) is possible.

[0090] Following the third receiving or generating step 74, the processor receives the audio sample {u m 0 The processor receives}. Next, in the fourth model generation step 76 of the first state, the processor generates the subject-specific speech model θ of the first state. 0 Generate θ 0 θ is the same as the reference speech model of the first state with respect to a second set of parameters. To perform this adaptation of the reference model of the first state, the processor can use an algorithm similar to the Baum-Welch algorithm, which is described, for example, in L. Rabiner and BH. Juang, "Fundamentals of Speech Recognition," Prentice Hall, 1993, and is incorporated herein by reference. In detail, the processor sets θ to have the parameters of the reference model of the first state. 0 First, it can be initialized. Next, the processor will create an audio sample {u m 0 Each feature vector of} is θ 0 Each of these can be mapped to a state. The processor can then recalculate a first set of state parameters for each state using the feature vector mapped to that state. The processor can then remap the feature vector to the state. This process can be repeated until convergence occurs, i.e., until the mapping no longer changes.

[0091] Following the fourth model generation step 76 for the first state, the processor, in the fourth model generation step 78 for the second state, generates the subject-specific speech model θ for the second state. 1 The subject-specific voice model θ for the first state with respect to the first set of parameters 0 θ is identical to and identical to the reference speech model of the second state with respect to the second set of parameters. 1 Generates.

[0092] (Neural network discriminators) In another embodiment, the processor synthesizes a subject-specific neural network discriminator rather than a multi-model discriminator. Specifically, the processor first receives or generates a reference discriminator containing a neural network associated with several parameters. Subsequently, the processor adjusts some of these parameters, as described below, thereby fitting the network to subject 22 (Figure 1).

[0093] For further details regarding this technique, refer to Figure 5, which is a schematic diagram of a neural network discriminator according to several embodiments of the present invention.

[0094] Figure 5 illustrates how the reference neural network 80 is adapted to a specific subject. The reference neural network 80 is configured to receive speech-related inputs 82 based on one or more speech samples uttered by the subject. For example, the neural network may receive the speech sample itself and / or features such as Mel-frequency cepstrum coefficients (MFCCs) extracted from the sample. The reference neural network 80 may further receive text inputs 90, for example, which include instructions for the speech content of the speech sample. (The speech content can be predetermined or confirmed from the speech sample using speech recognition technology.) For example, if the neural network is trained on N different utterances numbered sequentially 0, ..., N-1, the text inputs 90 may be a permutation of bits indicating the serial numbers of the utterances uttered in the speech sample.

[0095] Given the aforementioned input, the neural network returns output 92 indicating the likelihood that the speech sample was uttered in a second state. For example, output 92 may explicitly include the likelihood that the speech sample was uttered in a second state. Alternatively, the output may explicitly include the likelihood that the speech sample was uttered in a first state, such that the output implicitly indicates the likelihood of the former. For example, if the output indicates a 30% probability of the first state, the output may effectively indicate a 70% probability of the second state. Yet another alternative is that the output may include scores for each of the two states, from which the likelihood of both can be calculated.

[0096] Typically, the reference neural network 80 includes multiple layers of neurons. For example, in an embodiment where the speech-related input 82 includes a raw speech sample (rather than features extracted therefrom), the neural network may include one or more acoustic layers 84 that generate acoustic layer outputs 83 in response to the speech-related response. In effect, the acoustic layers 84 extract feature vectors from the input speech sample by performing an acoustic analysis of the speech sample.

[0097] As another example, the neural network may include one or more speech layers 86 that generate a speech layer output 85 in response to an acoustic layer output 83 (or in response to similar features contained in a speech-related input 82). For example, a speech layer 86 may match the acoustic features of a speech sample specified by the acoustic layer output 83 with the expected speech content of the speech sample indicated by a text input 90. Alternatively, the network may be configured for a single pre-configured text, and therefore the speech layers 86 and text input 90 can be omitted.

[0098] As yet another example, the neural network may include one or more discrimination layers 88 that produce an output 92 in response to a speech layer output 85 (and optionally an acoustic layer output 83). The discrimination layer 88 may include, for example, one or more layers of neurons that compute features to distinguish between a first health state and a second health state, and an output layer that produces an output 92 based on these features. The output layer may include, for example, a first state output neuron that outputs a score indicating the likelihood of the first state, and a second state output neuron that outputs a score indicating the likelihood of the second state.

[0099] In some embodiments, the reference neural network 80 is a deep learning network in that the network incorporates a relatively large number of layers. Alternatively or additionally, the network may include special elements such as convolutional layers, skip layers, and / or recurrent neural network elements. The neurons within the neural network 80 can be associated with various types of activation functions.

[0100] To synthesize subject-specific neural network discriminators, the processor adjusts a subset of parameters associated with the reference neural network 80 to create a voice sample {u m 0 Minimize the error in output 92 for a series of input audio samples containing {u}. In other words, the processor minimizes the error in audio samples {u}. m 0 The input is given as an option, along with one or more audio samples uttered by the subject or another subject in a second state, and a subset of parameters is adjusted so that the error in output 92 is minimized.

[0101] For example, a processor can adjust some or all of the nerve weights of each nerve belonging to a neural network. In a specific example, a processor can adjust at least some of the weights associated with one neural layer without adjusting the weights associated with another layer of the neural network. For example, as shown in Figure 5, a processor can adjust the weights associated with the acoustic layer 84 and / or the discrimination layer 88, which are assumed to be subject-dependent, but not the weights associated with the speech layer 86.

[0102] In some embodiments, the neural network is associated with a speaker identification (or “subject ID”) parameter 94 that identifies the speaker of the speech sample used to generate the speech-related input 82. The speech refers to a subject used to train a reference neural network 80, and parameter 94 may include a sequence of R numbers. For each input 82 obtained from one of these subjects, the subject's serial number can be set to 1 in parameter 94, and other digits can be set to 0. Parameter 94 can be input to the acoustic layer 84, the speech layer 86, and / or the discrimination layer 88.

[0103] In such embodiments, the processor can adjust parameter 94 instead of, or in addition to, adjusting the nerve weights. By adjusting parameter 94, the processor can effectively approximate the subject's voice as a combination of the voices of some or all of the reference subjects. As a purely illustrative example, when R=10, the processor can adjust parameter 94 to the value [0.5 0 0 0 0.3 0 0 0 0.2 0]. This demonstrates that the processor can efficiently approximate the subject's voice as a combination of the voices of the 1st, 5th, and 9th reference subjects. (Thus, parameter 94 becomes associated with the network not simply by being a variable input to the network, but by being a fixed parameter of the network.)

[0104] To tune the parameters, the processor can use any suitable technique known in the art. One such technique is backpropagation, which repeatedly subtracts a vector of values ​​from the parameters that are multiples of the gradient of a deviation function with respect to the parameters. The deviation function quantifies the deviation between the output and the expected output of the network. Backpropagation can be performed for each sample in the set of input speech samples (multiple iterations for each sample are optional) until a suitable degree of convergence is reached.

[0105] Those skilled in the art will understand that the present invention is not limited to those specifically shown and described above. The scope of embodiments of the present invention includes both combinations and subcombinations of the various features described above, as well as both variations and modifications thereof not found in the prior art, which will be recalled by those skilled in the art who have read the foregoing. For example, the scope of embodiments of the present invention includes the synthesis of a single-model subject-specific discriminator, such as a neural network discriminator, from a reference discriminator including a first-state reference speech model and a second-state reference speech model.

[0106] Documents incorporated into this patent application by reference are considered an integral part of this application. In the event of any conflict between definitions made expressly or implicitly herein and definitions in these incorporated documents, only the definitions herein should be considered.

Claims

1. Communication interface; and Processor: A device having, The aforementioned processor is: Multiple voice samples {u} emitted by the subject while in a first state with respect to the disease are transmitted via the aforementioned communication interface. m 0 }, receive m=1...M, and Synthesizing subject-specific discriminators that are unique to the subject and are configured to produce an output indicating the likelihood that the subject is in a second state with respect to the disease, in response to one or more test utterances uttered by the subject. It is configured in such a way, The step of synthesizing the subject-specific discriminant is: A reference discriminant configured to discriminate between the first state and the second state, but not specific to the subject, from at least one reference discriminant, the voice sample {u m 0 The step of setting one or more parameters to fit the at least one reference discriminator to the subject, using} and without using any other voice samples emitted by the subject while the subject is in the second state, A device characterized by the following features.

2. The apparatus according to claim 1, characterized in that the first state is a stable state and the second state is an unstable state.

3. The apparatus according to claim 1, characterized in that the disease is selected from the group of diseases consisting of congestive heart failure (CHF), coronary heart disease, arrhythmia, chronic obstructive pulmonary disease (COPD), asthma, interstitial lung disease, edema, pleural effusion, Parkinson's disease, and depression.

4. The aforementioned processor is: A subject-specific voice model θ for a first state returns a first distance measure indicating a first similarity between an arbitrary voice sample s and the voice of the subject in the first state. 0 The steps to generate; and A second distance scale is provided for the subject in the second state, which returns a subject-specific voice model θ that indicates a second similarity between the voice sample s and the subject's voice in the second state. 1 The steps to generate; The subject-specific discriminant is synthesized by this method. The apparatus according to claim 1, characterized in that it is configured as follows.

5. K groups of one or more other subjects each emit K subsets of the first state of the reference audio sample of the first state. Each of the K groups emits K subsets of the second state of the reference audio sample of the second state. The at least one reference discriminant has K reference discriminants {φk}, k = 1...K, where the reference discriminants {φk} are: Each first distance {D k 0 (s)} is a reference speech model of each first state that returns the first distance {D k 0 (s)}, and the first distance {D k 1 (s)} indicates a first similarity between the speech sample s and a subset of each of the first states. A reference speech model of the first state; The second distance for each {D} k 1 A reference speech model of each second state that returns {s}, wherein the second distance {D} k 1 (s)} is a reference speech model of the second state that shows a second similarity between the speech sample s and each subset of the second state; It has, Here, the subject-specific voice model θ of the first state is shown. 0 is {D k 0 Apply the function to (s) to return the first distance measure, Here, the subject-specific voice model θ of the second state is shown. 1 is {D k 1 Apply the function to (s) to return the second distance measure. The apparatus according to feature 4.

6. The function is the first distance {D k 0 When applied to (s), {D'} k 0 Return the weighted average of (s), where D' k 0 (s) is the first distance D k 0 The apparatus according to claim 5, characterized in that (s) is a non-decreasing function.

7. The aforementioned weighted average is calculated using K weights {w k For}, k = 1...K Σ k=1 K lol k D' k 0 (s) and it is the aforementioned audio sample {u m 0 Minimize the sum of each distance scale for} and the audio sample {u m 0 Each audio sample belonging to} m The distance measure for Σ is k=1 K lol k D' k 0 (u m Based on, The apparatus according to feature 6.

8. The at least one reference discriminant is, A first distance D indicates a first similarity between the aforementioned audio sample s and the reference audio sample of the first state. 0 The first state reference speech model returns (s), and The second distance D indicates the second similarity between the aforementioned audio sample s and the reference audio sample of the second state. 1 The reference speech model of the second state that returns (s), The apparatus according to claim 4, characterized by having the following features.

9. The reference speech model of the first state is obtained by applying a first function to a series of feature vectors V(s) extracted from the speech sample s, thereby determining the first distance D 0 (s) returns, The reference speech model of the second state is obtained by applying the second function to the feature vector V(s) to obtain the second distance D 1 (s) returns, Here, the subject-specific voice model θ for the first state 0 and subject-specific voice model θ of the second state 1 The step of generating the feature vector {V(u)} under one or more predefined constraints m 0 Using the normalization transformation T that optimally transforms )}, the θ 0 and the θ 1 The step of generating The apparatus according to feature 8.

10. The aforementioned normalization transformation T has one constraint, Σ u∈{um0} Δ(T(V(u)), V(u 0 Minimize )) and Here, Δ is a third distance measure between two sets of features, u 0 is u ∈ {u m 0 This is a regular utterance for the content of}. The apparatus according to feature 9.

11. The apparatus according to claim 10, characterized in that the third distance measure Δ is a non-decreasing function of the dynamic time warping (DTW) distance.

12. The aforementioned normalization transformation T has one constraint, Σ u∈{um0} f' 0 Minimize (T(V(u))) f' 0 is a non-decreasing function of the first function, The apparatus according to feature 9.

13. The subject-specific voice model θ in the first state 0 This applies the first function to the normalization transformation T(V(s)) to return the first distance measure, The subject-specific voice model θ for the second state described above 1 This applies the second function to the normalization transformation T(V(s)) to return the second distance measure. The apparatus according to feature 9.

14. The subject-specific voice model θ in the first state 0 The step of generating the θ involves applying a denormalization transformation T' to the first parameter, which optimally transforms the first parameter of the reference speech model of the first state under one or more predefined constraints. 0 The process includes the step of generating, Subject-specific voice model θ for the second state described above 1 The step of generating the θ is performed by applying the denormalized transformation T' to the second parameter of the reference speech model of the second state. 1 The step of generating The apparatus according to feature 8.

15. The denormalization transformation T' is, under the predefined constraints, Σ u∈{um0} T'(D 0 Minimize (u), where T'(D 0 )(s) is the first distance returned by the reference speech model of the first state under the denormalization transformation, The apparatus according to feature 14.

16. The reference speech model of the first state includes a first hidden Markov model (HMM) comprising a plurality of first kernels, and the first parameters include the first kernel parameters of the first kernel. The reference speech model of the second state includes a second HMM comprising a plurality of second kernels, and the second parameters include the second kernel parameters of the second kernel. The apparatus according to feature 14.

17. The first kernel and the second kernel are Gaussian kernels. The aforementioned denormalization transformation T' is: Affine transformations acting on the mean vector of any one or more Gaussian kernels; and A quadratic transformation that acts on the covariance matrix of any one or more Gaussian kernels; The apparatus according to claim 16, characterized by having the following features.

18. The first state reference speech model includes a plurality of first reference frames, and the first parameter includes the first reference frame features of the first reference frame. The reference speech model of the second state includes a plurality of second reference frames, and the second parameter includes the second reference frame features of the second reference frame. The apparatus according to feature 14.

19. The reference audio sample of the first state is emitted by a first subset of other R subjects. The reference audio sample of the second state is emitted by a second subset of the other subjects, Here, the processor is: Each of the conversions for the other subjects mentioned above {T r A step in identifying {Φ}, r=1...R, wherein Tr is for each r-th subject of the other subjects under one or more predefined constraints. r This is a normalization transformation that optimally transforms {Φ}. r } is the union of (i) a reference voice sample of the first state uttered by the other subject, and (ii) a reference voice sample of the second state uttered by the other subject, in step; For each of the r-th other subject, the conversion T r The feature vector {V(Φ r The steps involve: applying this to} to compute the set of modified features; and The steps include: generating the reference discriminator from the set of modified features; By executing this, the system is further configured to obtain the aforementioned reference discriminator. The apparatus according to feature 8.

20. The reference speech model for the first state and the reference speech model for the second state are identical with respect to a first set of parameters, and differ from each other with respect to a second set of parameters. The processor provides a subject-specific voice model θ for the first state. 0 The θ is identical to the reference speech model of the first state with respect to the second set of parameters. 0 It is configured to generate, The processor provides a subject-specific voice model θ for the second state. 1 with respect to the first set of parameters θ 0 The θ is identical to the reference speech model of the second state with respect to the second set of parameters. 1 Configured to generate, The apparatus according to feature 8.

21. The first state reference speech model and the second state reference speech model each include different hidden Markov models (HMMs), each of which includes a plurality of kernels, each having its own kernel weights. The first set of parameters includes kernel weights, The second set of parameters includes the kernel parameters of the kernel. The apparatus according to feature 20.

22. The at least one reference discriminator includes a reference neural network associated with a plurality of neural network parameters, the reference neural network returns an output indicating the likelihood that any one or more test speech samples were emitted in the second state, The processor uses the audio sample {u m 0 The subject-specific discriminator is synthesized by adjusting only one subset of the neural network parameters to minimize the error in the output that shows the likelihood for a set of input audio samples including}. The apparatus according to feature 1.

23. The apparatus according to claim 22, wherein the neural network parameters include weights for multiple nerves, and one subset of the neural network parameters includes one subset of the weights.

24. The apparatus according to claim 23, wherein the reference neural network comprises a plurality of layers, and one subset of the weights includes at least some of the weights associated with one of the layers, but does not include the weights associated with another layer.

25. The layer comprises (i) one or more acoustic layers of a nerve that generate an acoustic layer output in response to an input based on the test audio sample, (ii) one or more voice layers of a nerve that generate an audio layer output in response to the acoustic layer output, and (iii) one or more discrimination layers of a nerve that generate an output indicating the likelihood in response to the voice layer output. The one subset of the weights includes at least some of the weights associated with the acoustic layer and the discrimination layer, but does not include the weights associated with the speech layer. The apparatus according to feature 24.

26. The apparatus according to claim 22, characterized in that one subset of the neural network parameters includes a speaker identification parameter that identifies the speaker of the test voice sample.

27. The apparatus according to claim 22, characterized in that the set of input audio samples further comprises one or more input audio samples of the second state.

28. A method performed by a processor in a computer device, the method being: Multiple audio samples uttered by the subject while in the first state with respect to the disease {u m 0 }, the step of receiving m=1...M; and The steps include: synthesizing a subject-specific discriminator that is unique to the subject and configured to produce an output indicating the likelihood that the subject is in a second state with respect to the disease, in response to one or more test utterances uttered by the subject; It has, The step of synthesizing the subject-specific discriminant is: A reference discriminant configured to discriminate between the first state and the second state, but not specific to the subject, from at least one reference discriminant, the voice sample {u m 0 The step of setting one or more parameters to fit the at least one reference discriminator to the subject, using} and without using any other voice samples emitted by the subject while the subject is in the second state, A method characterized by the following:

29. The method according to 28, characterized in that the first state is a stable state and the second state is an unstable state.

30. The method according to 28, characterized in that the disease is selected from the group of diseases consisting of congestive heart failure (CHF), coronary heart disease, arrhythmia, chronic obstructive pulmonary disease (COPD), asthma, interstitial lung disease, edema, pleural effusion, Parkinson's disease, and depression.

31. The step of synthesizing the subject-specific discriminant is: A subject-specific voice model θ for a first state returns a first distance measure indicating a first similarity between an arbitrary voice sample s and the voice of the subject in the first state. 0 The steps to generate; and Return a second distance measure indicating a second similarity between the voice sample s and the voice of the subject in the second state, a voice model θ specific to the subject in the second state 1 generating step; The method according to 28, characterized by having the following:

32. K groups of one or more other subjects each emit K subsets of the first state of the reference audio sample of the first state, Each of the K groups emits K subsets of the second state of the reference audio sample of the second state. The at least one reference discriminant has K reference discriminants {φk}, k = 1...K, where the reference discriminants {φk} are: The first distance between each {D} k 0 A reference speech model of each first state that returns {s}, wherein the first distance {D} k 0 (s)} is a reference speech model of the first state that shows a first similarity between the speech sample s and each subset of the first state; Each second distance {D k 1 (s)} is a reference acoustic model of each second state that returns the second distance {D k 1 (s)} indicates a second similarity between the audio sample s and a subset of each of the second states, a reference acoustic model of the second state; It has, Here, the subject-specific voice model θ of the first state is shown. 0 is {D k 0 Apply the function to (s) to return the first distance measure, Here, the subject-specific voice model θ of the second state is shown. 1 is {D k 1 Apply the function to (s) to return the second distance measure. The method according to feature 31.

33. The function is the first distance {D k 0 When applied to (s), {D'} k 0 The weighted average of (s) is returned, where {D'} k 0 (s) is the first distance D k 0 The method according to 32, characterized in that (s) is a non-decreasing function.

34. The aforementioned weighted average is calculated using K weights {w k For}, k = 1...K Σ k=1 K lol k D' k 0 (s) and it is with respect to one constraint the aforementioned audio sample {u m 0 Minimize the sum of the distance measures for each of the} audio samples {u m 0 Each audio sample belonging to} m The distance measure for Σ is k=1 K lol k D' k 0 (u m Based on, The method according to feature 33.

35. The at least one reference discriminant is, A first distance D indicates a first similarity between the aforementioned audio sample s and the reference audio sample of the first state. 0 The first state reference speech model returns (s), and The second distance D indicates the second similarity between the aforementioned audio sample s and the reference audio sample of the second state. 1 The reference speech model of the second state that returns (s), The method according to 31, characterized by having the following:

36. The reference speech model of the first state is obtained by applying a first function to a series of feature vectors V(s) extracted from the speech sample s, thereby determining the first distance D 0 (s) returns, The reference speech model of the second state is obtained by applying the second function to the feature vector V(s) to obtain the second distance D 1 (s) returns, Here, the subject-specific voice model θ for the first state 0 and subject-specific voice model θ of the second state 1 The step of generating a feature vector {V(u)} under one or more predefined constraints m 0 Using the normalization transformation T that optimally transforms )}, the θ 0 and the θ 1 The step of generating The method according to 35, characterized by the features described above.

37. The aforementioned normalization transformation T has one constraint, Σ u∈{um0} Δ(T(V(u)), V(u 0 Minimize )) and Here, Δ is a third distance measure between two sets of features, u 0 is u ∈ {u m 0 This is a regular utterance for the content of}. The method according to the feature of 36.

38. The method according to 37, characterized in that the third distance measure Δ is a non-decreasing function of the dynamic time warping (DTW) distance.

39. The aforementioned normalization transformation T has one constraint, Σ u∈{um0} f' 0 Minimize (T(V(u))) f'0 is a non-decreasing function of the first function described above. The method according to the feature of 36.

40. Subject-specific voice model θ in the first state 0 This applies the first function to the normalization transformation T(V(s)) to return the first distance measure, The subject-specific voice model θ for the second state described above 1 This applies the second function to the normalization transformation T(V(s)) to return the second distance measure. The method according to the feature of 36.

41. The subject-specific voice model θ in the first state 0 The step of generating the θ involves applying a denormalization transformation T' to the first parameter, which optimally transforms the first parameter of the reference speech model of the first state under one or more predefined constraints. 0 The process includes the step of generating, Subject-specific voice model θ for the second state described above 1 The step of generating the θ is performed by applying the denormalized transformation T' to the second parameter of the reference speech model of the second state. 1 The step of generating The method according to 35, characterized by the features described above.

42. The denormalization transformation T' is, under the predefined constraints, Σ u∈{um0} T'(D 0 Minimize (u), where T'(D 0 )(s) is the first distance returned by the reference speech model of the first state under the denormalization transformation, The method according to feature 41.

43. The reference speech model of the first state includes a first hidden Markov model (HMM) comprising a plurality of first kernels, and the first parameters include the first kernel parameters of the first kernel. The reference speech model of the second state includes a second HMM comprising a plurality of second kernels, and the second parameters include the second kernel parameters of the second kernel. The method according to feature 41.

44. The first kernel and the second kernel are Gaussian kernels. The aforementioned denormalization transformation T' is: Affine transformations acting on the mean vector of any one or more Gaussian kernels; and A quadratic transformation that acts on the covariance matrix of any one or more Gaussian kernels; The method according to 43, characterized by having

45. The first state reference speech model includes a plurality of first reference frames, and the first parameter includes the first reference frame features of the first reference frame. The reference speech model of the second state includes a plurality of second reference frames, and the second parameter includes the second reference frame features of the second reference frame. The method according to feature 41.

46. The method according to claim 35: The reference audio sample of the first state is emitted by a first subset of other R subjects. The reference sound of the second state is emitted by a second subset of other subjects, The aforementioned method further: Each of the conversions for the other subjects mentioned above {T r A step in identifying {Φ}, r=1...R, wherein Tr is for each r-th subject of the other subjects under one or more predefined constraints. r This is a normalization transformation that optimally transforms {Φ}. r } is the union of (i) a reference voice sample of the first state uttered by the other subject, and (ii) a reference voice sample of the second state uttered by the other subject, in step; For each of the r-th other subject, the conversion T r The feature vector {V(Φ r The steps involve: applying this to} to compute the set of modified features; and The steps include: generating the reference discriminator from the set of modified features; Having, The method according to 35, characterized by the features described above.

47. The reference speech model for the first state and the reference speech model for the second state are identical with respect to a first set of parameters, and differ from each other with respect to a second set of parameters. Subject-specific voice model θ in the first state 0 The step of generating θ is 0 The θ is identical to the reference speech model of the first state with respect to the second set of parameters. 0 It is configured to generate, Said θ 1 The step of generating the subject-specific voice model θ of the second state is 1 with respect to the first set of parameters, the θ 0 The θ is identical to the reference speech model of the second state with respect to the second set of parameters. 1 Configured to generate, The method according to 35, characterized by the features described above.

48. The first state reference speech model and the second state reference speech model each include different hidden Markov models (HMMs), each of which includes a plurality of kernels, each having its own kernel weights. The first set of parameters includes kernel weights, The second set of parameters includes the kernel parameters of the kernel. The method according to feature 47.

49. The at least one reference discriminator includes a reference neural network associated with a plurality of neural network parameters, the reference neural network returns an output indicating the likelihood that any one or more test speech samples were emitted in the second state, The step of synthesizing the subject-specific discriminant is performed by the voice sample {u m 0 The step of adjusting only one subset of the neural network parameters to minimize the error in the output indicating the likelihood for a set of input audio samples including} The method according to feature 28.

50. The method according to 49, characterized in that the neural network parameters include weights for multiple nerves, and one subset of the neural network parameters includes one subset of the weights.

51. The method according to 50, characterized in that the reference neural network comprises a plurality of layers, and one subset of the weights includes at least some of the weights associated with one of the layers, but does not include the weights associated with another layer.

52. The layer comprises (i) one or more acoustic layers of a nerve that generate an acoustic layer output in response to an input based on the test audio sample, (ii) one or more voice layers of a nerve that generate an audio layer output in response to the acoustic layer output, and (iii) one or more discrimination layers of a nerve that generate an output indicating the likelihood in response to the voice layer output. The one subset of the weights includes at least some of the weights associated with the acoustic layer and the discrimination layer, but does not include the weights associated with the speech layer. The method according to 51, characterized by...

53. The method according to 49, characterized in that one subset of the neural network parameters includes a speaker identification parameter that identifies the speaker of the test audio sample.

54. The method according to 49, characterized in that the set of input audio samples further includes one or more input audio samples of the second state.

55. A computer software product comprising a tangible, non-transient, computer-readable medium containing program instructions, wherein, when read by a processor, the instructions are communicated to the processor as follows: Multiple audio samples uttered by the subject while in the first state with respect to the disease {u m 0 }, the step of receiving m=1...M; and The steps include: synthesizing a subject-specific discriminator that is unique to the subject and configured to produce an output indicating the likelihood that the subject is in a second state with respect to the disease, in response to one or more test utterances uttered by the subject; Make it run, The step of synthesizing the subject-specific discriminant is: A reference discriminant configured to discriminate between the first state and the second state, but not specific to the subject, from at least one reference discriminant, the voice sample {u m 0 The step of setting one or more parameters to fit the at least one reference discriminator to the subject, using} and without using any other voice samples emitted by the subject while the subject is in the second state, A computer software product characterized by the following features.

Citation Information

Patent Citations

  • Device and method for determining a medical health parameter of a subject using voice analysis

    DE102015218948A1

  • System and method for pulmonary condition monitoring and analysis

    US20200098384A1

  • Design of Stimuli for Symptom Detection

    US20200152226A1

  • Method and apparatus for processing voice data of speech

    US20200168230A1

  • Method and apparatus for speech recognition adapted to an individual speaker

    US5864810A