Synthesize patient-specific speech models

By synthesizing a subject-specific speech model and utilizing hidden Markov models and neural network technology, the problem of lacking training samples was solved, enabling accurate assessment of the subject's physiological state.

CN115768342BActive Publication Date: 2025-10-31CORDIO MEDICAL LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202180045274.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-06-29
Filing Date
2021-06-07
Publication Date
2025-10-31
Estimated Expiration
2041-06-07

AI Technical Summary

Technical Problem

Existing technologies struggle to generate accurate speech discriminators for assessing the stability and instability of subjects' physiological states when there are only a few or no training samples.

Method used

By using a subject-independent reference discriminator, a subject-specific speech model is generated. Then, using techniques such as Hidden Markov Models and Neural Networks, a subject-specific discriminator is synthesized to adapt to the subject's speech characteristics and perform state assessment.

Benefits of technology

This invention enables the generation of a discriminator that can accurately assess the physiological state of a subject even in the absence of unstable speech samples, thereby improving the accuracy and reliability of the assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115768342B_ABST
    Figure CN115768342B_ABST
Patent Text Reader

Abstract

An apparatus (40) includes a communication interface (26) and a processor (28). The processor is configured to: receive via the communication interface a plurality of speech samples (Formula I), m = 1…M, spoken by a subject (22) in a first state concerning a disease; and synthesize a subject-specific discriminator (44) using (Formula I) and at least one non-subject-specific reference discriminator, the subject-specific discriminator being subject-specific and configured to generate an output indicating the probability that the subject is in a second state concerning a disease in response to one or more test utterances by the subject. Other embodiments are also described.
Need to check novelty before this filing date? Find Prior Art

Description

Invention Field

[0001] This invention relates to the field of speech-signal processing, and particularly to its use for diagnostic purposes.

[0002] background

[0003] Sakoe and Chiba's paper, "Dynamic Programming Algorithm Optimization for Spoken Word Recognition," published in IEEE Transactions on Acoustics, Speech, and Signal Processing 26.2 (1978): 43-49, reports a time normalization algorithm based on optimal dynamic programming (DP) for spoken word recognition, which is incorporated herein by reference. First, the general principle of time normalization is given using a time-warping function. Then, two distance definitions for time normalization (referred to as the symmetric and asymmetric forms) are derived from this principle. These two forms are compared through theoretical discussion and experimental studies. The superiority of the symmetric form algorithm is established. A technique called slope constraint is introduced, where the slope of the warping function is restricted to improve the discriminative power between different categories of words.

[0004] Rabiner and Lawrence R.’s “A tutorial on hidden Markov models and selected applications in speech recognition”, published in Proceedings of the IEEE 77.2 (1989): 257-286, reviews the theoretical aspects of various types of statistical modeling and demonstrates how they are applied to selected problems in machine speech recognition.

[0005] U.S. Patent 5,864,810 describes a method and apparatus for automatic speech recognition that develops a transformation using adaptive data to adapt to a specific speaker. This transformation converts a speaker-independent model into a speaker-adaptive model. The speaker-adapted model is then used for speaker recognition, achieving better recognition accuracy than the non-adaptive model. In another embodiment, the transformation-based adaptation technique is combined with known Bayesian adaptation techniques.

[0006] U.S. Patent 9,922,641 describes a method comprising receiving input speech data from a speaker speaking a first language, and estimating a speaker transform representing speaker features associated with the input speech data based on a general speech model. The method further comprises acquiring a speaker-independent speech model for generating speech data in a second language different from the first language. The method also includes using the speaker transform to modify the speaker-independent speech model to obtain a speaker-specific speech model, and using the speaker-specific speech model to generate the second language speech data. Invention Overview

[0008] According to some embodiments of the present invention, an apparatus including a communication interface and a processor is provided. The processor is configured to receive multiple voice samples via the communication interface. m = 1...M, these multiple speech samples were spoken by the subjects in their initial state regarding the disease; and using... A subject-specific discriminator is synthesized using at least one non-subject-specific reference discriminator, which is specific to the subject and configured to generate an output indicating the probability that the subject is in a second state regarding the disease in response to one or more test utterances spoken by the subject.

[0009] In some embodiments, the first state is a stable state, and the second state is an unstable state.

[0010] In some embodiments, the disease is selected from the group of diseases consisting of: congestive heart failure (CHF), coronary artery disease, arrhythmia, chronic obstructive pulmonary disease (COPD), asthma, interstitial lung disease, pulmonary edema, pleural effusion, Parkinson's disease, and depression.

[0011] In some embodiments, the processor is configured to synthesize a subject-specific discriminator in the following manner:

[0012] Generate a first-state subject-specific speech model θ 0 For any speech sample s, the first-state subject-specific speech model θ 0 Return a first distance metric, which indicates a first similarity between s and the subject's first-state speech, and

[0013] Generate a second-state subject-specific speech model θ 1 The second-state subject-specific speech model θ 1 Return a second distance metric that indicates the second similarity between s and the subject's second-state speech.

[0014] In some embodiments,

[0015] At least one reference discriminator includes K reference discriminators k = 1...K, include:

[0016] The corresponding first-state reference speech model returns the corresponding first distance. The corresponding first distance The first similarity between the instruction s and the corresponding reference first-state speech spoken by one or more other subjects in group K, and

[0017] The corresponding second-state reference speech model returns the corresponding second distance. The corresponding second distance The second similarity between the instruction s and the corresponding reference second-state speech spoken by the group.

[0018] θ 0 By applying the function To return the first distance metric, and

[0019] θ 1 By applying this function Return the second distance metric.

[0020] In some embodiments, this function is applied when... When, return The weighted average, yes A non-decreasing function.

[0021] In some embodiments, the weighted average is about K weights {w k}, k = 1...K These K weights make for The sum of the corresponding distance metrics is minimized with respect to the constraints, for those belonging to Each speech sample u m The distance metric is based on

[0022] In some embodiments, at least one reference discriminator includes:

[0023] The first-state reference speech model returns a first distance D. 0 (s), the first distance D 0 (s) indicates the first similarity between s and the reference first-state speech, and

[0024] The second-state reference speech model returns the second distance D. 1 (s), the second distance D 1 (s) indicates the second similarity between s and the reference second-state speech.

[0025] In some embodiments,

[0026] The first-state reference speech model returns D by applying a first function to a set of feature vectors V(s) extracted from s. 0 (s),

[0027] The second-state reference speech model returns D by applying a second function to V(s). 1 (s), and

[0028] Generate θ 0 and θ 1 This includes generating θ using the normalization transformation T. 0 and θ 1 The normalization transformation T optimally transforms under one or more predefined constraints.

[0029] In some embodiments, T with respect to constraints makes Minimize, where Δ is the third distance metric between any two sets of features, and u0 is The standard language used to describe the content.

[0030] In some embodiments, Δ is a non-decreasing function of the dynamic time warp (DTW) distance.

[0031] In some embodiments, T with respect to constraints makes Minimize f′0, where f′0 is the non-decreasing function of the first function.

[0032] In some embodiments,

[0033] θ 0 The first distance metric is returned by applying the first function to T(V(s)), and

[0034] θ 1 The second distance metric is returned by applying the second function to T(V(s)).

[0035] In some embodiments,

[0036] Generate θ 0 This includes generating θ by applying a denormalizing transformation T′ to the first parameters of the first-state reference speech model. 0 The inverse normalization transformation T′ optimally transforms the first parameter under one or more predefined constraints, and

[0037] Generate θ 1 This includes generating θ by applying T′ to the second parameter of the second-state reference speech model. 1 .

[0038] In some embodiments, T′ under the constraint makes Minimize, T′(D 0 )(s) is the first distance returned by the first-state reference speech model under the transformation.

[0039] In some embodiments,

[0040] The first-state reference speech model includes a first Hidden Markov Model (HMM), which comprises multiple first kernels. The first parameters include the kernel parameters of the first kernels, and...

[0041] The second-state reference speech model includes a second HMM, which includes multiple second kernels, and the second parameters include the second kernel parameters of the second kernels.

[0042] In some embodiments, the first core and the second core are Gaussian kernels, and T′ includes:

[0043] Affine transformations that operate on the mean vectors of any one or more Gaussian kernels, and

[0044] A quadratic transformation that operates on the covariance matrix of any one or more Gaussian kernels.

[0045] In some embodiments,

[0046] The first-state reference speech model includes multiple first reference frames, and the first parameters include the first reference frame features of the first reference frames, and

[0047] The second-state reference speech model includes multiple second reference frames, and the second parameters include the second reference frame features of the second reference frames.

[0048] In some embodiments,

[0049] The reference first-state speech includes multiple first-state reference speech samples spoken by a first subset of R other subjects.

[0050] The reference second-state speech includes multiple second-state reference speech samples spoken by a second subset of the other subjects, and

[0051] The processor is also configured to:

[0052] For other subjects, identify the corresponding transformation {T}.r}, r = 1...R, for each r-th additional subject among these additional subjects, T r It is to optimally transform {Φ} under one or more predefined constraints. r The normalization transformation of}, {Φ r} is the union of (i) those samples in the first state reference speech sample spoken by the other subject and (ii) those samples in the second state reference speech sample spoken by the other subject.

[0053] For each r-th additional subject among the other subjects, by T r Applied to

[0054] {V(Φ r To calculate the modified feature set, and

[0055] A reference discriminator is generated from the modified feature set.

[0056] In some embodiments,

[0057] The first-state reference speech model and the second-state reference speech model are the same with respect to the first set of parameters, but different with respect to the second set of parameters.

[0058] The processor is configured to generate θ 0 This makes θ such that, with respect to the second set of parameters 0 Same as the first-state reference speech model, and

[0059] The processor is configured to generate θ 1 This makes θ such that, with respect to the first set of parameters 1 With θ 0 The same, and regarding the second set of parameters, θ 1 Same as the second-state reference speech model.

[0060] In some embodiments,

[0061] The first-state reference speech model and the second-state reference speech model each include different corresponding Hidden Markov Models (HMMs), each of which includes multiple kernels with corresponding kernel weights.

[0062] The first set of parameters includes kernel weights, and

[0063] The second set of parameters includes the kernel parameters.

[0064] In some embodiments,

[0065] At least one reference discriminator includes a reference neural network associated with multiple parameters, which, for any one or more speech samples, returns another output indicating the likelihood that the speech sample will be spoken in a second state.

[0066] The processor is configured to adjust a subset of parameters to make it suitable for including The error of the other output of a set of input speech samples is minimized, thereby synthesizing a subject-specific discriminator by synthesizing a subject-specific neural network.

[0067] In some embodiments, the parameters include multiple neuron weights, and a subset of the parameters includes a subset of the weights.

[0068] In some embodiments, the reference neural network includes multiple layers, and a subset of weights includes at least some of the weights associated with one of these layers, but excludes any weights associated with another of these layers.

[0069] In some embodiments,

[0070] These layers include: (i) an acoustic layer with one or more neurons that generates an acoustic layer output in response to input based on speech samples; (ii) a speech layer with one or more neurons that generates a phonetic layer output in response to the acoustic layer output; and (iii) a discrimination layer with one or more neurons that generates another output in response to the speech layer output.

[0071] The subset of weights includes at least some of the weights associated with the acoustic layer and the discrimination layer, but excludes any weights associated with the speech layer.

[0072] In some embodiments, a subset of the parameters includes speaker recognition parameters that identify the speaker of the speech sample.

[0073] In some embodiments, the set of input speech samples may also include one or more second-state speech samples.

[0074] According to some embodiments of the present invention, a method is also provided, the method comprising receiving a plurality of speech samples. m = 1...M, these multiple speech samples The subject speaks about the disease in their initial state. The method also includes using... A subject-specific discriminator is synthesized from at least one non-subject-specific reference discriminator, which is specific to the subject and configured to generate an output indicating the probability that the subject is in a second state regarding the disease in response to one or more test utterances by the subject.

[0075] According to some embodiments of the present invention, a computer software product is also provided, comprising a tangible, non-transitory computer-readable medium in which program instructions are stored. When a processor reads the instructions, the instructions cause the processor to receive a plurality of speech samples. m = 1...M, these multiple speech samples Spoken by the subject in their initial state regarding the disease, and used A subject-specific discriminator is synthesized from at least one non-subject-specific reference discriminator, which is specific to the subject and configured to generate an output indicating the probability that the subject is in a second state regarding the disease in response to one or more test utterances by the subject.

[0076] The invention will be more fully understood from the following detailed description of embodiments thereof, taken in conjunction with the accompanying drawings, in which: Brief description of the attached diagram

[0078] Figure 1 This is a schematic diagram of a system for assessing the physiological state of a subject according to some embodiments of the present invention;

[0079] Figures 2-4 This is a flowchart of a technique for generating subject-specific speech models according to some embodiments of the present invention; and

[0080] Figure 5 This is a schematic diagram of a neural network discriminator according to some embodiments of the present invention. Detailed Implementation

[0081] Glossary

[0082] In the context of this application, including the claims, a subject is referred to as being in an “unstable state” with respect to a physiological condition (or “disease”) if the subject is experiencing an acute deterioration of that condition. Otherwise, the subject is referred to as being in a “stable state” with respect to that condition.

[0083] In the context of this application, including the claims, a “speech model” refers to a computer-implemented function configured to map speech samples to outputs indicating the attributes of the samples. For example, given a speech sample s spoken by a subject, a speech model may return a distance metric D(s) indicating the similarity between s and a reference speech of the subject or another subject.

[0084] In the context of this application, including the claims, a “discriminator” refers to a set of one or more models, typically machine learning models, configured to distinguish between various states. For example, given a set of states for a particular physiological condition, such as “stable” and “unstable,” the discriminator can generate an output indicating the probability that the subject is in one of these states based on a sample of the subject’s speech.

[0085] Overview

[0086] For subjects with physiological conditions, it may be desirable to train a discriminator configured to determine whether the subject is in a stable or unstable state regarding that condition based on their speech. However, a challenge is that obtaining a sufficient number of training samples for each state can be difficult. For example, for a generally stable subject, a sufficient number of speech samples spoken in a stable state may be available, but obtaining a sufficient number of speech samples spoken in an unstable state may be difficult. For other subjects (e.g., after hospitalization), collecting a sufficient number of unstable state samples may be straightforward, but collecting a sufficient number of stable state samples may be less straightforward.

[0087] To address this challenge, embodiments of the present invention generate a subject-specific discriminator (i.e., configured to discriminate against the subject) from a non-subject-specific reference discriminator. To generate the subject-specific discriminator, the processor modifies or adjusts the reference discriminator using speech samples spoken by the subject in one of the states. This process is referred to as the “synthesis” of the subject-specific discriminator, and advantageously, speech samples spoken by the subject in the other state are not required.

[0088] The techniques described herein can be used to synthesize discriminators for any suitable physiological condition, such as congestive heart failure (CHF), coronary artery disease, atrial fibrillation or any other type of arrhythmia, chronic obstructive pulmonary disease (COPD), asthma, interstitial lung disease, pulmonary edema, pleural effusion, Parkinson's disease, or depression.

[0089] System Description

[0090] First refer to Figure 1 , Figure 1 This is a schematic diagram of a system 20 for assessing the physiological state of a subject 22 according to some embodiments of the present invention.

[0091] System 20 includes an audio receiving device 32 used by subject 22, such as a mobile phone, tablet computer, laptop computer, desktop computer, or voice-controlled personal assistant (e.g., Amazon Echo). TM Or Google HomeTM Device 32 may be an audio sensor 38 (e.g., a microphone) that converts sound waves into analog electrical signals, an analog-to-digital (A / D) converter 42, a processor 36, and a network interface (e.g., a network interface controller (NIC) 34). Typically, device 32 also includes a storage device (e.g., a solid-state drive), a screen (e.g., a touchscreen), and / or other user interface components, such as a keyboard and a speaker. In some embodiments, the audio sensor 38 (and optionally, the A / D converter 42) is an external unit to device 32. For example, the audio sensor 38 may be part of a headset that is connected to device 32 via a wired or wireless connection (e.g., Bluetooth).

[0092] System 20 also includes a server 40, which includes circuitry including a processor 28, a storage device 30 (such as a hard disk drive or flash drive), and a network interface (such as a network interface controller (NIC) 26). Server 40 may also include a screen, keyboard, and / or any other suitable user interface components. Typically, server 40 is located remotely from device 32, for example, in a control center, and server 40 and device 32 communicate with each other via their respective network interfaces through a network 24, which may include a cellular network and / or the Internet.

[0093] System 20 is configured to assess a subject's physiological state by processing one or more speech signals (referred to herein as "speech samples") received from the subject. Typically, processor 36 of device 32 and processor 28 of server 40 cooperate in receiving and processing at least some speech samples. For example, when a subject speaks to device 32, the sound waves of the subject's speech can be converted into an analog signal by audio sensor 38, which can then be sampled and digitized by A / D converter 42. (Typically, the subject's speech can be sampled at any suitable rate (e.g., between 8 kHz and 45 kHz). The resulting digital speech signal can be received by processor 36. Processor 36 can then transmit the speech signal to server 40 via NIC 34, so that processor 28 receives the speech signal via NIC 26. Processor 28 can then process the speech signal.

[0094] To process the subject's speech signals, processor 28 uses a subject-specific discriminator 44, which is specific to subject 22 and stored in storage device 30. Based on each input speech signal, the subject-specific discriminator generates an output indicating the probability that the subject is in a particular physiological state. For example, regarding physiological state, the output may indicate the probability that the subject is in a stable state and / or the probability that the subject is in an unstable state. Alternatively or additionally, the output may include a score indicating the degree to which the subject's state appears unstable. As described in detail below with reference to the following figures, processor 28 is also configured to synthesize subject-specific discriminator 44 before using the subject-specific discriminator.

[0095] In response to the output from the subject-specific discriminator, the processor can generate any appropriate audio or visual output to the subject and / or another individual (e.g., the subject's doctor). For example, processor 28 can transmit this output to processor 36, and processor 36 can then transmit the output to the subject, for example, by displaying a message on the screen of device 32. Alternatively or additionally, in response to the subject-specific discriminator outputting a relatively high probability that the subject's state is unstable, the processor can generate an alert instructing the subject to take medication or see a doctor. Such an alert can be transmitted by placing a call to the subject, the subject's doctor, and / or monitoring center, or by sending a message (e.g., a text message). Alternatively or additionally, in response to the output from the discriminator, the processor can control the medication administration device to adjust the amount of medication administered to the subject.

[0096] In other embodiments, after synthesizing the subject-specific discriminator, processor 28 transmits the subject-specific discriminator to processor 36, and processor 36 then stores the discriminator in a storage device belonging to device 32. Subsequently, processor 36 can use the discriminator to assess the physiological state of subject 22. Alternatively, the synthesis of the subject-specific discriminator may even be performed by processor 36. (Nevertheless, for simplicity, the remainder of this specification generally assumes that processor 28—hereinafter referred to simply as "processor"—performs the synthesis.)

[0097] In some embodiments, device 32 includes an analog telephone that does not include an A / D converter or processor. In such an embodiment, device 32 transmits analog audio signals from audio sensor 38 to server 40 via a telephone network. Typically, in the telephone network, the audio signals are digitized, transmitted digitally, and then converted back to analog signals before reaching server 40. Therefore, server 40 may include an A / D converter that converts incoming analog audio signals received via a suitable telephone network interface into digital voice signals. Processor 28 receives the digital voice signals from the A / D converter and then processes the signals as described above. Alternatively, server 40 may receive the signals from the telephone network before the signals are converted back to analog signals, so that the server does not necessarily include an A / D converter.

[0098] As further described below with reference to the accompanying figures, processor 28 synthesizes a subject-specific discriminator 44 using training speech samples spoken by subject 22 in a known physiological state. Each of these samples can be received via a network interface as described above or via any other suitable communication interface (e.g., a flash drive interface). Similarly, processor 28 can receive at least one reference discriminator not specific to subject 22 or training samples from another subject that can be used to generate the reference discriminator, which is also used to synthesize the subject-specific discriminator, via any suitable communication interface.

[0099] Processor 28 may be a single processor or a networked or clustered group of processors. For example, a control center may include multiple interconnected servers, each including a corresponding processor, which collaboratively execute the techniques described herein. In some embodiments, processor 28 is a virtual machine.

[0100] In some embodiments, the functionality of processor 28 and / or processor 36 as described herein is implemented solely in hardware, such as using one or more fixed-function integrated circuits or general-purpose integrated circuits, application-specific integrated circuits (ASICs), and / or field-programmable gate arrays (FPGAs). Alternatively, the functionality may be implemented at least partially in software. For example, processor 28 and / or processor 36 may be embodied as a programmable processor including, for example, a central processing unit (CPU) and / or a graphics processing unit (GPU). Program code and / or data, including software programs, may be loaded for execution and processing by the CPU and / or GPU. For example, the program code and / or data may be downloaded to the processor electronically via a network. Alternatively or additionally, the program code and / or data may be provided and / or stored on a non-transitory tangible medium (e.g., magnetic storage, optical storage, or electronic storage). Such program code and / or data, when provided to the processor, create a machine or special-purpose computer configured to perform the tasks described herein.

[0101] Synthetic subject-specific discriminator

[0102] As described in the review above, conventional techniques for generating discriminators to distinguish between two states typically require a sufficient number of training samples for each state. However, in some cases, the processor may only have enough training samples for one of the states. To address this, the processor synthesizes a subject-specific discriminator.

[0103] To perform the synthesis, the processor first receives multiple speech samples. m = 1…M, these multiple speech samples This is stated by the subject while they are in a primary state regarding the disease (e.g., stable state). Next, the processor uses... A subject-specific discriminator is synthesized using at least one non-subject-specific reference discriminator. Advantageously, although the processor has few or no speech samples spoken by the subject in a second state (e.g., unstable state) regarding the disease, the subject-specific discriminator can generate an output indicating the probability that the subject is in a second state in response to one or more test utterances spoken by the subject.

[0104] Multi-model discriminator

[0105] In some embodiments, the subject-specific discriminator includes a first-state subject-specific speech model θ 0 Second-state subject-specific speech model θ 1 For any speech sample s, θ 0 Return the first distance metric, which indicates the similarity between s and the subject's first-state speech, while θ1 The system returns a second distance metric indicating the similarity between s and the subject's second-state speech. In such an embodiment, the subject-specific discriminator can generate output based on a comparison between the two distance metrics. For example, assuming a convention where a larger distance indicates a smaller similarity, the subject-specific discriminator can generate output indicating that the subject is likely in a first state in response to the ratio between the first and second distance metrics being less than a threshold. Alternatively, the subject-specific discriminator can output the corresponding probability for both states based on the distance metrics, or simply output both distance metrics.

[0106] Various techniques can be used to synthesize such a multi-model discriminator. (See reference below.) Figures 2-4 Examples describing these technologies.

[0107] (i) The first technology

[0108] Now for reference Figure 2 , Figure 2 This is for generating θ according to some embodiments of the present invention. 0 and θ 1 The flowchart of the first technology 46.

[0109] Technique 46 begins with the first receiving or generating step 48, where the processor receives or generates K≥1 reference discriminators. k = 1...K. (Note that the processor can receive some discriminators and generate others simultaneously.) This includes a corresponding first-state reference speech model and a corresponding second-state reference speech model, which are specific to one or more additional subjects in the same K groups, referred to herein as "reference subjects". In other words, for any speech sample s, the first-state reference speech model returns a corresponding first distance. k = 1...K, these are the corresponding first distances The similarity between the indicator s and the corresponding reference first-state speech spoken by the K groups is calculated, while the second-state reference speech model returns the corresponding second distance. k = 1...K, these corresponding second distances The similarity between the indicator s and the corresponding reference second-state speech spoken by the K groups. In some embodiments, each reference speech model includes a parametric statistical speech model, such as a Hidden Markov Model (HMM).

[0110] Subsequently, at voice sample receiving step 50, the processor retrieves data from subject 22 (…). Figure 1 ) Receive one or more first-state speech samples Next, at step 52 of the first state model generation, the processor computes function "f", which is used to set the distances. Transformed into a single transformed distance Makes it possible for Another function of the transformed distance is minimized with respect to one or more suitable constraints. The processor thus generates θ. 0 This makes it possible to obtain any speech sample s by applying the function "f" to the speech sample s. To calculate by θ 0 The returned distance metric.

[0111] For example, the processor can identify and make The function "f" that minimizes q≥0 with respect to constraints. Alternatively, the function "f" can be a weighted sum. Regarding constraint minimization, in such an embodiment, the weight β for each speech sample... m It can be a function of sample quality, as higher-quality samples can be assigned a larger weight. Alternatively or additionally, those speech samples whose transformed distance is greater than a predefined threshold (such as a specific percentage of the transformed distance) can be assumed to be outliers and therefore can be assigned a weight of 0.

[0112] Subsequently, at step 54 of the first second-state model generation, the processor applies the same function... Generate θ 1 In other words, the processor generates θ. 1 Such that for any speech sample s, by θ 1 The returned distance metric is equal to

[0113] In fact, in technique 46, the processor uses the subject's first-state speech samples to learn a way in which the subject's voice in the first state can be optimally approximated as a function of the voices of K sets of reference subjects in the first state. The processor then assumes the same approximation applies to the second state, such that for θ... 0 The function can also be used for θ 1 .

[0114] As a specific example, the function calculated in step 52 of the first state model generation is applied when... When this happens, the function can return... The weighted average, yes Non-decreasing functions, for example, for p≥1 In other words, for any speech sample s, by θ 0The returned distance metric can be equal to the distance for K weights {w}. k}, k = 1...K Similarly, in such an embodiment, by θ 1 The returned distance metric can be equal to yes It is also a non-decreasing function. In fact, this function approximates the subject's voice as a weighted average of the voices of K groups of reference subjects.

[0115] In such an embodiment, in order to calculate K weights in the first state model generation step 52, the processor can make the processor capable of calculating the weights for... The sum of the corresponding distance metrics with respect to constraints (e.g., Minimize, for those belonging to Each speech sample u m The distance metric is based on the transformed distance. For example, the processor can enable operations for q≥0. Regarding the minimization of validity constraints. (For which...) In some embodiments, q is typically equal to 1 / p. As described above, for example, the transformed distance can be weighted in response to changes in the quality of the samples.

[0116] In some embodiments, to simplify subject-specific models, the processor makes relatively low weights (such as less than {w}) k A specific percentage and / or weights less than a predefined threshold are nullified. The processor can then rescale the remaining non-zero weights so that the sum of the weights is 1. For example, the processor can nullify weights excluding the largest weight w. 最大 All weights except θ are set to zero, so that by θ 0 The returned distance metric is equal to Where k 最大 It is w 最大 The index. Therefore, in practice, the subject's voice can be approximated by the voice of a single reference subject in K groups of reference subjects, while ignoring the other K-1 groups.

[0117] (ii) The second technology

[0118] Now for reference Figure 3 , Figure 3 This is for generating θ according to some embodiments of the present invention. 0 and θ 1 The flowchart of the second technology 56.

[0119] Technique 56 begins with the second receiving or generating step 58, where the processor receives or generates a first state reference speech model and a second state reference speech model (each of which is not subject-specific). This is consistent with Technique 46 ( Figure 2 Each first-state reference model in ) is similar; the first-state reference speech model in Technique 56 returns a first distance D. 0 (s), the first distance D 0 (s) indicates the similarity between any speech sample s and the reference first-state speech. Similarly, similar to each second-state reference model in technique 46, the second-state reference speech model in technique 56 returns a second distance D. 1 (s), the second distance D 1 (s) indicates the similarity between s and the reference second-state speech.

[0120] For example, the first-state reference speech model can return D by applying a first function f0 to a set of feature vectors V(s) extracted from s. 0 (s)(i.e., D) 0 (s) may be equal to f0(V(s)), while the second-state reference speech model can return D by applying the second function f1 to V(s). 1 (s)(i.e., D) 1 (s) may be equal to f1(V(s)). Each reference speech model can include a parametric statistical speech model, such as a Hidden Markov Model (HMM).

[0121] However, unlike in technique 46, the two reference models are not necessarily generated from the same set of subjects' reference speech. For example, the first-state reference speech model may be generated from reference first-state speech from one or more subjects in a group, while the second-state reference speech model may be generated from reference second-state speech from another group of one or more subjects. Alternatively, one or both of these models may be generated from artificial speech generated by a speech synthesizer. Therefore, as will be described in detail below, technique 56 differs from technique 46.

[0122] After performing the second receiving or generating step 58, at the voice sample receiving step 50, the processor receives... Next, in some embodiments, at transformation calculation step 60, the processor calculates transformation T, which optimally transforms under one or more predefined constraints. T can be called a "feature normalization" transformation because T transforms the features of the subject's speech sample in order to neutralize the subject's vocal tract specificity; that is, T makes the speech sample more general or standardized.

[0123] For example, T can make Regarding constraint minimization, f′0 is a non-decreasing function of f0. (For example, for p≥1, f′0(*) can be equal to |f0(*)|.) p Alternatively, T can make Minimize under one or more predefined validity constraints, where Δ is the distance metric between any two sets of feature vectors, and for any set of feature vectors belonging to... For each sample u, u0 is the canonical utterance of the content of u, such as the synthesized utterance of the content. In some embodiments, Δ is a non-decreasing function of the Dynamic Time Warp (DTW) distance, which can be computed as described in the references of Sakoe and Chiba cited in the background, which are incorporated herein by reference. For example, Δ(T(V(u)),V(u0)) can be equal to |DTW(T(V(u)),V(u0))|p, where DTW(V1,V2) is the DTW distance between two sets of feature vectors V1 and V2, and p≥1.

[0124] (Note that, typically, the DTW distance between two sets of feature vectors is calculated by mapping each feature vector in one set to a corresponding feature vector in the other set, such that the sum of the corresponding local distances between the feature vector pairs is minimized. The local distance between each pair of vectors can be calculated by summing the squared differences between the corresponding components of the vectors, or by using any other suitable function.)

[0125] Typically, the processor extracts N overlapping or non-overlapping frames from each received speech sample *s*, where N is a function of the predefined length of each frame. V(s) therefore comprises N feature vectors {v...} n}, n = 1...N, one feature vector per frame. (Each feature vector may include, for example, a set of cepstral coefficients and / or a set of linear prediction coefficients with respect to the frame.) Typically, T includes transformations that operate independently on each feature vector, i.e., T(V(s)) = {T(v... n )}, n=1...N. For example, T can include affine transformations that operate independently on each eigenvector, i.e., T(V(s)) can be equal to {Av} n Let A be an L×L matrix, and b be an L×1 vector, where L is the vector of each vector v. n The length.

[0126] After calculating T, at step 62 of the second first-state model generation, the processor generates θ. 0 (For the first-state model of the subject) such that for any speech sample s, θ 0 Return f0(T(V(s))). Similarly, at step 64 of the second second-state model generation, the processor generates θ.1 Make θ 1 Returns f1(T(V(s))).

[0127] In other embodiments, instead of computing T, the processor computes an alternative transformation T′ in an alternative transformation computation step 66, which optimally transforms the parameters of the first-state reference speech model under one or more predefined constraints. For example, the processor may compute T′ such that T′ under constraints makes Minimize, T′(D 0 (s) is the distance returned by the first-state reference speech model under this transformation. Alternatively, after computing T, the processor can derive T′ from T such that applying T′ to the model parameters has the same effect as applying T to the characteristics of the subject's speech samples. T′ can be called a “parameter denormalization” transformation because T′ transforms the parameters of the reference model to better match the subject's vocal tract specificity, i.e., T′ makes the reference model more subject-specific.

[0128] In such an embodiment, after calculating T′, at the third first-state model generation step 68, the processor generates θ by applying T′ to the parameters of the first-state reference speech model. 0 Similarly, in the third second-state model generation step 70, the processor generates θ by applying T′ to the parameters of the second-state reference speech model. 1 In other words, the processor generates θ. 0 Such that for any speech sample s, θ 0 Return T′(D) 0 (s)=f0′(V(s)), where f0′ is different from f0 because the parameters of the first-state reference speech model modified by T′ are used; similarly, the processor generates θ 1 Make θ 1 Return T′(D) 1 (s)=f1′(V(s)), where f1′ is different from f1 because the parameters of the second-state reference speech model modified by T′ are used. (For the embodiment where T′ is derived from T as described above, f0′(V(s))=f0(T(V(s))) and f1′(V(s))=f1(T(V(s))).

[0129] For example, in the case where each reference speech model includes an HMM with multiple kernels, each subject-specific model can input T(V(s)) into the corresponding kernel of the reference speech model according to the former embodiment. Alternatively, according to the latter embodiment, the parameters of the kernel can be transformed using T′, and then V(s) can be input into the transformed kernel.

[0130] As a concrete example, each reference HMM may include multiple Gaussian kernels for each state, each kernel having the form v is any eigenvector belonging to V(s), μ is the mean vector, and σ is the covariance matrix with determinant |σ|. For example, assuming state x has J kernels, the local distance between v and x can be calculated as follows: Where g(v; μ) x,j ,ρ x,j ) is the j-th Gaussian kernel belonging to state x for j = 1…J, w x,j Here, L is the weight of the kernel, and L is any suitable scalar function, such as the identity function or the negative logarithmic function. In this case, T′ can include an affine transformation operating on the mean vector of any one or more kernels and a quadratic transformation operating on the covariance matrix of any one or more kernels. In other words, T′ can be derived by using μ′ = A -1 (μ+b) replace μ and use σ'=A -1 σA T Replace σ with the Gaussian kernel so that, for example, each local distance is calculated as (For the embodiment where T′ is derived from T as described above, g(v; μ′) x,j ,σ′ x,j ) equals g(T(v); μ x,j ,σ x,j ), where T(v) = Av + ​​b.

[0131] Alternatively, each reference speech model may include multiple reference frames. In such an embodiment, for each speech sample s, the distance returned by each reference speech model can be calculated (e.g., using DTW) by: taking each feature vector v n Mapping to one of the reference frames minimizes the sum of the corresponding local distances between the feature vector and the reference frames to which the feature vector is mapped. In this case, according to the former embodiment, each subject-specific model can map {T(v n )} is mapped to the reference frame of the corresponding reference model, where n = 1...N, such that the sum of local distances is minimized. Alternatively, according to the latter embodiment, the features of the reference frame can be transformed using T′, and then {v} can be mapped to the reference frame of the corresponding reference model. n} is mapped to the transformed reference frame, where n = 1...N.

[0132] Regardless of whether T is applied to the subject's speech sample or T′ is applied to the reference model, it is generally advantageous for the reference model to be as normal or independent of the subject as possible. Therefore, in some embodiments, particularly if the reference speech used to generate the reference model comes from a relatively small number of other subjects, the processor normalizes the reference speech before generating the reference model during receiving or generating step 58.

[0133] For example, the processor may first receive a first state reference speech sample spoken by a first subset of R other subjects, and a second state reference speech sample spoken by a second subset of the same other subjects. (The subsets may overlap, i.e., at least one of the other subjects may provide both the first and second state reference speech samples.) Next, for each r-th other subject, the processor may identify the union {Φ} of (i) the samples in the first state reference speech sample spoken by the r-th other subject and (ii) the samples in the second state reference speech sample spoken by the r-th other subject. r Subsequently, for other subjects, the processor can identify the corresponding transformation {T}. r}, r=1...R,T r The optimal transformation of {Φ} under the constraints described above. r Another normalization transformation of}. For example, T r Under predefined validity constraints, Minimize, Φ0 is the canonical (e.g., synthesized) discourse of the content of Φ. Next, for each r-th other subject in this additional subject, the processor can, by T r Applied to {V(Φ r The processor calculates the modified sets of features using a process called}. Finally, the processor generates a reference discriminator—comprising two reference models—based on the modified feature sets.

[0134] (iii) The third technology

[0135] Now for reference Figure 4 , Figure 4 This is for generating θ according to some embodiments of the present invention. 0 and θ 1 The flowchart of the third technology 72.

[0136] Similar to technology 56 ( Figure 3Technique 72 can handle situations where the first-state reference speech and the second-state reference speech come from different corresponding groups of subjects. Technique 72 only requires that the two reference models be identical with respect to the first set of parameters, but that the two reference models differ with respect to the second set of parameters, which are assumed to represent the influence of the subject's health status on the reference speech. Since this influence is assumed to be... Figure 1 The result is the same, so technique 72 generates θ. 0 and θ 1 , so that θ 0 and θ 1 The second set of parameters is the same as their corresponding reference models, while the first set of parameters is different.

[0137] Technique 72 begins with the third receiving or generating step 74, whereby the processor receives or generates a first state reference speech model and a second state reference speech model such that the two models are identical with respect to the first set of parameters, but different with respect to the second set of parameters.

[0138] For example, the processor may first receive or generate a first-state reference model. Subsequently, the processor can adapt the second-state reference model to the first-state reference model by modifying a second set of parameters (without modifying the first set) such that the sum of the corresponding distances returned by the second-state reference model for the second-state reference speech samples is minimized with respect to appropriate validity constraints. (Any suitable non-decreasing function (e.g., q ≥ 1 power of absolute value) can be applied to each distance in this summation.) Alternatively, the processor may first receive or generate the second-state reference model and then adapt the first-state reference model according to the second-state reference model.

[0139] In some embodiments, the reference models include different corresponding Hidden Markov Models (HMMs), each HMM including multiple kernels with corresponding kernel weights. In such embodiments, the first set of parameters may include kernel weights. In other words, the two reference models may include the same states, and in each state, the same number of kernels have the same kernel weights. The first set of parameters may also include state transition distances or probabilities. The second set of parameters (which differs from each other in the reference models with respect to the second set of parameters) may include kernel parameters (e.g., mean and covariance).

[0140] For example, for a first-state reference model, the local distance between any state x and any feature vector v can be The second state reference model can include the same states as the first state reference model, and for any state x, the local distance can be

[0141] After the third receiving or generating step 74, at the speech sample receiving step 50, the processor receives... Next, at step 76 of the fourth first-state model generation, the processor generates θ. 0 Make θ 0 The second set of parameters is the same as that of the first-state reference speech model. To perform this adaptation of the first-state reference model, the processor can use an algorithm similar to the Baum-Welch algorithm, as described in Section 6.4.3 of L. Rabiner and BH. Juang's "Fundamentals of Speech Recognition" (Prentice Hall, 1993), which is incorporated herein by reference. Specifically, the processor can first... 0 Initialize the parameters to have a first-state reference model. Next, the processor can... Each feature vector in the vector is mapped to θ 0 The processor then considers the corresponding states in the algorithm. For each state, the processor can then recalculate the first set of parameters for that state using the feature vector mapped to that state. The processor can then remap the feature vector back to the state. This process can then be repeated until convergence, i.e., until the mapping no longer changes.

[0142] After the fourth first-state model generation step 76, at the fourth second-state model generation step 78, the processor generates θ. 1 Make θ 1 Regarding the first set of parameters and θ 0 The second set of parameters is the same as that of the second-state reference speech model.

[0143] Neural network discriminator

[0144] In an alternative embodiment, the processor synthesizes a subject-specific neural network discriminator instead of a multi-model discriminator. Specifically, the processor first receives or generates a reference discriminator comprising a neural network associated with multiple parameters. The processor then adjusts some of these parameters, as described below, to adapt the network to the subject 22 ( Figure 1 ).

[0145] For further details on this technology, please refer to [link / reference]. Figure 5 , Figure 5 This is a schematic diagram of a neural network discriminator according to some embodiments of the present invention.

[0146] Figure 5The diagram illustrates how a reference neural network 80 can adapt to a specific subject. The neural network 80 is configured to receive speech-related input 82 based on one or more speech samples spoken by the subject. For example, the neural network may receive the speech samples themselves and / or features extracted from the samples, such as Mel-frequency cepstral coefficients (MFCCs). The neural network 80 may also receive text input 90, which includes, for example, an indication of the phonetic content of the speech samples. (The phonetic content may be predetermined or determined from the speech samples using speech recognition techniques.) For example, if the neural network is trained on N different utterances consecutively numbered 0…N-1, the text input 90 may include a bit sequence indicating the sequence number of the utterance spoken in the speech samples.

[0147] Given the aforementioned input, the neural network returns an output 92, which indicates the probability that the speech sample was spoken in the second state. For example, output 92 could explicitly include the probability that the speech sample was spoken in the second state. Alternatively, the output could explicitly include the probability that the speech sample was spoken in the first state, such that the output implicitly indicates the former probability. For example, if the output indicates a 30% probability of the first state, then the output could effectively indicate a 70% probability of the second state. As yet another alternative, the output could include corresponding scores for the two states from which the two probabilities can be calculated.

[0148] Typically, the neural network 80 includes multiple layers of neurons. For example, in an embodiment where the speech-related input 82 includes raw speech samples (rather than features extracted from the raw speech samples), the neural network may include one or more acoustic layers 84 that generate acoustic layer outputs 83 in response to the speech-related input 82. In effect, the acoustic layers 84 extract feature vectors from the input speech samples by performing acoustic analysis of the speech samples.

[0149] As another example, the neural network may include one or more speech layers 86 that generate speech layer output 85 in response to acoustic layer output 83 (or in response to similar features contained in speech-related input 82). For example, speech layer 86 may match the acoustic features of a speech sample specified by acoustic layer output 83 with the expected speech content of a speech sample indicated by text input 90. Alternatively, the network may be configured for a single predefined text, and therefore speech layer 86 and text input 90 may be omitted.

[0150] As another example, the neural network may include one or more discriminative layers 88 that generate output 92 in response to speech layer output 85 (and optionally, acoustic layer output 83). The discriminative layers 88 may include, for example, one or more layers of neurons that compute features for distinguishing between a first health state and a second health state, followed by an output layer that generates output 92 based on these features. The output layer may include, for example, first-state output neurons and second-state output neurons, the first-state output neurons outputting a score indicating the probability of the first state, and the second-state output neurons outputting another score indicating the probability of the second state.

[0151] In some embodiments, neural network 80 is a deep learning network because it incorporates a relatively large number of layers. Alternatively or additionally, the network may include specialized elements such as convolutional layers, skip layers, and / or recurrent neural network components. Neurons in neural network 80 may be associated with various types of activation functions.

[0152] To synthesize a subject-specific neural network discriminator, the processor adjusts a subset of parameters associated with network 80 to make the parameters, including... The processor input is minimized by reducing the error of 92% of the output from a set of input speech samples. In other words, the processor input... And optionally, one or more speech samples spoken by the subject or another subject while in the second state, and adjusting a subset of parameters to minimize the error of the output 92.

[0153] For example, a processor can adjust some or all of the weights of the corresponding neurons belonging to the network. As a concrete example, a processor can adjust at least some weights associated with one layer of neurons, without adjusting any weights associated with another layer. For example, as... Figure 5 As shown, the processor can adjust the weights associated with acoustic layer 84 and / or with discrimination layer 88, which are assumed to be relevant to the subject, but the processor will not adjust the weights associated with speech layer 86.

[0154] In some embodiments, the neural network is associated with a speaker identification (or “subject ID”) parameter 94, which identifies the speaker of the speech sample used to generate the speech-related input 82. For example, given R consecutively numbered reference subjects whose speech is used to train network 80, parameter 94 may include a sequence of R digits. For each input 82 taken from one of these subjects, in parameter 94, the subject’s sequence number may be set to 1, and the other digits may be set to 0. Parameter 94 may be input to acoustic layer 84, speech layer 86, and / or discrimination layer 88.

[0155] In such an embodiment, the processor can adjust parameter 94, alternatively or additionally, the neuron weights. By adjusting parameter 94, the processor can effectively approximate the subject's voice as a combination of the corresponding voices of some or all reference subjects. As an illustrative example only, for R = 10, the processor can adjust parameter 94 to the value [0.5 0 0 0.3 0 0 0.2 0], indicating that the subject's voice is approximated by a combination of the corresponding voices of the first, fifth, and ninth reference subjects. (Therefore, parameter 94 becomes associated with the network because it is a fixed parameter of the network, rather than merely as a variable input to the network.)

[0156] To tune the parameters, the processor can use any suitable technique known in the art. One such technique is backpropagation, which iteratively subtracts a vector of values ​​from the parameters, which is a multiple of the gradient of the bias function with respect to the parameters, quantifying the deviation between the network's output and the expected output. Backpropagation can be performed on each sample in a set of input speech samples (optionally iterating over the samples multiple times) until a suitable degree of convergence is reached.

[0157] Those skilled in the art will recognize that the present invention is not limited to what has been specifically shown and described above. Rather, the scope of embodiments of the invention includes combinations and sub-combinations of the various features described above, as well as variations and modifications thereof that would occur to those skilled in the art upon reading the above description and that are not found in the prior art. For example, embodiments of the invention include synthesizing a single-model subject-specific discriminator, such as a neural network discriminator, from a reference discriminator comprising a first-state reference speech model and a second-state reference speech model.

[0158] Documents incorporated herein by reference shall be considered part of this application, and the definitions in this specification shall be taken into account only, except that any terms are defined in those incorporated documents in a manner that conflicts to some extent with the definitions expressly or implicitly made herein.

Claims

1. An apparatus for speech signal processing, comprising: Communication interface; as well as Processor, the processor being configured to: Multiple voice samples are received via the communication interface. m = 1...M, the multiple speech samples The subject's statement when in the initial state regarding the disease, and use A subject-specific discriminator is synthesized using at least one reference discriminator that is not specific to the subject and without even using any additional speech samples spoken by the subject in a second state regarding the disease. This subject-specific discriminator is specific to the subject and is configured to generate an output indicating the probability that the subject is in the second state in response to one or more test utterances spoken by the subject. The processor is configured to synthesize the subject-specific discriminator in the following manner: Generate a first-state subject-specific speech model θ 0 For any speech sample s, the first state subject-specific speech model θ 0 Returns a first distance metric, which indicates a first similarity between s and the subject's first-state speech, and Generate a second-state subject-specific speech model θ 1 The second state subject-specific speech model θ 1 Return a second distance metric, which indicates a second similarity between s and the subject's second-state speech.

2. The apparatus according to claim 1, wherein, The first state is a stable state, and the second state is an unstable state.

3. The apparatus according to claim 1, wherein, The diseases mentioned are selected from the group consisting of the following: congestive heart failure (CHF), coronary artery disease, arrhythmia, chronic obstructive pulmonary disease (COPD), asthma, interstitial lung disease, pulmonary edema, pleural effusion, Parkinson's disease, and depression.

4. The apparatus according to claim 1, in, The at least one reference discriminator includes K reference discriminators. k = 1...K, include: The corresponding first-state reference speech model returns the corresponding first distance. The corresponding first distance The first similarity between the instruction s and the corresponding reference first-state speech spoken by one or more other subjects in group K, and The corresponding second-state reference speech model returns the corresponding second distance. The corresponding second distance The second similarity between the instruction s and the corresponding reference second-state speech spoken by the group. Where, θ 0 By applying the function To return the first distance metric value, and Where, θ 1 By applying the function Return the second distance metric value.

5. The apparatus according to claim 4, wherein, The function is applied when When, return The weighted average, yes A non-decreasing function.

6. The apparatus according to claim 5, wherein, The weighted average is about K weights {w k }, k = 1...K The K weights make for... The sum of the corresponding distance metrics is minimized with respect to the constraints, for those belonging to Each speech sample u m The distance metric is based on 7. The apparatus according to claim 1, wherein, The at least one reference discriminator includes: The first state reference speech model returns a first distance D. 0 (s), the first distance D 0 (s) indicates the first similarity between s and the reference first-state speech, and The second-state reference speech model returns the second distance D. 1 (s), the second distance D 1 (s) indicates the second similarity between s and the reference second-state speech.

8. The apparatus according to claim 7, in, The first state reference speech model returns D by applying a first function to a set of feature vectors V(s) extracted from s. 0 (s), The second state reference speech model returns D by applying a second function to V(s). 1 (s), and Among them, θ is generated 0 and θ 1 This includes generating θ using the normalization transformation T. 0 and θ 1 The normalization transformation T optimally transforms under one or more predefined constraints.

9. The apparatus according to claim 8, wherein, T regarding the constraint makes Minimize, where Δ is the third distance metric between any two sets of features, and u0 is The standard discourse of the content.

10. The apparatus according to claim 9, wherein, Δ is a non-decreasing function of the dynamic time warp (DTW) distance.

11. The apparatus according to claim 8, wherein, T regarding the constraint makes Minimize f′0, where f′0 is a non-decreasing function of the first function.

12. The apparatus according to claim 8, in, θ 0 The first distance metric is returned by applying the first function to T(V(s)), and Where, θ 1 The second distance metric is returned by applying the second function to T(V(s)).

13. The apparatus according to claim 7, in, Generate θ 0 This includes generating θ by applying an inverse normalization transformation T′ to the first parameters of the first state reference speech model. 0 The inverse normalization transformation T′ optimally transforms the first parameter under one or more predefined constraints, and Among them, θ is generated 1 This includes generating θ by applying T′ to the second parameter of the second state reference speech model. 1 .

14. The apparatus according to claim 13, wherein, T′ under the constraint makes T′ Minimize, T′(D 0 (s) is the first distance returned by the first state reference speech model under the transformation.

15. The apparatus according to claim 13, in, The first state reference speech model includes a first hidden Markov model (HMM) containing multiple first kernels, and the first parameters include the first kernel parameters of the first kernels, and The second state reference speech model includes a second HMM containing multiple second kernels, and the second parameters include the second kernel parameters of the second kernels.

16. The apparatus according to claim 15, wherein, The first kernel and the second kernel are Gaussian kernels, and wherein T′ includes: Affine transformations that operate on the mean vector of any one or more Gaussian kernels, and A quadratic transformation that operates on the covariance matrix of any one or more Gaussian kernels.

17. The apparatus according to claim 13, in, The first state reference speech model includes multiple first reference frames, and the first parameters include the first reference frame features of the first reference frames, and The second state reference speech model includes multiple second reference frames, and the second parameters include the second reference frame features of the second reference frames.

18. The apparatus according to claim 7, in, The reference first-state speech includes multiple first-state reference speech samples spoken by a first subset of R other subjects. The reference second-state speech includes multiple second-state reference speech samples spoken by a second subset of the other subjects, and The processor is further configured to: For the other subjects, identify the corresponding transformation {T}. r }, r = 1...R, for each r-th additional subject among the other subjects, T r It is to optimally transform {Φ} under one or more predefined constraints. r The normalization transformation of}, {Φ r } is the union of (i) the samples in the first state reference speech sample spoken by the other subject and (ii) the samples in the second state reference speech sample spoken by the other subject. For each r-th additional subject among the other subjects, by T r Applied to {V(Φ r To calculate the modified feature set, and The reference discriminator is generated from the modified feature set.

19. The apparatus according to claim 7, in, The first state reference speech model and the second state reference speech model are the same with respect to the first set of parameters, but different with respect to the second set of parameters. The processor is configured to generate θ 0 This makes θ such that, with respect to the second set of parameters 0 Same as the first state reference speech model, and The processor is configured to generate θ 1 So that with respect to the first set of parameters, θ 1 With θ 0 The same applies to the second set of parameters, θ 1 It is the same as the reference speech model in the second state.

20. The apparatus according to claim 19, in, The first state reference speech model and the second state reference speech model include different corresponding Hidden Markov Models (HMMs), each HMM including multiple kernels with corresponding kernel weights. The first set of parameters includes the kernel weights, and The second set of parameters includes the kernel parameters of the kernel.

21. An apparatus for speech signal processing, comprising: Communication interface; as well as Processor, the processor being configured to: Multiple voice samples are received via the communication interface. m = 1...M, the multiple speech samples The subject's statement when in the initial state regarding the disease, and use A subject-specific discriminator is synthesized using at least one reference discriminator that is not specific to the subject and without even using any additional speech samples spoken by the subject in a second state regarding the disease. This subject-specific discriminator is specific to the subject and is configured to generate an output indicating the probability that the subject is in the second state in response to one or more test utterances spoken by the subject. The at least one reference discriminator includes a reference neural network associated with multiple parameters comprising the weights of multiple neurons. For any one or more speech samples, the reference neural network returns another output indicating the probability that the speech sample is spoken in the second state. The processor is configured to adjust a subset of the parameters to make the parameters include The error of the other output of a set of input speech samples is minimized, thereby synthesizing the subject-specific discriminator by synthesizing a subject-specific neural network. The reference neural network includes: (i) one or more neuron acoustic layers that generate acoustic layer outputs in response to input based on the speech sample; (ii) one or more neuron speech layers that generate speech layer outputs in response to the acoustic layer outputs; and (iii) one or more neuron discrimination layers that generate the other output in response to the speech layer outputs. The subset of parameters includes at least some of the neuron weights associated with the acoustic layer and the discrimination layer, but excludes any neuron weights associated with the speech layer.

22. The apparatus according to claim 21, wherein, The subset of parameters also includes speaker recognition parameters that identify the speaker of the speech sample.

23. An apparatus for speech signal processing, comprising: Communication interface; as well as Processor, the processor being configured to: Multiple voice samples are received via the communication interface. m = 1...M, the multiple speech samples The subject's statement when in the initial state regarding the disease, and use A subject-specific discriminator is synthesized using at least one reference discriminator that is not specific to the subject and without even using any additional speech samples spoken by the subject in a second state regarding the disease. This subject-specific discriminator is specific to the subject and is configured to generate an output indicating the probability that the subject is in the second state in response to one or more test utterances spoken by the subject. The at least one reference discriminator includes a reference neural network associated with multiple parameters, which, for any one or more speech samples, returns another output indicating the probability that the speech sample is spoken in the second state. The processor is configured to adjust a subset of the parameters to make the parameters include The error of the other output of a set of input speech samples and one or more second-state speech samples is minimized, thereby synthesizing the subject-specific discriminator by synthesizing a subject-specific neural network.

24. A method for speech signal processing, comprising: Receive multiple voice samples m = 1...M, the multiple speech samples It was stated by the subject in their initial state regarding the disease; and use A subject-specific discriminator is synthesized using at least one reference discriminator that is not specific to the subject, without using any additional speech samples spoken by the subject in a second state related to the disease. This subject-specific discriminator is specific to the subject and is configured to generate an output indicating the probability that the subject is in the second state in response to one or more test utterances spoken by the subject. The synthesized subject-specific discriminator includes: Generate a first-state subject-specific speech model θ 0 For any speech sample s, the first state subject-specific speech model θ 0 Return a first distance metric, which indicates a first similarity between s and the subject's first-state speech; and Generate a second-state subject-specific speech model θ 1 The second state subject-specific speech model θ 1 Return a second distance metric, which indicates a second similarity between s and the subject's second-state speech.

25. The method according to claim 24, wherein, The first state is a stable state, and the second state is an unstable state.

26. The method according to claim 24, wherein, The diseases mentioned are selected from the group consisting of the following: congestive heart failure (CHF), coronary artery disease, arrhythmia, chronic obstructive pulmonary disease (COPD), asthma, interstitial lung disease, pulmonary edema, pleural effusion, Parkinson's disease, and depression.

27. The method according to claim 24, in, The at least one reference discriminator includes K reference discriminators. k = 1...K, include: The corresponding first-state reference speech model returns the corresponding first distance. The corresponding first distance The first similarity between the instruction s and the corresponding reference first-state speech spoken by one or more other subjects in group K, and The corresponding second-state reference speech model returns the corresponding second distance. The corresponding second distance The second similarity between the instruction s and the corresponding reference second-state speech spoken by the group. Where, θ 0 By applying the function To return the first distance metric value, and Where, θ 1 By applying the function Return the second distance metric value.

28. The method according to claim 27, wherein, The function is applied when When, return The weighted average, yes A non-decreasing function.

29. The method according to claim 28, wherein, The weighted average is about K weights {w k }, k = 1...K The K weights make for... The sum of the corresponding distance metrics is minimized with respect to the constraints, for those belonging to Each speech sample u m The distance metric is based on 30. The method according to claim 24, wherein, The at least one reference discriminator includes: The first state reference speech model returns a first distance D. 0 (s), the first distance D 0 (s) indicates the first similarity between s and the reference first-state speech, and The second-state reference speech model returns the second distance D. 1 (s), the second distance D 1 (s) indicates the second similarity between s and the reference second-state speech.

31. The method according to claim 30, in, The first state reference speech model returns D by applying a first function to a set of feature vectors V(s) extracted from s. 0 (s), The second state reference speech model returns D by applying a second function to V(s). 1 (s), and Among them, θ is generated 0 and θ 1 This includes generating θ using the normalization transformation T. 0 and θ 1 The normalization transformation T optimally transforms under one or more predefined constraints.

32. The method according to claim 31, wherein, T regarding the constraint makes Minimize, where Δ is the third distance metric between any two sets of features, and u0 is The standard discourse of the content.

33. The method according to claim 32, wherein, Δ is a non-decreasing function of the dynamic time warp (DTW) distance.

34. The method according to claim 31, wherein, T regarding the constraint makes Minimize f′0, where f′0 is a non-decreasing function of the first function.

35. The method according to claim 31, in, θ 0 The first distance metric is returned by applying the first function to T(V(s)), and Where, θ 1 The second distance metric is returned by applying the second function to T(V(s)).

36. The method according to claim 30, in, Generate θ 0 This includes generating θ by applying an inverse normalization transformation T′ to the first parameters of the first state reference speech model. 0 The inverse normalization transformation T′ optimally transforms the first parameter under one or more predefined constraints, and Among them, θ is generated 1 This includes generating θ by applying T′ to the second parameter of the second state reference speech model. 1 .

37. The method of claim 36, wherein, T′ under the constraint makes T′ Minimize, T′(D 0 (s) is the first distance returned by the first state reference speech model under the transformation.

38. The method according to claim 36, in, The first state reference speech model includes a first hidden Markov model (HMM) containing multiple first kernels, and the first parameters include the first kernel parameters of the first kernels, and The second state reference speech model includes a second HMM containing multiple second kernels, and the second parameters include the second kernel parameters of the second kernels.

39. The method according to claim 38, wherein, The first kernel and the second kernel are Gaussian kernels, and wherein T′ includes: Affine transformations that operate on the mean vector of any one or more Gaussian kernels, and A quadratic transformation that operates on the covariance matrix of any one or more Gaussian kernels.

40. The method according to claim 36, in, The first state reference speech model includes multiple first reference frames, and the first parameters include the first reference frame features of the first reference frames, and The second state reference speech model includes multiple second reference frames, and the second parameters include the second reference frame features of the second reference frames.

41. The method according to claim 30, in, The reference first-state speech includes multiple first-state reference speech samples spoken by a first subset of R other subjects. The reference second-state speech includes multiple second-state reference speech samples spoken by a second subset of the other subjects, and The method further includes: For the other subjects, identify the corresponding transformation {T}. r }, r = 1...R, for each r-th additional subject among the other subjects, T r It is to optimally transform {Φ} under one or more predefined constraints. r The normalization transformation of}, {Φ r } is the union of (i) the samples in the first state reference speech sample spoken by the other subject and (ii) the samples in the second state reference speech sample spoken by the other subject; For each r-th additional subject among the other subjects, by T r Applied to {V(Φ r To calculate the modified feature set; and The reference discriminator is generated from the modified feature set.

42. The method according to claim 30, in, The first state reference speech model and the second state reference speech model are the same with respect to the first set of parameters, but different with respect to the second set of parameters. Among them, θ is generated 0 Including generating θ 0 This makes θ such that, with respect to the second set of parameters 0 Same as the first state reference speech model, and Among them, θ is generated 1 Including generating θ 1 So that with respect to the first set of parameters, θ 1 With θ 0 The same applies to the second set of parameters, θ 1 It is the same as the reference speech model in the second state.

43. The method according to claim 42, in, The first state reference speech model and the second state reference speech model include different corresponding Hidden Markov Models (HMMs), each HMM including multiple kernels with corresponding kernel weights. The first set of parameters includes the kernel weights, and The second set of parameters includes the kernel parameters of the kernel.

44. A method for speech signal processing, comprising: Receive multiple voice samples m = 1...M, the multiple speech samples It was stated by the subject in their initial state regarding the disease; and use A subject-specific discriminator is synthesized using at least one reference discriminator that is not specific to the subject, without using any additional speech samples spoken by the subject in a second state related to the disease. This subject-specific discriminator is specific to the subject and is configured to generate an output indicating the probability that the subject is in the second state in response to one or more test utterances spoken by the subject. The at least one reference discriminator includes a reference neural network associated with multiple parameters comprising the weights of multiple neurons. For any one or more speech samples, the reference neural network returns another output indicating the probability that the speech sample is spoken in the second state. Specifically, synthesizing the subject-specific discriminator involves adjusting a subset of the parameters to make them suitable for subjects including... The error of the other output is minimized from a set of input speech samples, thereby synthesizing a subject-specific neural network. The reference neural network includes: (i) one or more neuron acoustic layers that generate acoustic layer outputs in response to input based on the speech sample; (ii) one or more neuron speech layers that generate speech layer outputs in response to the acoustic layer outputs; and (iii) one or more neuron discrimination layers that generate the other output in response to the speech layer outputs. The subset of parameters includes at least some of the neuron weights associated with the acoustic layer and the discrimination layer, but excludes any neuron weights associated with the speech layer.

45. The method according to claim 44, wherein, The subset of parameters also includes speaker recognition parameters that identify the speaker of the speech sample.

46. ​​A method for speech signal processing, comprising: Receive multiple voice samples m = 1...M, the multiple speech samples It was stated by the subject in their initial state regarding the disease; and use A subject-specific discriminator is synthesized using at least one reference discriminator that is not specific to the subject, without using any additional speech samples spoken by the subject in a second state related to the disease. This subject-specific discriminator is specific to the subject and is configured to generate an output indicating the probability that the subject is in the second state in response to one or more test utterances spoken by the subject. The at least one reference discriminator includes a reference neural network associated with multiple parameters, which, for any one or more speech samples, returns another output indicating the probability that the speech sample is spoken in the second state. Specifically, synthesizing the subject-specific discriminator involves adjusting a subset of the parameters to make them suitable for subjects including... The error of the other output of a set of input speech samples and one or more second-state speech samples is minimized, thereby synthesizing a subject-specific neural network.

47. A computer software product comprising a tangible, non-transitory computer-readable medium storing program instructions therein, the instructions, when read by a processor, causing the processor to: Receive multiple voice samples m = 1...M, the multiple speech samples The subject's statement when in the initial state regarding the disease, and use A subject-specific discriminator is synthesized using at least one reference discriminator that is not specific to the subject and without even using any additional speech samples spoken by the subject in a second state regarding the disease. This subject-specific discriminator is specific to the subject and is configured to generate an output indicating the probability that the subject is in the second state in response to one or more test utterances spoken by the subject. in, The instructions cause the processor to synthesize the subject-specific discriminator in the following manner: Generate a first-state subject-specific speech model θ 0 For any speech sample s, the first state subject-specific speech model θ 0 Returns a first distance metric, which indicates a first similarity between s and the subject's first-state speech, and Generate a second-state subject-specific speech model θ 1 The second state subject-specific speech model θ 1 Return a second distance metric, which indicates a second similarity between s and the subject's second-state speech.

Citation Information

Patent Citations

  • Method and apparatus for speech recognition adapted to an individual speaker

    US5864810A

  • Cross-lingual speaker adaptation for multi-lingual speech synthesis

    US9922641B1

  • System and method for pulmonary condition monitoring and analysis

    US20200098384A1