Paired neural networks for diagnosing health status via voice.

A neural network-based system processes voice samples across periods to improve health status diagnosis by extracting features and applying mathematical models, addressing variability in voices and enhancing diagnostic accuracy for various health conditions.

JP2026065038APending Publication Date: 2026-04-14CANARY SPEECH LLC
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
CANARY SPEECH LLC
Filing Date
2026-01-05
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Diagnosing health status through voice analysis is challenging due to the variability of individual voices and the need for improved accuracy beyond traditional methods.

Method used

A system utilizing paired neural networks processes voice samples from multiple periods to determine health status changes by extracting features, embedding voice data, and applying mathematical models to calculate health status labels, including acoustic and linguistic features, and using transformer or recurrent neural networks for analysis.

Benefits of technology

Enhances the accuracy of health status diagnosis by leveraging voice analysis, enabling the detection of conditions such as mental health issues, concussions, and heart failure, and predicting health-related events like hospital readmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026065038000001_ABST
    Figure 2026065038000001_ABST
Patent Text Reader

Abstract

To provide technology that improves the diagnosis of health conditions by processing human voice. [Solution] A person's health status or changes in health status can be determined by processing the person's speech using a neural network. Speech from two or more time periods may be processed, and in some embodiments, speech from certain time periods may be associated with health status labels. For each time period, a feature vector may be computed from the speech, and the feature vector may be processed using a neural network to obtain a speech embedding vector. In some embodiments, the feature vector may include wordpiece coding, and the neural network may be a transformer neural network. The speech embedding vector may be processed using a mathematical model to determine changes in health status between two time periods, or to determine health status labels for a particular time period.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001]

[0001] Improved diagnosis of health status has many advantages for society. For example, if the diagnosis of health status is improved, the quality of life is enhanced, the average life expectancy is extended, and furthermore, if early diagnosis and treatment are more effective than late diagnosis and treatment, medical costs may be reduced.

[0002]

[0002] Health status can be diagnosed in various ways. Some methods of diagnosing health status use a patient's voice. For example, a person's voice may be used to diagnose mental health status (stress, depression, anxiety), concussion, Alzheimer's disease, and congestive heart failure.

[0003]

[0003] In some cases, a person may listen to a person's voice and use the nature of that voice when making a diagnosis. In some cases, a mathematical model (such as a neural network) can process the voice to make a diagnosis and may provide a more accurate diagnosis than a trained medical professional. Improved techniques for diagnosing health status using mathematical models can provide many additional benefits to society.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Means for Solving the Problems

[0005]

[0004] The following detailed description of the present invention and its specific embodiments can be understood by referring to the following drawings.

Brief Description of the Drawings

[0006] [Figure 1]

[0005] Figure 1A is a diagram of an exemplary system for processing speech using a mathematical model to determine a health status label.

[0006] Figure 1B is a diagram of an exemplary system for processing a first speech from a first period and a second speech from a second period using a mathematical model to determine a change in health status between the first and second periods.

[0007]

[0007] Figure 1C is a diagram of an exemplary system for determining a second health status label for a second period by processing a first voice from a first period, a first health status label from a first period, and a second voice from a second period using a mathematical model.

[0008] Figure 1D is a diagram of an exemplary system for determining a health status label for the current time by processing multiple previous pairs of speech and health status labels from previous periods and a current speech sample from the current time, using a mathematical model. [Figure 2]

[0009] Figure 2 is a diagram illustrating an exemplary system for determining changes in health status between the first and second periods by processing a first voice from a first period and a second voice from a second period using a mathematical model. [Figure 3]

[0010] Figure 3 is a diagram illustrating an exemplary system that uses a mathematical model to process audio from two time periods and determine changes in health status using element-wise differences. [Figure 4]

[0011] Figure 4 is a diagram illustrating an exemplary system for using a mathematical model to process a first voice and first health status label from a first period, and a second voice from a second period, to determine a second health status label for the second period. [Figure 5]

[0012] Figure 5 is a diagram illustrating an exemplary system that uses a mathematical model to process multiple pairs of previous audio and health status labels from previous periods, and current audio from the current period, to determine the current health status label for the current period. [Figure 6]

[0013] Figure 6 is a flowchart illustrating an exemplary method for processing audio from two time periods using a mathematical model and determining changes in health status using element-wise differences. [Figure 7]

[0014] Figure 7 is a flowchart illustrating an exemplary method for using a mathematical model to process a first voice and first health status label from a first period, and a second voice from a second period, to determine a second health status label for the second period. [Figure 8]

[0015] Figure 8 shows components of one embodiment of a computing device 800 for carrying out any of the technologies described herein. [Modes for carrying out the invention]

[0008]

[0016] Voices vary from person to person, possessing a wide range of different properties and characteristics. Because voices differ from person to person, diagnosing a person's health can sometimes be difficult. For example, in the case of person 1, their voice may normally sound smooth, but after speaking for a long time, it may become hoarse and they may lose their voice altogether. However, in the case of person 2, their voice may always be hoarse, and this may be their normal way of speaking.

[0009]

[0017] To improve the diagnosis of health status through the processing of human voices, samples of a person's voice from multiple periods may be used. Continuing the example above, a sample of a person's voice from a first period in which the voice is not hoarse would help determine whether the person lost their voice in a second period. This specification describes techniques for improving the diagnosis of health status by processing a person's voice from two or more periods.

[0010]

[0018] Any appropriate health condition can be diagnosed using the techniques described herein. For example, health conditions may include mental health conditions (e.g., stress, depression, anxiety, and post-traumatic stress disorder), concussions, Parkinson's disease, Alzheimer's disease, and congestive heart failure. In some embodiments, health conditions may include the likelihood of health-related events occurring, such as the likelihood of readmission after a patient is discharged from a hospital, or the likelihood of readmission after receiving treatment for heart failure.

[0011]

[0019] As used herein, periods may be divided by any appropriate interval used by a healthcare professional in treating a patient. Depending on the circumstances, periods may be separated by intervals of several months or several years, but depending on the circumstances, multiple periods may fall on the same day.

[0012]

[0020] As used herein, "sound" includes any sound produced by the human vocal tract, and these sounds do not necessarily have to include sounds intended as intelligible speech or spoken language. For example, "sound" may include sighs, breaths, or grunts.

[0013]

[0021] Figures 1A–1D show exemplary architectures for processing speech using mathematical models to diagnose health conditions. The mathematical models in Figures 1A–1D can include any suitable mathematical model, such as a neural network.

[0014]

[0022] Figure 1A shows an exemplary system 100 for processing speech using a mathematical model component 110 to determine a health status label. The health status label may include any appropriate label related to a medical diagnosis, such as a Boolean value (indicating whether a person has a condition or not), selected from a set of labels (e.g., "mild," "moderate," or "severe"), an integer value (e.g., on a scale of 1 to 10), or a floating-point value (e.g., a temperature of 98.6 degrees).

[0015]

[0023] FIG. 1B is an exemplary system 102 for processing a first voice from a first period and a second voice from a second period using a mathematical model component 112 to determine a change in health state between the first period and the second period. The change in health state may be any suitable value that can be used to indicate the change in health state, such as a boolean value (indicating the presence or absence of a change), an integer value, or a floating-point value.

[0016]

[0024] FIG. 1C is an exemplary system 104 for processing a first voice from a first period, a first health state label from the first period, and a second voice from a second period using a mathematical model component 114 to determine a second health state label for the second period. The first health state label may be determined using any suitable technique, such as being determined by a person or a mathematical model. The first and second health state labels may include any of the labels described herein.

[0017]

[0025] FIG. 1D is an exemplary system 106 for processing a plurality of previous pairs of voice and health state labels from previous periods and a current voice sample from the current time using a mathematical model component 116 to determine a health state label for the current time. The example of FIG. 1D shows N previous pairs of voice and health state labels, where N may be any number greater than 1. The previous health state labels may be determined using any suitable technique, such as being determined by a person or a mathematical model. The previous and current health state labels may include any of the labels described herein. The current period may include any suitable period for which it is desired to calculate a health state label, and the processing of system 106 need not be performed at the time when the current voice is received.

[0018]

[0026] Here, further details of the embodiments of FIGS. 1A - 1D are described.

[0027] FIG. 2 is an exemplary system 200 for processing a first voice from a first period and a second voice from a second period using a mathematical model to determine a change in health state between the first period and the second period.

[0019]

[0028] In FIG. 2, the first voice is processed by a feature extraction component 210 to calculate a first feature vector (or optionally a first sequence of feature vectors), and the second voice is processed by a feature extraction component 212 to calculate a second feature vector (or optionally a second sequence of feature vectors). The feature extraction component 210 and the feature extraction component 212 can calculate the same type of features or different types of features. The feature vector can include any suitable type of features, including but not limited to the features described in U.S. Patent No. 10,152,988, which is incorporated herein by reference.

[0020]

[0029] The features can include acoustic features, which are any features calculated from audio data that do not involve or depend on performing speech recognition on the audio data (e.g., acoustic features do not use information about the words spoken in the audio data). For example, acoustic features can include Mel-frequency cepstral coefficients, perceptual linear prediction features, Wav2Vec features, prosodic features (such as pitch, energy, or probability of voicing), voice quality features (such as jitter, jitter of jitter, shimmer, or harmonic-to-noise ratio), or entropy.

[0021]

[0030] Features can include linguistic features, which are computed using recognized text obtained by automatic speech recognition. For example, linguistic features may include the words spoken in the speech, the speaking speed (e.g., the number of vowels or syllables per second), the number of pause fillers (e.g., "ums" and "ahs"), the difficulty of the words (e.g., less common words), or the part of speech of the words following the pause fillers. In some embodiments, linguistic features may include a determination of whether a person answered a question correctly. For example, a person might be asked what year it is or who the US president is. The person's speech can also be processed to determine what the person said in response to the question and whether the person answered the question correctly.

[0022]

[0031] In some embodiments, feature extraction components 210 and 212 may perform speech recognition to obtain text corresponding to speech, and then output the tokenized text as features such as wordpiece coding, byte pair coding, or sentencepiece coding. The tokenized text may be combined with any of the other features described herein.

[0023]

[0032] Next, a pair of mathematical models can process the feature vectors. The speech embedding component 220 can process the first feature vector computed by the feature extraction component 210 and compute the first speech embedding vector. Similarly, the speech embedding component 222 can process the second feature vector computed by the feature extraction component 212 and compute the second speech embedding vector.

[0024]

[0033] As used herein, a speech embedding vector is a representation of a corresponding speech in a vector space, and the position of the speech embedding vector in the vector space corresponds to information, properties, or other aspects of the speech. For example, in some embodiments, the position of the speech embedding vector can correspond to the meaning of words in the speech such that speech embedding vectors of speeches with similar meanings are close to each other in the vector space (e.g., "hello" and "good morning").

[0025]

[0034] The voice embedding component 220 and the voice embedding component 222 may have the same architecture and parameters, the same architecture and different parameters, or different architectures and different parameters.

[0026]

[0035] The speech embedding components 220 and 222 may be implemented using any suitable technique, such as transformer neural networks (Bidirectional Encoder Representation from Transformer, i.e., BERT neural networks), fully connected neural networks (e.g., multilayer perceptrons), recurrent neural networks, convolutional neural networks, or any combination of the aforementioned neural networks. In some embodiments, the speech embedding components 220 and 222 may include one or more feedforward neural network layers and one or more self-aware neural network layers.

[0027]

[0036] The mathematical model component 240 processes the first speech embedding vector computed by the speech embedding component 220 and the second speech embedding vector computed by the speech embedding component 222, and computes change values ​​indicating changes in health status, such as any of the health status change values ​​described herein. In some embodiments, the mathematical model component 240 can concatenate the first speech embedding vector with the second speech embedding vector and process the concatenated vector using a mathematical model. The mathematical model component 240 may be implemented using any suitable mathematical model, such as a linear model (e.g., matrix-vector multiplication or dot product) or a neural network (e.g., a fully connected neural network, a feedforward fully connected neural network, a multilayer perceptron, a transformer neural network, a recurrent neural network, a convolutional neural network, or any other neural network described herein).

[0028]

[0037] In some embodiments, the mathematical model component 240 may be implemented using a transformer neural network, such as a BERT neural network. For example, first and second speech coding vectors and a first label may be concatenated to form the input vector of the transformer neural network. The change values ​​may be elements of the output vector of the transformer neural network, or the output of the transformer neural network may be followed by one or more layers (e.g., linear layers) for calculating the change values ​​from the output of the transformer neural network.

[0029]

[0038] In some embodiments, the mathematical model component 240 may be implemented using a recurrent neural network. For example, the first and second speech coding vectors may be processed sequentially by the recurrent neural network (in any suitable order, and optionally with separator tokens). The change values ​​may be elements of the output vector of the recurrent neural network, or the output of the recurrent neural network may be followed by one or more layers (e.g., linear layers) for calculating the change values ​​from the output of the recurrent neural network.

[0030]

[0039] Figure 3 shows an exemplary system 300 for processing audio from two time periods using a mathematical model and determining health status change values ​​using element-wise differences. In Figure 3, feature extraction component 210, feature extraction component 212, audio embedding component 220, and audio embedding component 222 may be implemented as described above.

[0031]

[0040] The difference component 330 receives the first audio embedding from the audio embedding component 220 and the second audio embedding from the audio embedding component 222, and calculates a difference vector, which is the element-wise difference between the two audio embedding vectors. For example, if the first element of the first audio embedding vector is "a" and the first element of the second audio embedding vector is "b", then the first element of the difference vector is "ab".

[0032]

[0041] The mathematical model component 340 processes the difference vector calculated by the difference component 330 and calculates change values ​​indicating changes in health status, such as any of the health status change values ​​described herein. The mathematical model component 340 may be implemented using any suitable technique, such as any of the techniques described above for the mathematical model component 240.

[0033]

[0042] In some embodiments, the mathematical model component 340 can compute health change values ​​that are antisymmetric given an input, meaning that if the audio input is swapped, the output health change will be the same magnitude but with the opposite sign (for example, the health change switches from +3 to -3). For example, if the mathematical model component 340 computes the health change by computing the dot product of the difference vector and the parameter vector, the calculation of the health change is antisymmetric.

[0034]

[0043] Figure 4 shows an exemplary system 400 for using a mathematical model to process a first voice and a first health status label from a first period, and a second voice from a second period, to determine a second health status label for the second period. The first and second health status labels may include any appropriate labels, such as any of the labels described herein.

[0035]

[0044] In Figure 4, the feature extraction component 210, the feature extraction component 212, the speech embedding component 220, and the speech embedding component 222 may be implemented as described above.

[0036]

[0045] The mathematical model component 440 processes a first speech embedding vector computed by the speech embedding component 220, a second speech embedding vector computed by the speech embedding component 222, a first health status label corresponding to a first period, and computes a second health status label corresponding to a second period. The first health status label may be concatenated with the first and second speech embedding vectors using any suitable technique such as concatenation (e.g., by a transformer neural network) or sequential processing (e.g., by a recurrent neural network). The mathematical model component 440 may be implemented using any suitable technique, such as any of the techniques described above for the mathematical model component 240.

[0037]

[0046] In some embodiments, the mathematical model component 440 may implement regression techniques such as linear regression, nonlinear regression, multiple regression, multivariate regression, semiparametric regression, or any combination of nonparametric regressions (e.g., using nearest neighbors, regression trees, kernel regression, local regression, multivariate adaptive regression splines, neural networks, support vector regression, or smoothing splines).

[0038]

[0047] In some embodiments, the mathematical model component 440 can compute a difference vector as an element-wise difference between a first speech embedding and a second speech embedding, and then use the difference vector to compute a value indicating the change in health between the first and second periods. The mathematical model component 440 can then compute a second health label for the second period using the first label and the value indicating the change in health. For example, the second health label may be computed by adding the first health label and the change.

[0039]

[0048] Figure 5 shows an exemplary system 500 for determining the current health status label for the current period by processing multiple pairs of previous audio and health status labels from previous periods and the current audio from the current period using a mathematical model. The health status label may include any appropriate label, such as any of the labels described herein.

[0040]

[0049] Figure 5 shows N voice inputs and N labels, where N corresponds to any number greater than 1. The N voice inputs and labels may be obtained from the patient's medical records and may correspond to the patient's previous consultations at different time periods.

[0041]

[0050] The N+1th voice input may be any voice input for which a health status label is desired, and the N+1th voice input may correspond to the patient's most recent consultation or the current time. System 500 processes the N pairs of voice inputs and labels, along with the N+1th voice input, to calculate the N+1th health status label.

[0042]

[0051] In Figure 5, feature extraction components 210, 212, 214, and 216 may be implemented using any of the feature extraction techniques described herein. Each instance of a feature extraction component can compute features of the same type or different types.

[0043]

[0052] Speech embedding components 220, 222, 224, and 226 can compute speech embedding vectors from feature vectors using any of the techniques described herein. Each instance of a speech embedding component may use the same or different neural network architectures and parameters.

[0044]

[0053] The mathematical model component 540 processes N labels and N+1 speech embedding vectors to compute the N+1 health status label for the N+1 speech input. The mathematical model component 540 may be implemented using any suitable technique. For example, the mathematical model component 540 may be implemented using either of the techniques described above for the mathematical model component 240 or the mathematical model component 440, and these techniques may be adapted to additional input values ​​using techniques known to those skilled in the art.

[0045]

[0054] Figure 6 is a flowchart illustrating an exemplary method for processing audio from two time periods using a mathematical model and determining changes in health status using element-wise differences.

[0055] In step 610, a first audio signal corresponding to a first period is received, the first audio signal containing the voice of a first person. The first audio signal may be received using any appropriate technique, such as via an API call or by retrieving it from storage.

[0046]

[0056] In step 620, a first feature vector is calculated from the first audio signal. The first feature vector may include any suitable features, such as any of the features described herein. In some embodiments, the features may include wordpiece coding corresponding to the text of speech in the audio signal determined using automatic speech recognition. In some embodiments, the features may include acoustic features.

[0047]

[0057] In step 630, a first speech embedding vector is computed by processing the first feature vector using a neural network. The neural network may be any suitable neural network, such as any of the neural networks described herein. In some embodiments, the neural network may include a transformer neural network. In some embodiments, the neural network may include one or more feedforward neural network layers and one or more self-aware neural network layers.

[0048]

[0058] In step 640, a second audio signal corresponding to a second period is received, the second audio signal including the voice of a first person. The first audio signal may be received as described above for step 610.

[0049]

[0059] In step 650, a second feature vector is calculated from the second audio signal. The second feature vector may be calculated as described above for step 620.

[0050]

[0060] In step 660, a second speech embedding vector is computed by processing the second feature vector using a neural network. The second speech embedding may be computed as described above for step 630.

[0051]

[0061] In step 670, an element-wise difference vector is calculated between the first audio embedding vector and the second audio embedding vector. The element-wise difference vector may be calculated using any of the techniques described herein.

[0052]

[0062] In step 680, a change value representing the change in health status is calculated by processing element-wise difference vectors using a mathematical model. The change value can represent the change in health status between a first period and a second period. The mathematical model may be any suitable mathematical model, such as any of the mathematical models described herein. In some embodiments, the change value may be an antisymmetric change value. In some embodiments, the mathematical model may use element-wise difference vectors to calculate an inner product or matrix-vector multiplication.

[0053]

[0063] In some embodiments, some of steps 610–630 may be performed in advance, and the first feature vector or first speech embedding may be stored for later use. For example, some of steps 610–630 may be performed immediately after the first consultation in which the first audio sample is taken. The first feature vector or first speech embedding may be stored so that it can be used when the person has a subsequent consultation. Steps 640–680 may be performed after a subsequent consultation, which may be several days, weeks, months, or years after the first consultation.

[0054]

[0064] In some embodiments, the steps in Figure 6 may be performed for multiple previous examinations in order to calculate multiple change values. For example, the third audio signal may be obtained from another previous examination. The second change value may be calculated using the third audio signal, the third feature vector, the third speech embedding vector, and a second element-wise difference vector calculated using the second and third speech embedding vectors.

[0055]

[0065] In some embodiments, the first health status label may be obtained in correspondence with a first period. The first health status label may be determined using any suitable technique, such as any of the techniques described herein. The second health status label for a second period may then be calculated using the first health status label and the change value. The second health status label may be calculated using any suitable technique, such as by adding the first health status label and the change value.

[0056]

[0066] Figure 7 is a flowchart illustrating an exemplary method for using a mathematical model to process a first voice and first health status label from a first period, and a second voice from a second period, to determine a second health status label for the second period.

[0057]

[0067] In Figure 7, steps 710 to 760 may be carried out as described above for steps 610 to 660 in Figure 6.

[0068] In step 770, a first health status label corresponding to a first period is acquired. The first health status label may be any of the health status labels described herein and may be acquired using any suitable technique. For example, the first health status label may be stored together with a first audio signal (or a first feature vector or a first audio embedding vector).

[0058]

[0069] In step 780, a mathematical model is used to process the first health status label and the first and second speech embedding vectors to compute a second health status label corresponding to the second period. The mathematical model may be any suitable mathematical model, such as a neural network. In some embodiments, the mathematical model may compute the second health status label using linear or nonlinear regression techniques.

[0059]

[0070] For any of the techniques described herein, the parameters of a mathematical model (including a neural network) may be learned or trained using a training process. Any suitable training process may be used, such as supervised or unsupervised training. The training process may include a training corpus of audio files, which may include training labels indicating health status labels corresponding to the audio files, or training labels indicating change values ​​corresponding to pairs of audio files. The parameters of the mathematical model may be learned through an iterative training process. For example, the training process may include a forward pass that processes the training data to compute predicted values ​​(e.g., health status labels or change values), error values ​​may be computed using the predicted values ​​and training labels, and a backward pass may be performed to update the mathematical model parameters using the error values ​​(e.g., using stochastic gradient descent). The training process may continue until a desired convergence criterion is obtained.

[0060]

[0071] Figure 8 shows components of one embodiment of a computing device 800 for implementing any of the technologies described herein. In Figure 8, the components are shown as being on a single computing device, but the components may be distributed across multiple computing devices, such as a system of computing devices including, for example, end-user computing devices (e.g., smartphones or tablets) and / or server computers (e.g., cloud computing).

[0061]

[0072] The computing device 800 may include any components typical of a computing device, such as volatile or non-volatile memory 810, one or more processors 811, and one or more network interfaces 812. The computing device 800 may also include any input and output components, such as a display, keyboard, and touchscreen. The computing device 800 may also include various components or modules that provide specific functions, and these components or modules may be implemented in software, hardware, or a combination thereof. The computing device 800 may include one or more non-temporary computer-readable media that, when executed, contain computer-executable instructions that cause the processor to perform operations corresponding to any of the technologies described herein. Some examples of components are described below for one exemplary embodiment, and other embodiments may include additional components or omit some of the components described below.

[0062]

[0073] The computing device 800 may have a feature extraction component 820 that can compute a feature vector from an audio signal using any of the techniques described herein. The computing device 800 may have a speech embedding component 821 that can compute a speech embedding vector from a feature vector using any of the techniques described herein. The computing device 800 may have a mathematical modeling component 822 that can compute a health status label or a health status change value using any of the techniques described herein. The computing device 800 may have an element-wise difference component 823 that can compute the element-wise difference between two vectors using any of the techniques described herein.

[0063]

[0074] The computing device 800 may include, or access, various data stores. The data stores may be any known storage technology, such as files, relational databases, non-relational databases, or any non-temporary computer-readable media. The computing device 800 may have a training data store 830 for storing training data that can be used to train any of the mathematical models described herein.

[0064]

[0075] The methods and systems described herein may be deployed in part or in whole through computer software, program code, and / or machines that execute instructions on a processor. As used herein, “processor” means including at least one processor, and unless the context explicitly indicates otherwise, the plural and singular forms should be understood to be interchangeable. Any aspect of this disclosure may be implemented as a computer implementation method on a machine, as a system or device as part of or related to a machine, or as a computer program product embodied in a computer-readable medium running on one or more machines. A processor may be part of a server, client, network infrastructure, mobile computing platform, stationary computing platform, or other computing platform. A processor may be any type of computing or processing device capable of executing program instructions, code, binary instructions, etc. A processor may be any variation of a signal processor, digital processor, embedded processor, microprocessor, or a coprocessor (such as a numerical coprocessor, graphics coprocessor, or communications coprocessor) that can directly or indirectly facilitate the execution of stored program code or program instructions, or may include such variations. In addition, a processor may enable the execution of multiple programs, threads, and code. Threads may run concurrently to improve processor performance and facilitate the simultaneous operation of applications. In embodiments, the methods, program code, program instructions, etc., described herein may be implemented with one or more threads. Threads can generate other threads, which may have assigned priorities associated with them, and the processor can execute these threads based on priorities based on instructions provided in the program code, or on any other order.A processor may include memory for storing methods, code, instructions, and programs described herein and elsewhere. A processor may have access to a storage medium via an interface that can store methods, code, and instructions described herein and elsewhere. A storage medium associated with a processor for storing methods, programs, code, program instructions, or other types of instructions executable by a computing or processing unit may include, but is not limited to, one or more of the following: CD-ROMs, DVDs, memory, hard disks, flash drives, RAM, ROMs, caches, etc.

[0065]

[0076] The processor may include one or more cores that can improve the speed and performance of the multiprocessor. In some embodiments, the processor may be a dual-core processor, a quad-core processor, or other chip-level multiprocessors that combine two or more independent cores (called dies).

[0066]

[0077] The methods and systems described herein may be deployed in part or whole through a server, client, firewall, gateway, hub, router, or other machine running computer software on such computer and / or networking hardware. The software program may be associated with a server, which may include file servers, print servers, domain servers, internet servers, intranet servers, and other variations such as secondary servers, host servers, and distributed servers. A server may include one or more of the following: memory, processors, computer-readable media, storage media, ports (physical and virtual), communication devices, and interfaces that allow access to other servers, clients, machines, and devices via wired or wireless media. The methods, programs, or code described herein and elsewhere may be executed by a server. In addition, other equipment necessary for executing the methods described herein may be considered part of the infrastructure associated with the server.

[0067]

[0078] The server can provide interfaces to other devices, including but not limited to clients, other servers, printers, database servers, print servers, file servers, communication servers, and distributed servers. Furthermore, this coupling and / or connection can facilitate the remote execution of programs over a network. Networking some or all of these devices can facilitate parallel processing of programs or methods at one or more locations without departing from the scope of this disclosure. In addition, any device attached to the server via an interface may include at least one storage medium capable of storing methods, programs, code, and / or instructions. A central repository can provide program instructions to be executed on different devices. In this embodiment, a remote repository can serve as a storage medium for program code, instructions, and programs.

[0068]

[0079] The software program may be associated with a client that includes file clients, print clients, domain clients, internet clients, intranet clients, and other variations such as secondary clients, host clients, and distributed clients. The client may include one or more of the following: memory, processors, computer-readable media, storage media, ports (physical and virtual), communication devices, and interfaces that can access other clients, servers, machines, and devices via wired or wireless media. The methods, programs, or code described herein and elsewhere may be executed by the client. In addition, other devices necessary for executing the methods described herein may be considered part of the infrastructure associated with the client.

[0069]

[0080] The client can provide interfaces to other devices, including but not limited to servers, other clients, printers, database servers, print servers, file servers, communication servers, and distributed servers. Furthermore, this coupling and / or connection can facilitate the remote execution of programs over a network. Networking some or all of these devices can facilitate parallel processing of programs or methods at one or more locations without departing from the scope of this disclosure. In addition, any device attached to the client via the interface may include at least one storage medium capable of storing methods, programs, applications, code, and / or instructions. A central repository can provide program instructions to be executed on different devices. In this embodiment, a remote repository can serve as a storage medium for program code, instructions, and programs.

[0070]

[0081] The methods and systems described herein may be deployed in part or entirely through a network infrastructure. The network infrastructure may include elements such as computing devices, servers, routers, hubs, firewalls, clients, personal computers, communication devices, routing devices, and other active and passive devices, modules, and / or components known in the art. Computing devices and / or non-computing devices associated with the network infrastructure may include storage media such as flash memory, buffers, stacks, RAM, and ROM, apart from other components. The processes, methods, program code, and instructions described herein and elsewhere may be executed by one or more of the network infrastructure elements.

[0071]

[0082] The methods, program code, and instructions described herein and elsewhere may be implemented on a cellular network having multiple cells. The cellular network may be either a frequency division multiple access (FDMA) network or a code division multiple access (CDMA) network. The cellular network may include mobile devices, cell sites, base stations, repeaters, antennas, towers, etc. The cell network may be GSM, GPRS, 3G, EVDO, mesh, or other network types.

[0072]

[0083] The methods, program code, and instructions described herein and elsewhere may be implemented on or via a mobile device. A mobile device may include navigation devices, cell phones, mobile phones, mobile personal digital assistants, laptops, palmtops, netbooks, pagers, e-book readers, music players, and the like. These devices may include, apart from other components, storage media such as flash memory, buffers, RAM, ROM, and one or more computing devices. A computing device associated with a mobile device may be configured to execute program code, methods, and instructions stored therein. Alternatively, a mobile device may be configured to execute instructions in cooperation with other devices. A mobile device may communicate with a base station interfaced with a server and configured to execute program code. A mobile device may communicate over a peer-to-peer network, a mesh network, or other communication network. The program code may be stored on storage media associated with a server and executed by a computing device embedded within the server. A base station may include a computing device and storage media. The storage device may store program code and instructions executed by a computing device associated with the base station.

[0073]

[0084] Computer software, program code, and / or instructions may be stored and / or accessed on machine-readable media, which may include computer components, devices, and recording media that hold digital data used for computing for a certain period of time, such as semiconductor storage known as random access memory (RAM), mass storage typically for more permanent storage, such as optical discs, hard disks, tapes, drums, cards and other types of magnetic storage, processor registers, cache memory, volatile memory, non-volatile memory, optical storage such as CDs and DVDs, flash memory (e.g., USB sticks or keys), floppy disks, magnetic tape, paper tape, punch cards, standalone RAM disks, Zip drives, removable mass storage, removable media such as offline, dynamic memory, static memory, read / write storage, variable storage, read-only, random access, sequential access, location addressable, file addressable, content addressable, network-attached storage, storage area networks, barcodes, magnetic ink and other computer memory.

[0074]

[0085] The methods and systems described herein can transform physical and / or intangible items from one state to another. The methods and systems described herein can also transform data representing physical and / or intangible items from one state to another.

[0075]

[0086] Elements described and depicted herein, including flowcharts and block diagrams throughout the drawings, represent logical boundaries between elements. However, in accordance with the conventions of software or hardware engineering, the depicted elements and their functions may be implemented on a machine through a computer executable medium having a processor capable of executing stored program instructions, as a monolithic software structure, as a standalone software module, or as a module employing external routines, code, services, etc., or any combination thereof, and all such embodiments may be within the scope of this disclosure. Examples of such machines include, but are not limited to, personal information terminals, laptops, personal computers, mobile phones, other handheld computing devices, medical devices, wired or wireless communication devices, transducers, chips, calculators, satellites, tablet PCs, e-books, gadgets, electronic devices, devices with artificial intelligence, computing equipment, networking equipment, servers, routers, etc. Furthermore, elements shown in flowcharts and block diagrams or any other logical components may be implemented on a machine capable of executing program instructions. Accordingly, while the aforementioned drawings and descriptions illustrate functional aspects of the disclosed system, specific software configurations for implementing these functional aspects should not be inferred from these descriptions unless explicitly stated or otherwise evident from the context. Similarly, it will be understood that the various steps identified and described above may be modified, and the order of the steps may be adapted to specific uses of the technology disclosed herein. All such variations and modifications are intended to fall within the scope of this disclosure. Therefore, the depiction and / or description of the order of the various steps should not be understood as requiring a specific order of execution of those steps unless required by a particular use, or unless explicitly stated or otherwise evident from the context.

[0076]

[0087] The methods and / or processes described above, and their steps, may be implemented in hardware, software, or any combination of hardware and software suitable for a particular application. Hardware may include general-purpose computers and / or dedicated computing devices or specific computing devices, or specific embodiments or components of specific computing devices. Processes may be implemented in one or more microprocessors, microcontrollers, embedded microcontrollers, programmable digital signal processors, or other programmable devices, along with internal and / or external memory. Processes may also be implemented in application-specific integrated circuits, programmable gate arrays, programmable array logic, or any other devices or combinations of devices that can be configured to process electronic signals. Furthermore, it will be understood that one or more of the processes may be implemented as computer executable code that can be run on machine-readable media.

[0077]

[0088] Computer executable code may be written using structured programming languages ​​such as C, object-oriented programming languages ​​such as C++, or any other high-level or low-level programming languages ​​(including assembly languages, hardware description languages, and database programming languages ​​and techniques), which may be stored, compiled, or interpreted to run on one of the above devices, as well as heterogeneous combinations of processors, processor architectures, or different hardware and software combinations, or any other machine capable of executing program instructions.

[0078]

[0089] Accordingly, in one embodiment, each of the methods and combinations thereof described above may be embodied in computer executable code that performs the step when executed on one or more computing devices. In another embodiment, the method may be embodied in a system that performs the step, may be distributed across devices in several ways, or all of the functionality may be integrated into a dedicated standalone device or other hardware. In yet another embodiment, the means for performing the steps associated with the processes described above may include any of the hardware and / or software described above. All such permutations and combinations are intended to fall within the scope of this disclosure.

[0079]

[0090] Although the present invention is disclosed in relation to preferred embodiments illustrated and described in detail, various modifications and improvements thereto will be readily apparent to those skilled in the art. Therefore, the spirit and scope of the invention should not be limited by the examples described herein, but should be understood in the broadest sense permitted by law.

[0080]

[0091] All documents referenced herein are incorporated herein in their entirety by reference. [Explanation of Symbols]

[0081] 100 Systems 102 System 104 System 106 System 110 Mathematical Model Components 112 Mathematical Model Components 114 Mathematical Model Components 116 Mathematical Model Components 200 Systems 210 Feature Extraction Components 212 Feature Extraction Components 214 Feature Extraction Components 216 Feature Extraction Components 220 Audio Embedding Components 222 Audio Embedding Component 224 Audio Embedding Components 226 Audio Embedding Components 240 Mathematical Model Components 300 Systems 330 Differential Components 340 Mathematical Model Components 400 System 440 Mathematical Model Components 500 Systems 540 Mathematical Model Components 800 Computing Devices 810 Volatile or non-volatile memory 811 Processor 812 Network Interfaces 820 Feature Extraction Components 821 Audio Embedding Component 822 Mathematical Model Components 823 Difference components per element 830 Training Data Store

Claims

1. A computer implementation method, Using at least one first model, determine a second speech embedding vector for a second audio signal containing human speech during a second period, Using at least one second model, determine whether a change in the health status of the person occurred between a first period and a second period, at least partially based on (i) a first speech embedding vector corresponding to a first audio signal including the person's voice during the first period and (ii) the second speech embedding vector corresponding to the second audio signal. Computer implementation methods including

2. The computer implementation method according to claim 1, wherein the at least one first model includes a neural network, and the at least one second model includes a mathematical model.

3. The computer implementation method according to claim 1, wherein determining the change in health status includes calculating a change value indicating the change in health status.

4. Determining the change in health status means Calculate the element-wise difference between the first audio embedding vector and the second audio embedding vector, Using the aforementioned at least one second model, the difference for each element is processed. The computer implementation method according to claim 1, including the method described in claim 1.

5. The computer implementation method is Obtaining a first health status label indicating the health status during the first period, A second health status label indicating the health status during the second period is calculated by processing the first health status label and at least one value indicating the change between the first voice embedding vector and the second voice embedding vector. The computer implementation method according to claim 1, further comprising:

6. The computer implementation method according to claim 5, wherein calculating the second health status label includes adding the first health status label and the at least one value indicating the change.

7. Determining the second speech embedding vector comprises generating a feature vector, which comprises (i) obtaining recognized text by performing speech recognition on the second audio signal, and (ii) obtaining wordpiece coding corresponding to the recognized text. The computer implementation method according to claim 1, wherein the at least one first model includes a plurality of feedforward neural network layers and a plurality of self-aware neural network layers.

8. The computer implementation method according to claim 1, wherein the at least one second model includes a neural network.

9. The computer implementation method according to claim 1, wherein the at least one second model includes a fully connected neural network.

10. The computer implementation method according to claim 1, wherein the health status corresponds to stress, depression, anxiety, post-traumatic stress disorder, concussion, Parkinson's disease, Alzheimer's disease, or congestive heart failure.

11. The computer implementation method further includes calculating the element-wise difference between the first audio embedding vector and the second audio embedding vector, The computer implementation method according to claim 1, wherein determining whether the change in health status has occurred includes calculating at least one value indicating a change between the first voice embedding vector and the second voice embedding vector, and calculating the at least one value includes at least partially calculating an antisymmetric change value by processing the element-wise difference using the at least one second model.

12. A system comprising at least one computer, The aforementioned at least one computer is Using at least one first model, determine a second speech embedding vector for a second audio signal containing human speech during a second period, Using at least one second model, determine whether a change in the health status of the person occurred between a first period and a second period, at least in part on (i) a first speech embedding vector corresponding to a first audio signal including the person's voice during the first period and (ii) the second speech embedding vector corresponding to the second audio signal. A system configured to perform the following actions.

13. The system according to claim 12, wherein the at least one first model includes a neural network.

14. The system according to claim 12, wherein the at least one second model includes a mathematical model.

15. The at least one computer is Using the at least one first model, determine a third speech embedding vector for a third audio signal including the human voice during a third period. Using the at least one second model, it is determined whether a change in the health status of the person occurred between the first period and the third period, at least partially based on (i) the first speech embedding vector, (ii) the second speech embedding vector, and (iii) the third speech embedding vector. The system according to claim 12, further configured to perform the following:

16. The system according to claim 12, wherein determining the change in health status includes calculating a change value indicating the change in health status.

17. The system according to claim 12, wherein determining the second speech embedding vector comprises generating a feature vector, the generating of the feature vector comprising (i) obtaining recognized text by performing speech recognition on the second audio signal, and (ii) obtaining wordpiece coding corresponding to the recognized text.

18. One or more computer-readable media comprising computer-executable instructions, wherein, when the computer-executable instructions are executed, Using at least one first model, determine a second speech embedding vector for a second audio signal containing human speech during a second period, Using at least one second model, determine whether a change in the health status of the person occurred between a first period and a second period, at least partially based on (i) a first speech embedding vector corresponding to a first audio signal including the person's voice during the first period and (ii) the second speech embedding vector corresponding to the second audio signal. One or more computer-readable media that cause at least one processor to perform an operation including the following.

19. One or more computer-readable media according to claim 18, wherein the at least one first model includes a neural network, and the at least one second model includes a mathematical model.

20. The method for determining the change in health status includes calculating a change value that indicates the change in health status, for one or more computer-readable media according to claim 18.

21. The one or more computer-readable media according to claim 18, wherein determining the second speech embedding vector comprises generating a feature vector, the generating of the feature vector comprising (i) performing speech recognition on the second audio signal to obtain recognized text, and (ii) obtaining wordpiece coding corresponding to the recognized text.

Citation Information

Patent Citations

  • Selecting speech features for building models for detecting medical conditions

    US10152988B2