Systems and methods for machine learning voice authentication for improved medical device security
A machine learning-based voice authentication system for renal care devices enables secure, touch-free operation, addressing the challenges of manual control and unauthorized access, enhancing user comfort and operational efficiency.
Patent Information
- Application Number
- PCT/US2025/016462
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-12
- Filing Date
- 2025-02-19
- Publication Date
- 2025-12-18
AI Technical Summary
Medical devices, particularly renal care systems, require manual operation and lack secure voice control, leading to user discomfort and unauthorized access.
Implementing a machine learning-based voice authentication system that uses trained voice recognition models to authenticate users and execute commands, enhancing security and accessibility by allowing voice-controlled operation.
The system provides secure, touch-free control, preventing unauthorized access and improving user comfort and healthcare professional productivity by leveraging voice recognition technology.
Smart Images

Figure US2025016462_18122025_PF_FP_ABST
Abstract
Description
SYSTEMS AND METHODS FOR MACHINE LEARNING VOICE AUTHENTICATION FOR IMPROVED MEDICAL DEVICE SECURITYCLAIM TO PRIORITY
[0001] This application claims priority to India Provisional Application 202411045352, titled “SYSTEMS AND METHODS FOR MACHINE LEARNING VOICE AUTHENTICATION FOR IMPROVED MEDICAL DEVICE SECURITY” and filed June 12, 2024.FIELD OF TECHNOLOGY
[0002] The present disclosure generally relates to computer-based systems and methods for machine learning (ML) based voice authentication and recognition for secured voice control of systems and / or devices, including medical devices such as renal care or other devices.BACKGROUND OF TECHNOLOGY
[0003] Medical devices, such as renal care systems and devices, as well as other clinical and / or at-home devices are typically operated by hand via physical controls and / or screen-based interfaces. As a result, a user may have to manually control the device(s) while undergoing therapy. Moreover, people other than the user may attempt to use the device without authorization.SU MARY
[0004] In some aspects, the techniques described herein relate to a system including: at least one renal care device configured to provide at least one renal therapy to a user; and a controller in communication with the at least one renal care device, the controller configured to: receive a voice signal from an audio source; input the voice signal into at least one trained voice recognition machine learning model to output a voice authentication determination and a device command based at least in part on trained parameters of the at least one trained voice recognition model; wherein the trained parameters of the at least one trained voice recognition model are configured to determine a measure of a match between the voice signal and at least one training vocal pattern of a training vocal pattern dataset; and instruct the at least one renal care device with the device command based at least in part on the measure of the match of the voice signal exceeding a predetermined threshold.
[0005] In some aspects, the techniques described herein relate to a system, wherein the at least one trained voice recognition machine learning model includes a single voice recognition machine learning model configured to output the voice authentication determination and the device command.
[0006] In some aspects, the techniques described herein relate to a system, wherein the at least one trained voice recognition machine learning model includes: at least one trained voice command recognition machine learning model configured to output the device command; and at least one trained voice authentication machine learning model configured to output the voice authentication determination.
[0007] In some aspects, the techniques described herein relate to a system, wherein the at least one trained voice recognition machine learning model includes at least one clustering model.
[0008] In some aspects, the techniques described herein relate to a system, wherein the at least one trained voice recognition machine learning model includes at least one neural network model.
[0009] In some aspects, the techniques described herein relate to a system, wherein the controller is further configured to: receive a first voice input associated with a user vocalizing an authentication utterance; input the first voice input into at least one trained voice authentication recognition machine learning model of the at least one trained voice recognition machine learning model to output the voice authentication determination based at least in part on trained authentication parameters; wherein the trained authentication parameters are configured to determine the voice authentication determination based at least in part on an authentication similarity measure between the first voice input and at least one training authentication vocal pattern of a training authentication vocal pattern dataset; generate, based on the authentication similarity measure exceeding a predetermined authentication similarity measure threshold, a prompt to the user for the device command; receive a second voice input associated with the user vocalizing a device command utterance; input the second voice input into at least one trained voice command recognition machine learning model of the at least one trained voice recognition machine learning model to output the device command based at least in part on trained device command parameters; and wherein the trained device command parameters are configured to determine the device command based at least in part on a device command similarity measure between the second voice input and at least one training device command vocal pattern of a training device command vocal pattern dataset.
[0010] In some aspects, the techniques described herein relate to a system, wherein the controller is further configured to: generate a voice signal pattern representative of at least one characteristic of the voice signal based at least in part on at least one pre-processing algorithm; and input the voice signal pattern into at least one trained voice recognition machine learning model to output the voice authentication determination and the device command based at least in part on the trained parameters of the at least one trained voice recognition model.
[0011] In some aspects, the techniques described herein relate to a method including: at least one renal care device configured to provide at least one renal therapy to a user; and a controller in communication with the at least one renal care device, the controller configured to: receiving, by at least one processor, a voice signal from an audio source; inputting, by the at least one processor, the voice signal into at least one trained voice recognition machine learning model to output a voice authentication determination and a device command based at least in part on trained parameters of the at least one trained voice recognition model; wherein the trained parameters of the at least one trained voice recognition model are configured to determine a measure of a match between the voice signal and at least one training vocal pattern of a training vocal pattern dataset; and instructing, by the at least one processor, the at least one renal care device with the device command based at least in part on the measure of the match of the voice signal exceeding a predetermined threshold.
[0012] In some aspects, the techniques described herein relate to a method, wherein the at least one trained voice recognition machine learning model includes a single voice recognition machine learning model configured to output the voice authentication determination and the device command.
[0013] In some aspects, the techniques described herein relate to a method, wherein the at least one trained voice recognition machine learning model includes: at least one trained voice command recognition machine learning model configured to output the device command; and at least one trained voice authentication machine learning model configured to output the voice authentication determination.
[0014] In some aspects, the techniques described herein relate to a method, wherein the at least one trained voice recognition machine learning model includes at least one clustering model.
[0015] In some aspects, the techniques described herein relate to a method, wherein the at least one trained voice recognition machine learning model includes at least one neural network model.
[0016] In some aspects, the techniques described herein relate to a method, further including: receiving, by the at least one processor, a first voice input associated with a user vocalizing an authentication utterance; inputting, by the at least one processor, the first voice input into at least one trained voice authentication recognition machine learning model of the at least one trained voice recognition machine learning model to output the voice authentication determination based at least in part on trained authentication parameters; wherein the trained authentication parameters are configured to determine the voice authentication determination based at least in part on a authentication similarity measure between the first voice input and at least one training authentication vocal pattern of a training authentication vocal pattern dataset; generate, based on the authentication similarity measure exceeding a predetermined authentication similarity measure threshold, a prompt to the user for the device command; receiving, by the at least one processor, a second voice input associated with the user vocalizing a device command utterance; inputting, by the at least one processor, the second voice input into at least one trained voice command recognition machine learning model of the at least one trained voice recognition machine learning model to output the device command based at least in part on trained device command parameters; and wherein the trained device command parameters are configured to determine the device command based at least in part on a device command similarity measure between the second voice input and at least one training device command vocal pattern of a training device command vocal pattern dataset.
[0017] In some aspects, the techniques described herein relate to a method, further including: generating, by the at least one processor, a voice signal pattern representative of at least one characteristic of the voice signal based at least in part on at least one pre-processing algorithm; and inputting, by the at least one processor, the voice signal pattern into at least one trained voice recognition machine learning model to output the voice authentication determination and the device command based at least in part on the trained parameters of the at least one trained voice recognition model.
[0018] In some aspects, the techniques described herein relate to a non-transitory computer readable medium having executable instructions stored thereon, wherein the executable instructions are configured to cause at least one processor to perform steps including: receiving a voice signal from an audio source; inputting the voice signal into at least one trained voice recognition machine learning model to output a voice authentication determination and a devicecommand based at least in part on trained parameters of the at least one trained voice recognition model; wherein the trained parameters of the at least one trained voice recognition model are configured to determine a measure of a match between the voice signal and at least one training vocal pattern of a training vocal pattern dataset; and instructing at least one renal care device with the device command based at least in part on the measure of the match of the voice signal exceeding a predetermined threshold.
[0019] In some aspects, the techniques described herein relate to a non-transitory computer readable medium, wherein the at least one trained voice recognition machine learning model includes a single voice recognition machine learning model configured to output the voice authentication determination and the device command.
[0020] In some aspects, the techniques described herein relate to a non-transitory computer readable medium, wherein the at least one trained voice recognition machine learning model includes: at least one trained voice command recognition machine learning model configured to output the device command; and at least one trained voice authentication machine learning model configured to output the voice authentication determination.
[0021] In some aspects, the techniques described herein relate to a non-transitory computer readable medium, wherein the at least one trained voice recognition machine learning model includes at least one of: at least one clustering model, or at least one neural network model.
[0022] In some aspects, the techniques described herein relate to a non-transitory computer readable medium, wherein the executable instructions are further configured to cause the at least one processor to perform steps including: receiving a first voice input associated with a user vocalizing an authentication utterance; inputting the first voice input into at least one trained voice authentication recognition machine learning model of the at least one trained voice recognition machine learning model to output the voice authentication determination based at least in part on trained authentication parameters; wherein the trained authentication parameters are configured to determine the voice authentication determination based at least in part on a authentication similarity measure between the first voice input and at least one training authentication vocal pattern of a training authentication vocal pattern dataset; generate, based on the authentication similarity measure exceeding a predetermined authentication similarity measure threshold, a prompt to the user for the device command; receiving a second voice input associated with the user vocalizing a device command utterance; inputting the second voice input into at least onetrained voice command recognition machine learning model of the at least one trained voice recognition machine learning model to output the device command based at least in part on trained device command parameters; and wherein the trained device command parameters are configured to determine the device command based at least in part on a device command similarity measure between the second voice input and at least one training device command vocal pattern of a training device command vocal pattern dataset.
[0023] In some aspects, the techniques described herein relate to a non-transitory computer readable medium, wherein the executable instructions are further configured to cause the at least one processor to perform steps including: generating a voice signal pattern representative of at least one characteristic of the voice signal based at least in part on at least one pre-processing algorithm; and inputting the voice signal pattern into at least one trained voice recognition machine learning model to output the voice authentication determination and the device command based at least in part on the trained parameters of the at least one trained voice recognition model.BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Various embodiments of the present disclosure can be further explained with reference to the attached drawings, wherein like structures are referred to by like numerals throughout the several views. The drawings shown are not necessarily to scale, with emphasis instead generally being placed upon illustrating the principles of the present disclosure. Therefore, specific structural and functional details disclosed herein are not to be interpreted as limiting, but merely as a representative basis for teaching one skilled in the art to variously employ one or more illustrative embodiments.
[0025] FIG. 1 depicts a block diagram of a controller having ML-based voice authentication for improved security in controlling a renal care system in accordance with aspects of one or more embodiments of the present disclosure.
[0026] FIG. 2 depicts a block diagram for training an ML model for ML-based voice authentication for improved security in controlling a renal care system in accordance with aspects of one or more embodiments of the present disclosure.
[0027] FIG. 3 depicts a block diagram for executing an ML model, trained in accordance with FIG. 2, for ML-based voice authentication for improved security in controlling a renal care system in accordance with aspects of one or more embodiments of the present disclosure.
[0028] FIG. 4 depicts a flow chart for ML-based voice authentication for improved security in controlling a renal care system in accordance with aspects of one or more embodiments of the present disclosure.
[0029] FIG. 5 depicts illustrative schematics of an exemplary implementation of the cloud computing / architecture(s) in which embodiments of a system for ML-based voice authentication for controlling a renal care system may be specifically configured to operate in accordance with some embodiments of the present disclosure.
[0030] FIG. 6 depicts illustrative schematics of another exemplary implementation of the cloud computing / architecture(s) in which embodiments of a system for ML-based voice authentication for controlling a renal care system be specifically configured to operate in accordance with some embodiments of the present disclosure.DETAILED DESCRIPTION
[0031] Various detailed embodiments of the present disclosure, taken in conjunction with the accompanying FIGs., are disclosed herein; however, it is to be understood that the disclosed embodiments are merely illustrative. In addition, each of the examples given in connection with the various embodiments of the present disclosure is intended to be illustrative, and not restrictive.
[0032] Throughout the specification, the following terms take the meanings explicitly associated herein, unless the context clearly dictates otherwise. The phrases “in one embodiment” and “in some embodiments” as used herein do not necessarily refer to the same embodiment(s), though it may. Furthermore, the phrases “in another embodiment” and “in some other embodiments” as used herein do not necessarily refer to a different embodiment, although it may. Thus, as described below, various embodiments may be readily combined, without departing from the scope or spirit of the present disclosure.
[0033] In addition, the term "based on" is not exclusive and allows for being based on additional factors not described, unless the context clearly dictates otherwise. In addition, throughout the specification, the meaning of "a," "an," and "the" include plural references. The meaning of "in" includes "in" and "on."
[0034] As used herein, the terms “and” and “or” may be used interchangeably to refer to a set of items in both the conjunctive and disjunctive in order to encompass the full description of combinations and alternatives of the items. By way of example, a set of items may be listed withthe disjunctive “or”, or with the conjunction “and.” In either case, the set is to be interpreted as meaning each of the items singularly as alternatives, as well as any combination of the listed items.
[0035] FIGs. 1 through 6 illustrate systems and methods of controlling and securing control of medical devices such as renal care products, including dialysis systems for clinical and / or at-home use. The following embodiments provide technical solutions and technical improvements that overcome technical problems, drawbacks and / or deficiencies in the technical fields involving renal care systems and control system thereof, including software, hardware and / or a combination thereof.
[0036] As explained in more detail, below, technical solutions and technical improvements herein include aspects of improved accessibility and security of control of renal care systems, among other clinical and / or home medical apparatuses, that leverage hardware and / or software-based voice analysis and control to concurrently authenticate and instruct components of the renal care system(s). In so doing, the hardware and / or software for voice analysis and / or control enables improved operation of the renal care systems, including improved accessibility for control by patients undergoing care, as well as improved security beyond traditional physical and / or touchscreen based interfaces.
[0037] Renal care systems typically require users / operators to engage in physical interaction for access during therapy, which may cause discomfort. Additionally, existing systems employ sophisticated hardware to ensure security. Automatic Speech Recognition (ASR) enhances comprehensibility and accessibility of information, simultaneously fortifying system security against unauthorized access. This dual benefit not only elevates patient health outcomes but also amplifies the productivity of healthcare professionals by equipping renal care systems with new technical functionalities including increased security measures and new user interfaces.
[0038] The Voice Al (Artificial Intelligence) feature serves as a means of enhancing device security by preventing unauthorized access. A "dataset" may be trained using a diverse range of voice patterns and languages to enable it to carry out specific functions within the device. Once the pattern receives input (voice signal and language), the Voice Al compares the input with the “dataset” and makes the decision regarding whether the received input is from the authorized person or not. If the machine learning algorithm confirms the input voice authenticity, the device proceeds to execute the corresponding function.
[0039] Incorporating the Voice AI-ML algorithm into a renal care system enhances security by preventing unauthorized access and enable users to communicate in any language. Further, the Voice AI-ML algorithm may have additional intelligence for identifying mimic voices attempting to impersonate the user. Hence, unauthorized access will be restricted. This integration of the Voice AI-ML algorithm in the renal care system also introduces a contactless mode of operation.
[0040] Based on such technical features, further technical benefits become available to users and operators of these systems and methods. Moreover, various practical applications of the disclosed technology are also described, which provide further practical benefits to users and operators that are also new and useful improvements in the art.
[0041] Referring to FIG. 1, a controller 110 having ML-based voice authentication for improved security in controlling a renal care system 200 is illustrated in accordance with aspects of one or more embodiments of the present disclosure.
[0042] In some embodiments, a renal care system 200 may be controlled by a controller 110 that is equipped with one or more speech machine learning (ML) model(s) of a speech processing pipeline 112 to provide secure, touch-free control of the renal care system 200 in home and / or clinical use. To do so, the speech processing pipeline 112 may include trained model parameters configured to authenticate a speaker based on learned vocal patterns, as well as identify a device command based on the learned vocal patterns. Where the voice input is authenticated as a permissioned user, the device command associated with the voice input may be used to control the renal care system 200 to start, end, adjust, configure or otherwise control a renal care procedure.
[0043] In some embodiments, the controller 110 may include hardware components such as one or more processing device(s) 114 and one or more storage device(s) 118. In some embodiments, the processing device(s) 114 may include any type of data processing capacity, such as a hardware logic circuit, for example an application specific integrated circuit (ASIC) and a programmable logic, or such as a computing device, for example, a microcomputer or microcontroller that include a programmable microprocessor. In some embodiments, the processing device(s) 114 may include data-processing capacity provided by the microprocessor. In some embodiments, the microprocessor may include memory, processing, interface resources, controllers, and counters. In some embodiments, the microprocessor may also include one or more programs stored in memory.
[0044] Similarly, the controller 1 10 may include storage device(s) 118, such as one or more local and / or remote data storage solutions such as, e.g., local hard-drive, solid-state drive, flash drive, database or other local data storage solutions or any combination thereof, and / or remote data storage solutions such as a server, mainframe, database or cloud services, distributed database or other suitable data storage solutions or any combination thereof. In some embodiments, the storage device(s) 118 may include, e.g., a suitable non-transient computer readable medium such as, e.g., random access memory (RAM), read only memory (ROM), one or more buffers and / or caches, among other memory devices or any combination thereof.
[0045] In some embodiments, the controller 110 may implement computer engines for implementing the speech processing pipeline 112 for user authentication and device command generation from voice inputs. In some embodiments, the terms “computer engine” and “engine” identify at least one software component and / or a combination of at least one software component and at least one hardware component which are designed / programmed / configured to manage / control other software and / or hardware components (such as the libraries, software development kits (SDKs), objects, etc.).
[0046] Examples of hardware elements may include processors, microprocessors, circuits, circuit elements (e.g., transistors, resistors, capacitors, inductors, and so forth), integrated circuits, application specific integrated circuits (ASIC), programmable logic devices (PLD), digital signal processors (DSP), field programmable gate array (FPGA), logic gates, registers, semiconductor device, chips, microchips, chip sets, and so forth. In some embodiments, the one or more processors may be implemented as a Complex Instruction Set Computer (CISC) or Reduced Instruction Set Computer (RISC) processors; x86 instruction set compatible processors, multicore, or any other microprocessor or central processing unit (CPU). In various implementations, the one or more processors may be dual-core processor(s), dual -core mobile processor(s), and so forth.
[0047] Examples of software may include software components, programs, applications, computer programs, application programs, system programs, machine programs, operating system software, middleware, firmware, software modules, routines, subroutines, functions, methods, procedures, software interfaces, application program interfaces (API), instruction sets, computing code, computer code, code segments, computer code segments, words, values, symbols, or any combination thereof. Determining whether an embodiment is implemented using hardwareelements and / or software elements may vary in accordance with any number of factors, such as desired computational rate, power levels, heat tolerances, processing cycle budget, input data rates, output data rates, memory resources, data bus speeds and other design or performance constraints.
[0048] In some embodiments, to perform user authentication and device command generation from voice inputs, the controller 110 may include computer engines including, e.g., one or more hardware and / or software components for instantiating and running the speech processing pipeline 112. In some embodiments, the speech processing pipeline 112 may include dedicated and / or shared software components, hardware components, or a combination thereof. For example, the speech processing pipeline 112 may include a dedicated processor and storage. However, in some embodiments, the speech processing pipeline 112 may share hardware resources, including the processing device(s) 114 and storage device(s) 118 of the controller 110 via, e.g., a bus.
[0049] In some embodiments, the controller 110 may receive a voice signal 103 from an audio source such as one or more voice input device(s) 101. The voice input device 101 may include one or more recording devices and / or components, including a microphone, digital-to-analog converter (DAC), audio encoder, audio decoder, audio encoder-decoder (codec), or other devices and / or components or any combination thereof. In some embodiments, the voice input device 101 may include a storage and / or memory storing a recording of voice input, such as a memory associated with the controller, an external memory, a remote storage device / system (e.g., cloud storage), or other non-transitory computer readable medium or any combination thereof.
[0050] A computer-readable medium may include any medium and / or mechanism for storing or transmitting information in a form readable by a machine (e.g., a computing device). For example, a machine-readable medium may include read only memory (ROM); random access memory (RAM); magnetic disk storage media; optical storage media; flash memory devices; electrical, optical, acoustical or other forms of propagated signals (e.g., carrier waves, infrared signals, digital signals, etc.), and others.
[0051] In some embodiments, the voice input device 101 may be a part of the controller 110, the renal care system 200 and / or an external device. The external device may include a user’s device, such as, e.g., at least one personal computer (PC), laptop computer, ultra-laptop computer, tablet, touch pad, portable computer, handheld computer, palmtop computer, a mobile phone, Personal Digital Assistant (PDA), Blackberry ™, Pager, Smartphone, smart watch, television, smart device(e g., smart phone, smart tablet or smart television), mobile internet device (MID), messaging device, data communication device, and so forth.
[0052] In some embodiments, the voice input device 101 may provide the voice signal 103 to the controller 110. In some embodiments, the controller 110 may communicate with the voice input device 101 via a network connection. In some embodiments, the voice input device 101 may interact with the controller 110 using one or more suitable local and / or network communication protocols, such as, e.g., a messaging protocol, a networking protocol, one or more application programming interfaces (APIs), or other suitable technique for communicating between computing systems or any combination thereof. For example, the voice input device 101 may interact with the controller 110 over a network including the Internet using the HyperText Transport Protocol (HTTP) to communicate one or more API requests to cause the controller 110 to perform voice-based authentication and control of the renal care system 200. In another example, the controller 110 is connected to the renal care system 200 via a local network, such as, e.g., Ethernet, Local Area Network (LAN), wireless LAN (WLAN), WiFi, Bluetooth, or other suitable networking technology or any combination thereof, and communicate via API requests and / or database queries in a suitable database query language (e.g., JSONiq, LDAP, Object Query Language (OQL), Object Constraint Language (OCL), PTXL, QUEL, SPARQL, SQL, XQuery, Cypher, DMX, FQL, Contextual Query Language (CQL), AQL, among suitable database query languages). In another example, the controller 110 may be local to the renal care system 200, such as, e.g., a software program installed on the renal care system 200 and configured to employ computing hardware of the renal care system 200 to perform the voice-based authentication and control of the renal care system 200.
[0053] In some embodiments, the voice signal 103 may include one or more representations of at least one utterance voiced by a user. In some embodiments, the voice signal 103 may include a digital audio recording in a suitable audio format, such as, e.g., MP3, MP4, AAC, Ogg, FLAC, ALAC, WMA, among others or any combination thereof. In some embodiments, the voice signal 103 may include a waveform representing the utterance(s), such as, e.g., a waveform formed from quantization of an analog audio recording of the utterance(s).
[0054] In some embodiments, the controller 110 may be configured to receive the voice signal 103 to determine whether the utterance(s) are authentic, where authentic indicates that the utterance was produced by the user permissioned to control the renal care system 200. To do so,the controller 1 10 may employ a speech processing pipeline 112 to recognize a command associated with the utterance(s), and authenticate the utterance(s) as actually from the permissioned user.
[0055] In some embodiments, the speech processing pipeline 112 may use a learned speech pattern library 120 to determine whether the voice signal 103 matches to a permissioned user and / or a valid device command based on a learned speech pattern library 120 and a permissions configuration(s) 122 stored in the storage device(s) 118.
[0056] In some embodiments, the permissions configuration(s) 122 may establish which users are permissioned for voice control of the renal care system 200. Accordingly, the permissions configuration(s) 122 may include an index of identities associated with the permissioned users, including, e.g., one or more login information for each permissioned user, such as a username or user identifier (ID) and a password, passkey, PIN, cryptographic key, or other credential(s) or any combination thereof.
[0057] In some embodiments, the permissions configuration(s) 122 may further include a permission level or other definition of functionalities accessible by each permissioned user. For example, the controller 110 and / or renal care system 200 may have multiple functionalities, where each functionality is associated with at least one permission level. For example, an administrator or root access permission level may include permissions to access all functions, features and / or data of the controller 110 and / or renal care system 200. In another example, a health care provider permission level may include permissions to access certain patient data and / or settings associated with each patient using the renal care system 200. In another example, a user, such as a patient, may have a user permission level may include permissions to access only the functions, features and / or data associated with operation of the renal care system 200 for their own treatment.
[0058] In some embodiments, the learned speech pattern library 120 may include a library of speech patterns generated through a calibration and training stage (see, e.g., FIG. 2) of the speech processing pipeline 112. The learned speech pattern library 120 may include one or more utterances by a permissioned user that represent the vocal patterns of the permissioned user. The learned speech patterns of the learned speech pattern library 120 may include one or more measured and / or derived features associated with the vocal patterns of the permissioned user as will be described in greater detail below. Such features are unique to each individual and form a vocal fingerprint of the individual’s voice. As such, the learned speech patterns of a permissioneduser may be used to authenticate the permissioned user through speech. Thus, the learned speech patterns may be associated with the identity data in the permissions configuration(s) 122 as a mechanism for authenticating the permissioned user’s identity, e.g., in addition to or instead of the credentials associated with the identity.
[0059] In some embodiments, the learned speech pattern library 120 may include learned speech patterns associated with device commands for controlling the renal care system 200. Examples of such device commands may include, e.g., powering the renal care system 200 and / or component(s) thereof on or off, adjusting a flow rate, performing a calibration, initiating cartridge recharging, starting and / or stopping a treatment, defining treatment parameters, clear (e.g., remove, delete, deprecate, mark as resolved, etc.) the errors already stored in the memory after correction, send a treatment report (e.g., by e-mail, in-app notification, API integration, text message, internet message communication, etc.) (such as where the device is connected with ethernet, WiFi, Bluetooth, USB, or other wired and / or wireless connection), enabling / disabling child lock feature, status of the recharge remaining time and therapy remaining time, among other commands or any combination thereof. Similar to the vocal patterns of the permissioned user, each command may include a unique set of features. Thus, each command may be identified based on an associated learned speech pattern that represents the vocal patterns of one or more utterances of the command.
[0060] In some embodiments, the learned speech pattern library 120 may include a set of learned speech patterns for each permissioned user, where the set includes a learned speech pattern for a particular permissioned user for each device command. Thus, the calibration process may include the particular permissioned user providing one or more training or calibration utterance for each device command. Thus, the speech processing pipeline 112 may simultaneously authenticate the permissioned user and identify the device command based on a new voice signal 103 and the learned speech patterns.
[0061] In some embodiments, the learned speech pattern library 120 may include a set of learned speech patterns for each device command generalized to any user, and one or more learned speech patterns associated with each permissioned user’s vocal patterns. Thus, the speech processing pipeline 112 may separately use the set of learned speech patterns for each device command to determine whether a new voice signal 103 is associated with a particular device command, and the one or more learned speech patterns of a particular permissioned user to authenticate the user providing the new voice signal 103.
[0062] In some embodiments, the reliability of voice authentication may be increased by requiring a permissioned user to utter an authentication utterance, such as a password or other unique word or phrase. Thus, the one or more learned speech patterns used to authenticate the permissioned user may be specific to the authentication utterance.
[0063] In some embodiments, the speech processing pipeline 112 may include one or more learning models. In some embodiments, machine learning model(s) may include one or more machine learning architectures, including, but not limited to, a decision tree, random forest, an Adaboost, a K Nearest-neighbor, a Support Vector Machine (SVM), a Gaussian Mixture Model (GMM), a Deep Neural Network (DNN), a convolution neural network (CNN), a recurrent neural network (RNN). For example, a DNN approach may facilitate end-to-end speech recognition with or without feature extraction, an RNN approach may facilitate interpreting sequential data such as as follow-up or a series of commands, and a CNN approach may be utilized where the voice signal 103 includes or is used to form a spectrogram and / or spectrogram -based features.
[0064] In some embodiments and, optionally, in combination of any embodiment described above or below, an exemplary implementation of Neural Network may be executed as follows: a. define Neural Network architecture / model, b. transfer the input data to the exemplary neural network model, c. train the exemplary model incrementally, d. determine the accuracy for a specific number of timesteps, e. apply the exemplary trained model to process the newly-received input data, f. optionally and in parallel, continue to train the exemplary trained model with a predetermined periodicity.
[0065] In some embodiments and, optionally, in combination of any embodiment described above or below, the exemplary trained neural network model may specify a neural network by at least a neural network topology, a series of activation functions, and connection weights. For example, the topology of a neural network may include a configuration of nodes of the neural network and connections between such nodes. In some embodiments and, optionally, in combination of any embodiment described above or below, the exemplary trained neural network model may also be specified to include other parameters, including but not limited to, bias values / functions and / or aggregation functions. For example, an activation function of a node may be a step function, sine function, continuous or piecewise linear function, sigmoid function, hyperbolic tangent function,or other type of mathematical function that represents a threshold at which the node is activated. In some embodiments and, optionally, in combination of any embodiment described above or below, the exemplary aggregation function may be a mathematical function that combines (e.g., sum, product, etc.) input signals to the node. In some embodiments and, optionally, in combination of any embodiment described above or below, an output of the exemplary aggregation function may be used as input to the exemplary activation function. In some embodiments and, optionally, in combination of any embodiment described above or below, the bias may be a constant value or function that may be used by the aggregation function and / or the activation function to make the node more or less likely to be activated.
[0066] In some embodiments, the speech processing pipeline 112 may ingest the voice signal 103 and, where the speaker can be authenticated as a permissioned user, output an authenticated device command 116. To do so, in some embodiments, the speech processing pipeline 112 may process the voice signal 103 to extract and / or generate one or more features to quantitatively represent the voice signal 103. Such features may be encoded in a feature vector and input into one or more machine learning models, such as those detailed above. In some embodiments, the machine learning model(s) may be trained to authenticate the speaker based on the features of the voice signal 103 and to recognize a device command uttered in the voice signal 103.
[0067] In some embodiments, the training may be based on the learned speech pattern library 120. For example, the features may be used to determine a pattern that represents one or more characteristics of the voice signal 103, and the pattern may be classified based on its similarity to the learned speech patterns, the similarity being indicative of a likelihood of the voice signal 103 having originated from a permissioned user and / or a most likely command represented by the voice signal 103. For example, where the similarity measure exceeds a predetermined threshold, the voice signal 103 may be determined to match the learned speech pattern, and as a result the voice signal 103 may be authenticated and / or a device command may be classified.
[0068] In some embodiments, one or more machine learning models may be used to generate the pattern of the voice signal 103. Alternatively, or in addition, the pattern may be generated by one or more statistical models and / or mathematical functions. Alternatively, or in addition, the pattern may be defined by the feature vector encoding the extracted features.
[0069] In some embodiments, the voice signal 103 may be classified using one or more machine learning models. The machine learning model(s) may include one or more supervised learningmodels trained on the learned speech patterns, such as, but not limited to, an RNN, a CNN, a DNN, a SVM, random forest, decision trees, among others or any combination thereof. The machine learning model(s) may include one or more unsupervised learning models trained on the learned speech patterns, such as, but not limited to, clustering, association, dimensionality reduction, Gaussian Mixture Model (GMM), autoencoders, among others or any combination thereof. For example, the unsupervised model(s) may include one or more clustering models trained to cluster the pattern of the voice signal 103 with the learned speech patterns so as to identify the nearest, and thus most likely, classification.
[0070] In some embodiments, whether supervised or unsupervised, the speech processing pipeline 112 may output a classification. The classification may include, e.g., an authentication classification, device command classification, or both. Where the voice signal 103 is confirmed as authentic and the device command is recognized, the speech processing pipeline 112 may output an authenticated device command 116 configured to cause the renal care system 200 to operate according to the device command. Thus, in some embodiments, one or more processing device(s) 114 may instruct the renal case system 200 with the authenticated device command 116 so as to control the renal care system 200 and perform the commanded function.
[0071] Referring to FIG. 2, training an ML model for ML-based voice authentication for improved security in controlling the renal care system 200 is illustrated in accordance with aspects of one or more embodiments of the present disclosure.
[0072] In some embodiments, the speech processing pipeline 112 may include training one or more machine learning models using signal processing 202, pattern generation 204 and classification via one or more machine learning-based (ML) classifier(s) 205.
[0073] In some embodiments, signal processing 202 may include analog-to-digital conversion. In some embodiments, the analog-to-digital conversion process includes quantization. In digital audio, an analog-to-digital converter captures from an analog recording thousands of audio samples per second at a specified sample rate and bit depth to reconstruct the original signal as a digital voice signal.
[0074] In some embodiments, the signal processing 202 may include audio preprocessing. Audio preprocessing may include techniques applied to the digital audio data to enhance quality, extract features, and prepare for further analysis or input into machine learning models. In some embodiments, the audio preprocessing may include resampling the audio data. In someembodiments, the voice input device 101 may include a recording device and / or D AC that captures the digital audio signal at a sampling rate that is different from the sampling rate used by the speech processing pipeline 112. Thus, the resampling may convert the digital audio signal from the original sampling rate to the sampling rate of the speech processing pipeline 112. For example, the ML classifier may use a sampling rate of 16 kHz, while the digital audio signal may include a different sampling rate (e.g., 8 kHz). To align the sampling rates, the audio data may be resampled to match the model’s expected rate, e.g., including calculating additional sample values to fill in between the existing ones.
[0075] In some embodiments, the signal processing 202 may include filtering. Filtering may include removing unwanted noise or artifacts from the audio data. For example, the filtering may include, e.g., low-pass filtering to removes high-frequency noise, high-pass filtering to removes low-frequency noise, band-pass filtering to retains a specific frequency range, notch filtering to eliminate specific frequencies (e.g., 50 Hz for power line interference), among other filtering techniques or any combination thereof.
[0076] In some embodiments, using the preprocessed digital audio data, the speech processing pipeline 112 may generate one or more patterns representative of the utterance(s) of the voice signal 103. In some embodiments, to train the ML classifier 205, the pattern(s) may be entered into the learned speech pattern library 120 as known speech data associated with an authentic permissioned user. Thus, the ML classifier 205 may be trained to classify patterns based on the known speech patterns in the learned speech pattern library 120. In some embodiments, the create the pattern(s), the pattern generation 204 may include feature extraction and / or pattern recognition.
[0077] In some embodiments, feature extraction may include extracting relevant features from the audio signal. For example, the pattern generation 204 may determine one or more spectrograms, Mel-frequency cepstral coefficients (MFCCs), pitch(es), energy / energies, or among other features or any combination thereof. In some embodiments, spectrograms capture visual representations of the audio spectrum over time. In some embodiments, MFCCs capture spectral features of the audio data. In some embodiments, pitch may include a fundamental frequency of the voice. In some embodiments, energy may include intensity of the sound.
[0078] Feature extraction may extract at least features of the utterance(s). Specifically, feature extraction may analyze, for the utterance(s), non-linguistic information, which is different from linguistic information indicating a specific semantic content of the utterance(s). Further, featureextraction generates a feature vector. Feature extraction may extract features other than the utterance(s) and generate a feature vector.
[0079] For example, in some embodiments, the feature extraction may include applying MFCCs. However, the scope of the present invention should not be limited to MFCCs only, other approaches that may achieve feature extraction can also be applied in the present invention (some other features may also suitable for acoustic analysis, e.g., log-power spectrogram, i-vector, X- vector, fundamental frequency, phonetic posteriorgrams, linear predictive coefficients, linear predictive cepstral coefficients, data-driven approach, etc.).
[0080] Mel frequency cepstrum coefficients (MFCCs) may include the linear transformation of the logarithmic energy spectrum based on the nonlinear mel scale of sound frequency. Past researches have pointed out that the sound features extracted using this MFCC feature are closer to the operation method of human cochlea's perception of sound, and have good effects in multiple sound situation recognition applications (e.g., speech recognition, speaker recognition . . . etc.). MFCC feature extraction includes seven steps, including: (1) Pre-emphasis; (2) Frame blocking; (3) Hamming window; (4) Discrete Fourier transform; (5) Triangular band-pass filter; (6) Discrete cosine transformation; and (7) Difference cepstrum coefficient.
[0081] Note that the non-linguistic information is information that is different from the linguistic information (the character string) of utterance(s) to be processed and includes at least one of prosodic information on the utterance(s) and response history information. The prosodic information is information indicating features of a voice waveform of utterance(s) such as a fundamental frequency, a sound pressure, a variation in frequency or the like, a band of variations, a maximum amplitude, an average amplitude, and so on.
[0082] Specifically, feature extraction analyzes prosodic information based on the voice waveform by performing a voice analysis or the like for the user voice data acquired by the voice input device 101. Then, feature extraction calculates a value indicating a feature quantity indicating the prosodic information. Note that feature extraction may calculate, for the user voice data, a fundamental frequency or the like for each of frames that are obtained by dividing the user voice data, for example, at the interval of 32 msec.
[0083] In some embodiments, the pattern generation 204 may include using the extracted features to build a fingerprint representing the voice signal 103. In some embodiments, the pattern generation 204 may include encoding the extracted features in a feature vector, generatingadditional features using, e.g., isolated word recognition, connected word recognition, continuous speech recognition, among other pattern recognition techniques to represent features of one or more attributes of speech or any combination thereof.
[0084] In some embodiments, isolated word recognition may include identifying individual words spoken in isolation. Example, but non-limiting, techniques for isolated word recognition may include aligning an input speech signal with reference templates to find the best match Dynamic Time Warping (DTW), modelling sequential data (e g., phonemes) and capture transitions between states using Hidden Markov Models (HMMs), comparing the input signal with predefined templates for each word such as with template matching.
[0085] In some embodiments, connected word recognition may include recognizing words spoken in natural sequences (e.g., not isolated). Example, but non-limiting, techniques for connected word recognition may include, e g., techniques include: probabilistic models based on word sequences such as N-gram, modeling transitions between connected words such as by using HMMs, learning techniques for incorporating context to improve recognition accuracy such as with language models, deep neural networks or others or any combination thereof.
[0086] In some embodiments, continuous speech recognition may include recognizing entire sentences or phrases spoken naturally. Example, but non-limiting, techniques for continuous speech recognition may include HMM-based systems for combining acoustic models with language models, neural network architectures (e.g., recurrent neural networks, transformer-based models) for handling long-range dependencies, feature extraction for transforming raw audio into relevant features (e.g., Mel-frequency cepstral coefficients).
[0087] In some embodiments, during the training stage, the patterns generated by the pattern generation 204 may be entered into a learned speech pattern library 120. In some embodiments, the learned speech pattern library 120 may provide calibration for the ML classifier(s) 205 to calibrate the ML classifier(s) 205 to the user’s voice and / or to the utterances of commands for the renal care system 200.
[0088] In some embodiments, the ML classifier(s) 205 may be trained using the learned speech pattern library 120 to authenticate the permissioned user by voice, recognition a device command uttered by the permissioned user, or both. Accordingly, the ML classifier(s) 205 may update parameters of a voice recognition model to correlate the learned speech patterns with the associated user and / or device command. For example, the voice recognition model may be configured todetermine a measure of a match between the voice signal 103 and the learned speech patterns. For example, the ML classified s) 205 may include one or more clustering model(s), neural networks, decision trees, language models, random forest(s), among others or any combination thereof.
[0089] In some embodiments, the parameters of the ML classified s) 205 may be trained based on known outputs. For example, the learned speech pattern(s) may be paired with a target classification or known classification to form a training pair, such as a historical learned speech pattern(s) and an observed result and / or human annotated classification denoting whether the historical learned speech pattem(s) is associated with a permissioned user and / or a device command. In some embodiments, the learned speech pattern(s) may be provided to the ML classifier(s) 205, e.g., encoded in a feature vector, to produce a predicted label. In some embodiments, an optimization function associated with the ML classifier(s) 205 may then compare the predicted label with the known output of a training pair including the historical learned speech pattern(s) to determine an error of the predicted label. In some embodiments, the optimization function may employ a loss function, such as, e.g., Hinge Loss, Multi-class SVM Loss, Cross Entropy Loss, Negative Log Likelihood, or other suitable classification loss function to determine the error of the predicted label based on the known output.
[0090] In some embodiments, the known output may be obtained after the ML classifier(s) 205 produces the prediction, such as in online learning scenarios. In such a scenario, the ML classifier(s) 205 may receive the learned speech pattern(s) and generate the model output vector to produce a label classifying the learned speech pattern(s). Subsequently, a user may provide feedback by, e.g., modifying, adjusting, removing, and / or verifying the label via a suitable feedback mechanism, such as a user interface device (e.g., keyboard, mouse, touch screen, user interface, or other interface mechanism of a user device or any suitable combination thereof). The feedback may be paired with the learned speech pattern(s) to form the training pair and the optimization function may determine an error of the predicted label using the feedback.
[0091] In some embodiments, based on the error, the optimization function may update the parameters of the ML classified s) 205 using a suitable training algorithm such as, e.g., backpropagation for a classifier machine learning model. In some embodiments, backpropagation may include any suitable minimization algorithm such as a gradient method of the loss function with respect to the weights of the classifier machine learning model. Examples of suitable gradient methods include, e.g., stochastic gradient descent, batch gradient descent, mini-batch gradientdescent, or other suitable gradient descent technique. As a result, the optimization function may update the parameters of the ML classifier(s) 205 based on the error of predicted labels in order to train the ML classifier(s) 205 to model the correlation between learned speech pattern(s) and permissioned user authentication and / or device command in order to produce more accurate labels of learned speech pattem(s).
[0092] In some embodiments, the ML classifier(s) 205 may utilize a clustering model. The clustering model may map the input feature vector encoding a speech pattern into a multidimensional space. Each learned speech pattern in the learned speech pattern library 120 may also be mapped to the multi-dimensional space. The optimization function may be applied to the clustering algorithm so as to fine-tune parameters to create clusters of learned speech patterns, each cluster being associated with a particular permissioned user, a particular device command, or a particular combination thereof. In some embodiments, the clustering algorithm may include, e.g., k-mean, k-nearest neighbor, or other suitable clustering algorithm that clusters based on similarity as measured by distance in the multi-dimensional space. For example, similarity may be measured using, e.g., Jaccard similarity, Jaro-Winkler similarity, Cosine similarity, Euclidean similarity, Overlap similarity, Pearson similarity, Approximate Nearest Neighbors, K-Nearest Neighbors, among other similarity measure or any combination thereof.
[0093] In some embodiments, the ML classifier(s) 205 may include a single voice recognition model that is trained to classify a speech pattern according to permissioned user and device command represented by the voice signal 103. In some embodiments, the ML classifier(s) 205 may include two separate voice recognition models: a first trained to authenticate the permissioned user by classifying a speech pattern as being associated with a voice signal 103 uttered by the permissioned user, and a second trained to classify the speech pattern or a second speech pattern as a particular device command uttered by the permissioned user in the voice signal 103 or a subsequent voice signal. In some embodiments, the second voice recognition model may be configured to only be triggered upon the first voice recognition model successfully classifying the speech pattern as associated with the permissioned user. In some embodiments, the separate voice recognition models may be trained on different learned speech patterns, e.g., learned speech patterns for authenticating the permissioned user, and learned speech patterns for available device commands.
[0094] In some embodiments, the user may be prompted to input separate voice signal 103, a first for authentication and a second for a device command. In some embodiments, the voice signal 103 for authentication may include a preset, selected or configured keyword, such as a passphrase or password that the permissioned user is prompted to utter so as to facilitate pattern matching by removing variability in different utterances.
[0095] Referring to FIG. 3, executing an ML model, trained in accordance with FIG. 2, for ML- based voice authentication for improved security in controlling the renal care system 200 is illustrated in accordance with aspects of one or more embodiments of the present disclosure.
[0096] In some embodiments, in the inference stage with the speech processing pipeline 112, the ML classifier(s) 205 may be deployed for new voice input signals 103. In some embodiments, the new voice signal 103 may undergo the signal processing 202 and the pattern generation 204 to generate a speech pattern for the new voice signal 103. Based on the generated speech pattern and the learned parameters as detailed above, the ML classifier(s) 205 may classifying the voice signal 103 as an authentication determination (e.g., authenticated or not authenticated) of the user that provided the voice input signal as the permissioned user, a command uttered to produce the voice signal 103, or both. As a result, the speech processing pipeline 112 may ingest a voice signal 103 to concurrently authenticated a user as a permissioned user and generate a command to the renal care system 200 to enable a permissioned user to securely control the renal care system 200 with voice.
[0097] Referring to FIG. 4, a flow chart for ML-based voice authentication for improved security in controlling the renal care system 200 is illustrated in accordance with aspects of one or more embodiments of the present disclosure.
[0098] In some embodiments, a user may utter a voiced command into a voice input device 101. The voice input device 101 produces the voice signal 103 representing the voiced command. In some embodiments, a signal processing and pattern generation 401 stage may generate a voice pattern 402 representing the voice signal 103. In some embodiments, the signal processing and pattern generation 401 may extract and / or generate extracted features to build a fingerprint representing the voice signal 103. In some embodiments, the signal processing and pattern generation 401 may include encoding the extracted features in a feature vector, generating additional features using, e.g., isolated word recognition, connected word recognition, continuousspeech recognition, among other pattern recognition techniques to represent features of one or more attributes of speech or any combination thereof.
[0099] In some embodiments, the voice pattern(s) 402 may be compared to learned speech pattern(s) in the learned speech pattern library 120 at stage 403. In some embodiments, the comparison may include, e.g., one or more ML classifier model(s), isolated word recognition, connected word recognition, continuous speech recognition, clustering, among other comparison techniques or any combination thereof.
[0100] In some embodiments, isolated word recognition may include identifying individual words spoken in isolation. Example, but non-limiting, techniques for isolated word recognition may include aligning an input speech signal with reference templates to find the best match Dynamic Time Warping (DTW), modelling sequential data (e.g., phonemes) and capture transitions between states using Hidden Markov Models (HMMs), comparing the input signal with predefined templates for each word such as with template matching.
[0101] In some embodiments, connected word recognition may include recognizing words spoken in natural sequences (e.g., not isolated). Example, but non-limiting, techniques for connected word recognition may include, e.g., techniques include: probabilistic models based on word sequences such as N-gram, modeling transitions between connected words such as by using HMMs, learning techniques for incorporating context to improve recognition accuracy such as with language models, deep neural networks or others or any combination thereof.
[0102] In some embodiments, continuous speech recognition may include recognizing entire sentences or phrases spoken naturally. Example, but non-limiting, techniques for continuous speech recognition may include HMM-based systems for combining acoustic models with language models, neural network architectures (e.g., recurrent neural networks, transformer-based models) for handling long-range dependencies, feature extraction for transforming raw audio into relevant features (e.g., Mel-frequency cepstral coefficients).
[0103] In some embodiments, the ML classified s) 205 may be configured to determine a measure of a match between the voice signal 103 and the learned speech patterns. For example, the ML classifier(s) 205 may include one or more clustering model(s), neural networks, decision trees, language models, random forest(s), among others or any combination thereof.
[0104] In some embodiments, the ML classifier(s) 205 may utilize a clustering model. The clustering model may map the input feature vector encoding a speech pattern into a multi-dimensional space. Each learned speech pattern in the learned speech pattern library 120 may also be mapped to the multi-dimensional space. The optimization function may be applied to the clustering algorithm so as to fine-tune parameters to create clusters of learned speech patterns, each cluster being associated with a particular permissioned user, a particular device command, or a particular combination thereof. In some embodiments, the clustering algorithm may include, e.g., k-mean, k-nearest neighbor, or other suitable clustering algorithm that clusters based on similarity as measured by distance in the multi-dimensional space. For example, similarity may be measured using, e.g., Jaccard similarity, Jaro-Winkler similarity, Cosine similarity, Euclidean similarity, Overlap similarity, Pearson similarity, Approximate Nearest Neighbors, K-Nearest Neighbors, among other similarity measure or any combination thereof.
[0105] In some embodiments, at test condition 404, where the ML classifier(s) does not identify a match between the voice pattern(s) 402 and a learned speech pattern, the process ends 407. In some embodiments, the end 407 may terminate. In some embodiments, the end 407 may cause an alert or notification to be presented, e.g., via display, audio, tactile output, or other suitable technique using an associated device (e.g., display, speaker(s), vibration motor, respectively) among others or any combination thereof. The alert or notification may inform the user of the failed match and prompt the user to try again.
[0106] In some embodiments, at test condition 404, where the ML classifier(s) does identify a match is found, the user is authenticated as a permissioned user and the command uttered by the user is recognized. As a result, a signal representing the command is generated and the command is issued at 405. As a result, a device command 406 is sent to the renal care system 200 to cause the renal care system 200 to perform at least one operation based on the device command 406. As a result, the user may securely provide voice based commands to the renal care system 200 to control the operation of the renal care system 200.
[0107] FIG. 5 depicts a block diagram of another exemplary computer-based system and platform 500 in accordance with one or more embodiments of the present disclosure. However, not all of these components may be required to practice one or more embodiments, and variations in the arrangement and type of the components may be made without departing from the spirit or scope of various embodiments of the present disclosure. In some embodiments, the client device 502a, client device 502b through client device 502n shown each at least includes a computer-readable medium, such as a random-access memory (RAM) 508 coupled to a processor 510 or FLASHmemory. In some embodiments, the processor 510 may execute computer-executable program instructions stored in memory 508. In some embodiments, the processor 510 may include a microprocessor, an ASIC, and / or a state machine. In some embodiments, the processor 510 may include, or may be in communication with, media, for example computer-readable media, which stores instructions that, when executed by the processor 510, may cause the processor 510 to perform one or more steps described herein. In some embodiments, examples of computer- readable media may include, but are not limited to, an electronic, optical, magnetic, or other storage or transmission device capable of providing a processor, such as the processor 510 of client device 502a, with computer-readable instructions. In some embodiments, other examples of suitable media may include, but are not limited to, a floppy disk, CD-ROM, DVD, magnetic disk, memory chip, ROM, RAM, an ASIC, a configured processor, all optical media, all magnetic tape or other magnetic media, or any other medium from which a computer processor can read instructions. Also, various other forms of computer-readable media may transmit or carry instructions to a computer, including a router, private or public network, or other transmission device or channel, both wired and wireless. In some embodiments, the instructions may comprise code from any computer-programming language, including, for example, C, C++, Visual Basic, Java, Python, Perl, JavaScript, and etc.
[0108] In some embodiments, client devices 502a through 502n may also comprise a number of external or internal devices such as a mouse, a CD-ROM, DVD, a physical or virtual keyboard, a display, or other input or output devices. In some embodiments, examples of client devices 502a through 502n (e.g., clients) may be any type of processor-based platforms that are connected to a network 506 such as, without limitation, personal computers, digital assistants, personal digital assistants, smart phones, pagers, digital tablets, laptop computers, Internet appliances, and other processor-based devices. In some embodiments, client devices 502a through 502n may be specifically programmed with one or more application programs in accordance with one or more principles / methodologies detailed herein. In some embodiments, client devices 502a through 502n may operate on any operating system capable of supporting a browser or browser-enabled application, such as Microsoft™, Windows™, and / or Linux. In some embodiments, client devices 502a through 502n shown may include, for example, personal computers executing a browser application program such as Microsoft Corporation's Internet Explorer™, Apple Computer, Inc.'s Safari™, Mozilla Firefox, and / or Opera. In some embodiments, through the member computingclient devices 502a through 502n, user 512a, user 512b through user 512n, may communicate over the exemplary network 506 with each other and / or with other systems and / or devices coupled to the network 506. As shown in FIG. 5, exemplary server devices 504 and 513 may include processor 505 and processor 514, respectively, as well as memory 517 and memory 516, respectively. In some embodiments, the server devices 504 and 513 may be also coupled to the network 506. In some embodiments, one or more client devices 502a through 502n may be mobile clients.
[0109] In some embodiments, at least one database of exemplary databases 507 and 515 may be any type of database, including a database managed by a database management system (DBMS). In some embodiments, an exemplary DBMS-managed database may be specifically programmed as an engine that controls organization, storage, management, and / or retrieval of data in the respective database. In some embodiments, the exemplary DBMS-managed database may be specifically programmed to provide the ability to query, backup and replicate, enforce rules, provide security, compute, perform change and access logging, and / or automate optimization. In some embodiments, the exemplary DBMS-managed database may be chosen from Oracle database, IBM DB2, Adaptive Server Enterprise, FileMaker, Microsoft Access, Microsoft SQL Server, MySQL, PostgreSQL, and a NoSQL implementation. In some embodiments, the exemplary DBMS-managed database may be specifically programmed to define each respective schema of each database in the exemplary DBMS, according to a particular database model of the present disclosure which may include a hierarchical model, network model, relational model, object model, or some other suitable organization that may result in one or more applicable data structures that may include fields, records, files, and / or objects. In some embodiments, the exemplary DBMS-managed database may be specifically programmed to include metadata about the data that is stored.
[0110] In some embodiments, the exemplary inventive computer-based systems / platforms, the exemplary inventive computer-based devices, and / or the exemplary inventive computer-based components of the present disclosure may be specifically configured to operate in a cloud computing / architecture 525 such as, but not limiting to: infrastructure a service (laaS) 610, platform as a service (PaaS) 608, and / or software as a service (SaaS) 606 using a web browser, mobile app, thin client, terminal emulator or other endpoint 604. FIG. 6 illustrates schematics of exemplary implementations of the cloud computing / architecture(s) in which the exemplary inventive computer-based systems / platforms, the exemplary inventive computer-based devices,and / or the exemplary inventive computer-based components of the present disclosure may be specifically configured to operate.
[0111] It is understood that at least one aspect / functionality of various embodiments described herein can be performed in real-time and / or dynamically. As used herein, the term “real-time” is directed to an event / action that can occur instantaneously or almost instantaneously in time when another event / action has occurred. For example, the “real-time processing,” “real-time computation,” and “real-time execution” all pertain to the performance of a computation during the actual time that the related physical process (e.g., a user interacting with an application on a mobile device) occurs, in order that results of the computation can be used in guiding the physical process.
[0112] As used herein, the term “dynamically” and term “automatically,” and their logical and / or linguistic relatives and / or derivatives, mean that certain events and / or actions can be triggered and / or occur without any human intervention. In some embodiments, events and / or actions in accordance with the present disclosure can be in real-time and / or based on a predetermined periodicity of at least one of nanosecond, several nanoseconds, millisecond, several milliseconds, second, several seconds, minute, several minutes, hourly, several hours, daily, several days, weekly, monthly, etc.
[0113] In some embodiments, exemplary inventive, specially programmed computing systems and platforms with associated devices are configured to operate in the distributed network environment, communicating with one another over one or more suitable data communication networks (e.g., the Internet, satellite, etc.) and utilizing one or more suitable data communication protocol s / m odes such as, without limitation, IPX / SPX, X.25, AX.25, AppleTalk(TM), TCP / IP (e.g., HTTP), near-field wireless communication (NFC), RFID, Narrow Band Internet of Things (NBIOT), 3G, 4G, 5G, GSM, GPRS, WiFi, WiMax, CDMA, satellite, ZigBee, and other suitable communication modes.
[0114] The material disclosed herein may be implemented in software or firmware or a combination of them or as instructions stored on a machine-readable medium, which may be read and executed by one or more processors. A machine-readable medium may include any medium and / or mechanism for storing or transmitting information in a form readable by a machine (e.g., a computing device). For example, a machine-readable medium may include read only memory (ROM); random access memory (RAM); magnetic disk storage media; optical storage media; flashmemory devices; electrical, optical, acoustical or other forms of propagated signals (e.g., carrier waves, infrared signals, digital signals, etc.), and others.
[0115] As used herein, the terms “computer engine” and “engine” identify at least one software component and / or a combination of at least one software component and at least one hardware component which are designed / programmed / configured to manage / control other software and / or hardware components (such as the libraries, software development kits (SDKs), objects, etc.).
[0116] Examples of hardware elements may include processors, microprocessors, circuits, circuit elements (e.g., transistors, resistors, capacitors, inductors, and so forth), integrated circuits, application specific integrated circuits (ASIC), programmable logic devices (PLD), digital signal processors (DSP), field programmable gate array (FPGA), logic gates, registers, semiconductor device, chips, microchips, chip sets, and so forth. In some embodiments, the one or more processors may be implemented as a Complex Instruction Set Computer (CISC) or Reduced Instruction Set Computer (RISC) processors; x86 instruction set compatible processors, multicore, or any other microprocessor or central processing unit (CPU). In various implementations, the one or more processors may be dual-core processor(s), dual -core mobile processor(s), and so forth.
[0117] Computer-related systems, computer systems, and systems, as used herein, include any combination of hardware and software. Examples of software may include software components, programs, applications, operating system software, middleware, firmware, software modules, routines, subroutines, functions, methods, procedures, software interfaces, application program interfaces (API), instruction sets, computer code, computer code segments, words, values, symbols, or any combination thereof. Determining whether an embodiment is implemented using hardware elements and / or software elements may vary in accordance with any number of factors, such as desired computational rate, power levels, heat tolerances, processing cycle budget, input data rates, output data rates, memory resources, data bus speeds and other design or performance constraints.
[0118] One or more aspects of at least one embodiment may be implemented by representative instructions stored on a machine-readable medium which represents various logic within the processor, which when read by a machine causes the machine to fabricate logic to perform the techniques described herein. Such representations, known as “IP cores” may be stored on a tangible, machine readable medium and supplied to various customers or manufacturing facilitiesto load into the fabrication machines that make the logic or processor. Of note, various embodiments described herein may, of course, be implemented using any appropriate hardware and / or computing software languages (e.g., C++, Objective-C, Swift, Java, JavaScript, Python, Perl, QT, etc.).
[0119] In some embodiments, one or more of illustrative computer-based systems or platforms of the present disclosure may include or be incorporated, partially or entirely into at least one personal computer (PC), laptop computer, ultra-laptop computer, tablet, touch pad, portable computer, handheld computer, palmtop computer, personal digital assistant (PDA), cellular telephone, combination cellular telephone / PDA, television, smart device (e.g., smart phone, smart tablet or smart television), mobile internet device (MID), messaging device, data communication device, and so forth.
[0120] As used herein, term “server” should be understood to refer to a service point which provides processing, database, and communication facilities. By way of example, and not limitation, the term “server” can refer to a single, physical processor with associated communications and data storage and database facilities, or it can refer to a networked or clustered complex of processors and associated network and storage devices, as well as operating software and one or more database systems and application software that support the services provided by the server. Cloud servers are examples.
[0121] In some embodiments, as detailed herein, one or more of the computer-based systems of the present disclosure may obtain, manipulate, transfer, store, transform, generate, and / or output any digital object and / or data unit (e.g., from inside and / or outside of a particular application) that can be in any suitable form such as, without limitation, a fde, a contact, a task, an email, a message, a map, an entire application (e.g., a calculator), data points, and other suitable data. In some embodiments, as detailed herein, one or more of the computer-based systems of the present disclosure may be implemented across one or more of various computer platforms such as, but not limited to: (1) FreeBSD, NetBSD, OpenBSD; (2) Linux; (3) Microsoft Windows™; (4) Open VMS™; (5) OS X (MacOS™); (6) UNIX™; (7) Android; (8) iOS™; (9) Embedded Linux; (10) Tizen™; (11) WebOS™; (12) Adobe AIR™; (13) Binary Runtime Environment for Wireless (BREW™); (14) Cocoa™ (API); (15) Cocoa™ Touch; (16) Java™ Platforms; (17) JavaFX™; (18) QNX™; (19) Mono; (20) Google Blink; (21) Apple WebKit; (22) Mozilla Gecko™; (23) Mozilla XUL; (24) NET Framework; (25) Silverlight™; (26) Open Web Platform; (27) OracleDatabase; (28) Qt™; (29) SAP NetWeaver™; (30) Smartface™; (31) Vexi™; (32) Kubernetes™ and (33) Windows Runtime (WinRT™) or other suitable computer platforms or any combination thereof. In some embodiments, illustrative computer-based systems or platforms of the present disclosure may be configured to utilize hardwired circuitry that may be used in place of or in combination with software instructions to implement features consistent with principles of the disclosure. Thus, implementations consistent with principles of the disclosure are not limited to any specific combination of hardware circuitry and software. For example, various embodiments may be embodied in many different ways as a software component such as, without limitation, a stand-alone software package, a combination of software packages, or it may be a software package incorporated as a “tool” in a larger software product.
[0122] For example, exemplary software specifically programmed in accordance with one or more principles of the present disclosure may be downloadable from a network, for example, a website, as a stand-alone product or as an add-in package for installation in an existing software application. For example, exemplary software specifically programmed in accordance with one or more principles of the present disclosure may also be available as a client-server software application, or as a web-enabled software application. For example, exemplary software specifically programmed in accordance with one or more principles of the present disclosure may also be embodied as a software package installed on a hardware device.
[0123] In some embodiments, illustrative computer-based systems or platforms of the present disclosure may be configured to handle numerous concurrent users that may be, but is not limited to, at least 100 (e.g., but not limited to, 100-999), at least 1,000 (e.g., but not limited to, 1,000- 9,999 ), at least 10,000 (e.g., but not limited to, 10,000-99,999 ), at least 100,000 (e.g., but not limited to, 100,000-999,999), at least 1,000,000 (e.g., but not limited to, 1,000,000-9,999,999), at least 10,000,000 (e.g., but not limited to, 10,000,000-99,999,999), at least 100,000,000 (e.g., but not limited to, 100,000,000-999,999,999), at least 1,000,000,000 (e.g., but not limited to, 1,000,000,000-999,999,999,999), and so on.
[0124] In some embodiments, illustrative computer-based systems or platforms of the present disclosure may be configured to output to distinct, specifically programmed graphical user interface implementations of the present disclosure (e.g., a desktop, a web app., etc.). In various implementations of the present disclosure, a final output may be displayed on a displaying screen which may be, without limitation, a screen of a computer, a screen of a mobile device, or the like.In various implementations, the display may be a holographic display. In various implementations, the display may be a transparent surface that may receive a visual projection. Such projections may convey various forms of information, images, or objects. For example, such projections may be a visual overlay for a mobile augmented reality (MAR) application.
[0125] As used herein, terms “cloud,” “Internet cloud,” “cloud computing,” “cloud architecture,” and similar terms correspond to at least one of the following: (1) a large number of computers connected through a real-time communication network (e.g., Internet); (2) providing the ability to run a program or application on many connected computers (e.g., physical machines, virtual machines (VMs)) at the same time; (3) network-based services, which appear to be provided by real server hardware, and are in fact served up by virtual hardware (e.g., virtual servers), simulated by software running on one or more real machines (e.g., allowing to be moved around and scaled up (or down) on the fly without affecting the end user).
[0126] In some embodiments, the illustrative computer-based systems or platforms of the present disclosure may be configured to securely store and / or transmit data by utilizing one or more of encryption techniques (e.g., private / public key pair, Triple Data Encryption Standard (3DES), block cipher algorithms (e.g., IDEA, RC2, RC5, CAST and Skipjack), cryptographic hash algorithms (e g., MD5, RIPEMD-160, RTRO, SHA-1, SHA-2, Tiger (TTH), WHIRLPOOL, RNGs).
[0127] As used herein, the term “user” shall have a meaning of at least one user. In some embodiments, the terms “user”, “subscriber” “consumer” or “customer” should be understood to refer to a user of an application or applications as described herein and / or a consumer of data supplied by a data provider. By way of example, and not limitation, the terms “user” or “subscriber” can refer to a person who receives data provided by the data or service provider over the Internet in a browser session, or can refer to an automated software application which receives the data and stores or processes the data.
[0128] The aforementioned examples are, of course, illustrative and not restrictive.
[0129] At least some aspects of the present disclosure will now be described with reference to the following numbered clauses.
[0130] Clause 1. A system comprising: at least one renal care device configured to provide at least one renal therapy to a user; and a controller in communication with the at least one renal care device, the controller configured to: receive a voice signal from an audio source; input the voicesignal into at least one trained voice recognition machine learning model to output a voice authentication determination and a device command based at least in part on trained parameters of the at least one trained voice recognition model; wherein the trained parameters of the at least one trained voice recognition model are configured to determine a measure of a match between the voice signal and at least one training vocal pattern of a training vocal pattern dataset; and instruct the at least one renal care device with the device command based at least in part on the measure of the match of the voice signal exceeding a predetermined threshold.
[0131] Clause 2. The system of clause 1, wherein the at least one trained voice recognition machine learning model comprises a single voice recognition machine learning model configured to output the voice authentication determination and the device command.
[0132] Clause 3. The system of clause 1, wherein the at least one trained voice recognition machine learning model comprises: at least one trained voice command recognition machine learning model configured to output the device command; and at least one trained voice authentication machine learning model configured to output the voice authentication determination.
[0133] Clause 4. The system of clause 1, wherein the at least one trained voice recognition machine learning model comprises at least one clustering model.
[0134] Clause 5. The system of clause 1, wherein the at least one trained voice recognition machine learning model comprises at least one neural network model.
[0135] Clause 6. The system of clause 1, wherein the controller is further configured to: receive a first voice input associated with a user vocalizing an authentication utterance; input the first voice input into at least one trained voice authentication recognition machine learning model of the at least one trained voice recognition machine learning model to output the voice authentication determination based at least in part on trained authentication parameters; wherein the trained authentication parameters are configured to determine the voice authentication determination based at least in part on an authentication similarity measure between the first voice input and at least one training authentication vocal pattern of a training authentication vocal pattern dataset; generate, based on the authentication similarity measure exceeding a predetermined authentication similarity measure threshold, a prompt to the user for the device command; receive a second voice input associated with the user vocalizing a device command utterance; input the second voice input into at least one trained voice command recognition machine learning model of the at least onetrained voice recognition machine learning model to output the device command based at least in part on trained device command parameters; and wherein the trained device command parameters are configured to determine the device command based at least in part on a device command similarity measure between the second voice input and at least one training device command vocal pattern of a training device command vocal pattern dataset.
[0136] Clause 7. The system of clause 1, wherein the controller is further configured to: generate a voice signal pattern representative of at least one characteristic of the voice signal based at least in part on at least one pre-processing algorithm; and input the voice signal pattern into at least one trained voice recognition machine learning model to output the voice authentication determination and the device command based at least in part on the trained parameters of the at least one trained voice recognition model.
[0137] Clause 8. A method comprising: at least one renal care device configured to provide at least one renal therapy to a user; and a controller in communication with the at least one renal care device, the controller configured to: receiving, by at least one processor, a voice signal from an audio source; inputting, by the at least one processor, the voice signal into at least one trained voice recognition machine learning model to output a voice authentication determination and a device command based at least in part on trained parameters of the at least one trained voice recognition model; wherein the trained parameters of the at least one trained voice recognition model are configured to determine a measure of a match between the voice signal and at least one training vocal pattern of a training vocal pattern dataset; and instructing, by the at least one processor, the at least one renal care device with the device command based at least in part on the measure of the match of the voice signal exceeding a predetermined threshold.
[0138] Clause 9. The method of clause 8, wherein the at least one trained voice recognition machine learning model comprises a single voice recognition machine learning model configured to output the voice authentication determination and the device command.
[0139] Clause 10. The method of clause 8, wherein the at least one trained voice recognition machine learning model comprises: at least one trained voice command recognition machine learning model configured to output the device command; and at least one trained voice authentication machine learning model configured to output the voice authentication determination.
[0140] Clause 1 1. The method of clause 8, wherein the at least one trained voice recognition machine learning model comprises at least one clustering model.
[0141] Clause 12. The method of clause 8, wherein the at least one trained voice recognition machine learning model comprises at least one neural network model.
[0142] Clause 13. The method of clause 8, further comprising: receiving, by the at least one processor, a first voice input associated with a user vocalizing an authentication utterance; inputting, by the at least one processor, the first voice input into at least one trained voice authentication recognition machine learning model of the at least one trained voice recognition machine learning model to output the voice authentication determination based at least in part on trained authentication parameters; wherein the trained authentication parameters are configured to determine the voice authentication determination based at least in part on a authentication similarity measure between the first voice input and at least one training authentication vocal pattern of a training authentication vocal pattern dataset; generate, based on the authentication similarity measure exceeding a predetermined authentication similarity measure threshold, a prompt to the user for the device command; receiving, by the at least one processor, a second voice input associated with the user vocalizing a device command utterance; inputting, by the at least one processor, the second voice input into at least one trained voice command recognition machine learning model of the at least one trained voice recognition machine learning model to output the device command based at least in part on trained device command parameters; and wherein the trained device command parameters are configured to determine the device command based at least in part on a device command similarity measure between the second voice input and at least one training device command vocal pattern of a training device command vocal pattern dataset.
[0143] Clause 14. The method of clause 8, further comprising: generating, by the at least one processor, a voice signal pattern representative of at least one characteristic of the voice signal based at least in part on at least one pre-processing algorithm; and inputting, by the at least one processor, the voice signal pattern into at least one trained voice recognition machine learning model to output the voice authentication determination and the device command based at least in part on the trained parameters of the at least one trained voice recognition model.
[0144] Clause 15. A non-transitory computer readable medium having executable instructions stored thereon, wherein the executable instructions are configured to cause at least one processor to perform steps comprising: receiving a voice signal from an audio source; inputting the voicesignal into at least one trained voice recognition machine learning model to output a voice authentication determination and a device command based at least in part on trained parameters of the at least one trained voice recognition model; wherein the trained parameters of the at least one trained voice recognition model are configured to determine a measure of a match between the voice signal and at least one training vocal pattern of a training vocal pattern dataset; and instructing at least one renal care device with the device command based at least in part on the measure of the match of the voice signal exceeding a predetermined threshold.
[0145] Clause 16. The non-transitory computer readable medium of clause 15, wherein the at least one trained voice recognition machine learning model comprises a single voice recognition machine learning model configured to output the voice authentication determination and the device command.
[0146] Clause 17. The non-transitory computer readable medium of clause 1 , wherein the at least one trained voice recognition machine learning model comprises: at least one trained voice command recognition machine learning model configured to output the device command; and at least one trained voice authentication machine learning model configured to output the voice authentication determination.
[0147] Clause 18. The non-transitory computer readable medium of clause 15, wherein the at least one trained voice recognition machine learning model comprises at least one of: at least one clustering model, or at least one neural network model.
[0148] Clause 19. The non-transitory computer readable medium of clause 15, wherein the executable instructions are further configured to cause the at least one processor to perform steps comprising: receiving a first voice input associated with a user vocalizing an authentication utterance; inputting the first voice input into at least one trained voice authentication recognition machine learning model of the at least one trained voice recognition machine learning model to output the voice authentication determination based at least in part on trained authentication parameters; wherein the trained authentication parameters are configured to determine the voice authentication determination based at least in part on a authentication similarity measure between the first voice input and at least one training authentication vocal pattern of a training authentication vocal pattern dataset; generate, based on the authentication similarity measure exceeding a predetermined authentication similarity measure threshold, a prompt to the user for the device command; receiving a second voice input associated with the user vocalizing a devicecommand utterance; inputting the second voice input into at least one trained voice command recognition machine learning model of the at least one trained voice recognition machine learning model to output the device command based at least in part on trained device command parameters; and wherein the trained device command parameters are configured to determine the device command based at least in part on a device command similarity measure between the second voice input and at least one training device command vocal pattern of a training device command vocal pattern dataset.
[0149] Clause 20. The non-transitory computer readable medium of clause 15, wherein the executable instructions are further configured to cause the at least one processor to perform steps comprising: generating a voice signal pattern representative of at least one characteristic of the voice signal based at least in part on at least one pre-processing algorithm; and inputting the voice signal pattern into at least one trained voice recognition machine learning model to output the voice authentication determination and the device command based at least in part on the trained parameters of the at least one trained voice recognition model. Publications cited throughout this document are hereby incorporated by reference in their entirety. While one or more embodiments of the present disclosure have been described, it is understood that these embodiments are illustrative only, and not restrictive, and that many modifications may become apparent to those of ordinary skill in the art, including that various embodiments of the inventive methodologies, the illustrative systems and platforms, and the illustrative devices described herein can be utilized in any combination with each other. Further still, the various steps may be carried out in any desired order (and any desired steps may be added and / or any desired steps may be eliminated).
Claims
CLAIMSWhat is claimed is:
1. A system comprising: at least one renal care device configured to provide at least one renal therapy to a user; and a controller in communication with the at least one renal care device, the controller configured to: receive a voice signal from an audio source; input the voice signal into at least one trained voice recognition machine learning model to output a voice authentication determination and a device command based at least in part on trained parameters of the at least one trained voice recognition model; wherein the trained parameters of the at least one trained voice recognition model are configured to determine a measure of a match between the voice signal and at least one training vocal pattern of a training vocal pattern dataset; and instruct the at least one renal care device with the device command based at least in part on the measure of the match of the voice signal exceeding a predetermined threshold.
2. The system of claim 1, wherein the at least one trained voice recognition machine learning model comprises a single voice recognition machine learning model configured to output the voice authentication determination and the device command.
3. The system of claim 1, wherein the at least one trained voice recognition machine learning model comprises: at least one trained voice command recognition machine learning model configured to output the device command; andat least one trained voice authentication machine learning model configured to output the voice authentication determination.
4. The system of claim 1, wherein the at least one trained voice recognition machine learning model comprises at least one clustering model.
5. The system of claim 1, wherein the at least one trained voice recognition machine learning model comprises at least one neural network model.
6. The system of claim 1, wherein the controller is further configured to: receive a first voice input associated with a user vocalizing an authentication utterance; input the first voice input into at least one trained voice authentication recognition machine learning model of the at least one trained voice recognition machine learning model to output the voice authentication determination based at least in part on trained authentication parameters; wherein the trained authentication parameters are configured to determine the voice authentication determination based at least in part on an authentication similarity measure between the first voice input and at least one training authentication vocal pattern of a training authentication vocal pattern dataset; generate, based on the authentication similarity measure exceeding a predetermined authentication similarity measure threshold, a prompt to the user for the device command; receive a second voice input associated with the user vocalizing a device command utterance; input the second voice input into at least one trained voice command recognition machine learning model of the at least one trained voice recognition machine learning model to output the device command based at least in part on trained device command parameters; andwherein the trained device command parameters are configured to determine the device command based at least in part on a device command similarity measure between the second voice input and at least one training device command vocal pattern of a training device command vocal pattern dataset.
7. The system of claim 1, wherein the controller is further configured to: generate a voice signal pattern representative of at least one characteristic of the voice signal based at least in part on at least one pre-processing algorithm; and input the voice signal pattern into at least one trained voice recognition machine learning model to output the voice authentication determination and the device command based at least in part on the trained parameters of the at least one trained voice recognition model.
8. A method comprising: at least one renal care device configured to provide at least one renal therapy to a user; and a controller in communication with the at least one renal care device, the controller configured to: receiving, by at least one processor, a voice signal from an audio source; inputting, by the at least one processor, the voice signal into at least one trained voice recognition machine learning model to output a voice authentication determination and a device command based at least in part on trained parameters of the at least one trained voice recognition model; wherein the trained parameters of the at least one trained voice recognition model are configured to determine a measure of a match between the voice signal and at least one training vocal pattern of a training vocal pattern dataset; andinstructing, by the at least one processor, the at least one renal care device with the device command based at least in part on the measure of the match of the voice signal exceeding a predetermined threshold.
9. The method of claim 8, wherein the at least one trained voice recognition machine learning model comprises a single voice recognition machine learning model configured to output the voice authentication determination and the device command.
10. The method of claim 8, wherein the at least one trained voice recognition machine learning model comprises: at least one trained voice command recognition machine learning model configured to output the device command; and at least one trained voice authentication machine learning model configured to output the voice authentication determination.
11. The method of claim 8, wherein the at least one trained voice recognition machine learning model comprises at least one clustering model.
12. The method of claim 8, wherein the at least one trained voice recognition machine learning model comprises at least one neural network model.
13. The method of claim 8, further comprising: receiving, by the at least one processor, a first voice input associated with a user vocalizing an authentication utterance; inputting, by the at least one processor, the first voice input into at least one trained voice authentication recognition machine learning model of the at least one trained voicerecognition machine learning model to output the voice authentication determination based at least in part on trained authentication parameters; wherein the trained authentication parameters are configured to determine the voice authentication determination based at least in part on a authentication similarity measure between the first voice input and at least one training authentication vocal pattern of a training authentication vocal pattern dataset; generate, based on the authentication similarity measure exceeding a predetermined authentication similarity measure threshold, a prompt to the user for the device command; receiving, by the at least one processor, a second voice input associated with the user vocalizing a device command utterance; inputting, by the at least one processor, the second voice input into at least one trained voice command recognition machine learning model of the at least one trained voice recognition machine learning model to output the device command based at least in part on trained device command parameters; and wherein the trained device command parameters are configured to determine the device command based at least in part on a device command similarity measure between the second voice input and at least one training device command vocal pattern of a training device command vocal pattern dataset.
14. The method of claim 8, further comprising: generating, by the at least one processor, a voice signal pattern representative of at least one characteristic of the voice signal based at least in part on at least one pre-processing algorithm; and inputting, by the at least one processor, the voice signal pattern into at least one trained voice recognition machine learning model to output the voice authentication determination and the device command based at least in part on the trained parameters of the at least one trained voice recognition model.
15. A non-transitory computer readable medium having executable instructions stored thereon, wherein the executable instructions are configured to cause at least one processor to perform steps comprising: receiving a voice signal from an audio source; inputting the voice signal into at least one trained voice recognition machine learning model to output a voice authentication determination and a device command based at least in part on trained parameters of the at least one trained voice recognition model; wherein the trained parameters of the at least one trained voice recognition model are configured to determine a measure of a match between the voice signal and at least one training vocal pattern of a training vocal pattern dataset; and instructing at least one renal care device with the device command based at least in part on the measure of the match of the voice signal exceeding a predetermined threshold.
16. The non-transitory computer readable medium of claim 15, wherein the at least one trained voice recognition machine learning model comprises a single voice recognition machine learning model configured to output the voice authentication determination and the device command.
17. The non-transitory computer readable medium of claim 15, wherein the at least one trained voice recognition machine learning model comprises: at least one trained voice command recognition machine learning model configured to output the device command; and at least one trained voice authentication machine learning model configured to output the voice authentication determination.
18. The non-transitory computer readable medium of claim 15, wherein the at least one trained voice recognition machine learning model comprises at least one of:at least one clustering model, or at least one neural network model.
19. The non-transitory computer readable medium of claim 15, wherein the executable instructions are further configured to cause the at least one processor to perform steps comprising: receiving a first voice input associated with a user vocalizing an authentication utterance; inputting the first voice input into at least one trained voice authentication recognition machine learning model of the at least one trained voice recognition machine learning model to output the voice authentication determination based at least in part on trained authentication parameters; wherein the trained authentication parameters are configured to determine the voice authentication determination based at least in part on a authentication similarity measure between the first voice input and at least one training authentication vocal pattern of a training authentication vocal pattern dataset; generate, based on the authentication similarity measure exceeding a predetermined authentication similarity measure threshold, a prompt to the user for the device command; receiving a second voice input associated with the user vocalizing a device command utterance; inputting the second voice input into at least one trained voice command recognition machine learning model of the at least one trained voice recognition machine learning model to output the device command based at least in part on trained device command parameters; and wherein the trained device command parameters are configured to determine the device command based at least in part on a device command similarity measure between the second voice input and at least one training device command vocal pattern of a training device command vocal pattern dataset.
0. The non-transitory computer readable medium of claim 15, wherein the executable instructions are further configured to cause the at least one processor to perform steps comprising: generating a voice signal pattern representative of at least one characteristic of the voice signal based at least in part on at least one pre-processing algorithm; and inputting the voice signal pattern into at least one trained voice recognition machine learning model to output the voice authentication determination and the device command based at least in part on the trained parameters of the at least one trained voice recognition model.
Citation Information
Patent Citations
Voice interface for a dialysis machine
US20190388599A1
Device and method for providing voice recognition service based on artificial intelligence
US20200005795A1