SKILL ACQUISITION FOR DYNAMIC TREATMENT CONCEPTS FURTHER APPLICATION INFORMATION

DE112023005296T5Pending Publication Date: 2025-10-23NEC LABORATORIES AMERICA INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
DE112023005296
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-14
Filing Date
2023-12-15
Publication Date
2025-10-23

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Methods and systems for training a machine learning model for healthcare include segmenting (304) a patient trajectory comprising a sequence of patient states and treatment actions. A machine learning model is trained (308) based on segments of the patient trajectory, including a prototype layer that learns prototype vectors representing respective classes of trajectory segments and an imitation learning layer that learns a policy for selecting a treatment action based on an input state and a skill embedding.
Need to check novelty before this filing date? Find Prior Art

Description

FURTHER APPLICATION INFORMATION

[0001] This application claims priority over U.S. Patent Application No. 63 / 434,133, filed on December 21, 2022, and U.S. Patent Application No. 18 / 539,506, filed on December 14, 2023, which are hereby incorporated in their entirety by reference. BACKGROUND Technical area

[0002] The present invention relates to machine learning systems and in particular systems for learning skills from historical records of medical treatments. Description of the state of the art

[0003] A dynamic treatment scheme is a set of sequential treatment decision rules that can be used to suggest treatment options for patients. While machine learning systems can be used to create a dynamic treatment scheme from historical records, the underlying reasons for these historical actions may be uninterpretable, posing challenges in real-world clinical scenarios. Furthermore, a single policy may not be sufficient to address a patient's evolving needs. SUMMARY

[0004] One method for training a healthcare machine learning model involves segmenting a patient history, which comprises a sequence of patient conditions and treatment interventions. A machine learning model is trained based on segments of the patient history, including a prototype layer that learns prototype vectors representing prototypes of the respective classes of history segments, and an imitation learning layer that learns a policy for selecting a treatment intervention based on an input condition and skill embedding.

[0005] A system for training a healthcare machine learning model comprises a hardware processor and memory that stores a computer program. When executed by the hardware processor, the computer program causes the hardware processor to segment a patient history, which includes a sequence of patient states and treatment interventions, and to train a machine learning model based on segments of the patient history. This includes a prototype layer that learns prototype vectors representing respective classes of history segments, and an imitation learning layer that learns a policy for selecting a treatment intervention based on an input state and skill embedding.

[0006] These and other features and advantages will become apparent from the following detailed description of the illustrative embodiments, which should be read in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] The disclosure is further explained in the following description of preferred embodiments with reference to the following figures, wherein Fig. 1 a block diagram of a health facility in which the imitation of skills is used to control treatment, according to an embodiment of the present invention; Fig. 2 a diagram of a patient with an automated treatment system according to an embodiment of the present invention; Fig. 3 a block diagram of a model for imitating skills which, according to an embodiment of the present invention, can operate on segments of a patient's progress; Fig. 4 a block / flow diagram of a method for training and using a model to mimic skills for patient treatment according to an embodiment of the present invention; Fig. Figure 5 is a block / flow diagram of a method for monitoring and treating a patient using a skill imitation model according to an embodiment of the present invention; Fig. Figure 6 is a block diagram of a computer system that can train and use a model to imitate skills for treating a patient, according to an embodiment of the present invention; Fig. Figure 7 is a diagram of a neural network architecture that can be used as part of a skill imitation model according to an embodiment of the present invention; and Fig. Figure 8 is a diagram of a deep neural network architecture that can be used as part of a skill imitation model according to an embodiment of the present invention. DETAILED DESCRIPTION OF PREFERRED EXECUTION FORMS

[0008] Imitation learning can be used to learn the association between states and actions when designing a treatment plan for a patient in healthcare. Imitation learning aims to replicate the behavior of experts, such as physicians' diagnostic and treatment actions, based on demonstrations from a series of records. To this end, an interpretable sequence modeling framework can be used to identify an expert's trajectory based on sequence data with temporal features. Learning and inference can be performed at the segment level, capturing temporal variations in states and identifying skills that are transferable across different trajectories. This offers flexibility compared to a trajectory-based formulation at the segment level, as trajectories can be represented by multiple segments.By aggregating the learned prototypes in each segment, the acquired skills serve as conditional information that implicitly guides a policy network to distinguish between patterns and provide precise treatment recommendations. These recommended treatments can then be interpreted at the segment level by tracing them back to the example segments.

[0009] Thus, an interpretable model for skill acquisition is provided for learning treatment policies, utilizing segment-level expert demonstrations and resulting in representative skills transferable across different trajectories. The model learns to capture example segments and construct a faithful skill embedding for imitation learning tasks.

[0010] With reference to Fig. Figure 1 shows a diagram of skill learning in the context of a healthcare facility 100. This can be used to monitor and treat multiple patients, for example, to respond to changes in their specific healthcare needs. The healthcare facility may include one or more healthcare professionals 102 who provide information on events and system status measurements to the skill simulation system 108. Treatment systems 104 can further monitor patient status to create medical records 106 and may be designed to automatically administer and adjust treatments as needed.

[0011] Based on information derived from at least the healthcare professionals 102, the treatment systems 104, and the medical records 106, the skill-imitation system 108 learns skills applied by the healthcare professionals 102 in response to evolving patient conditions. For example, the medical records 106 may contain the patient's historical health conditions (e.g., biometric information and a description of symptoms) and actions taken by the healthcare professionals 102 in response to those conditions.

[0012] The various elements of the healthcare facility 100 can communicate with each other via a network 110, for example, using a suitable wired or wireless communication protocol and medium. Thus, the competency simulation system 108 can access remotely stored medical records 106, communicate with the treatment systems 104, receive instructions, and send reports to healthcare professionals 102. In particular, the competency simulation system 108 can automatically trigger treatment changes for a patient by responding to new information from the medical records 106 and sending instructions to the treatment systems 104.

[0013] In some cases, the Skill Imitation System 108 can generate a specific treatment plan for the patient, including a prescription plan outlining medications to help treat the patient, a meal plan addressing the patient's nutritional needs, a rehabilitation plan detailing physical therapy and other activities necessary for the patient's recovery, and a discharge plan indicating whether the patient can return home, should remain in the facility, or should be transferred to another healthcare setting. The output of the Skill Imitation System 108 can therefore include one or a combination of the above-mentioned automated treatments and plan outputs.

[0014] With reference to Fig. Figure 2 depicts patient 202 within the context of a healthcare system. Patient 202 may, for example, be undergoing a hemodialysis session (also simply referred to as "dialysis"). During dialysis, a dialysis machine 204 automatically draws blood from the patient, processes and purifies it, and then returns the purified blood to the patient's body. Dialysis can last up to four hours and be performed every three days, although other durations and intervals are also possible. While dialysis is specifically considered here, it should be clear that any other suitable medical procedure or monitoring method could be used instead.

[0015] A health event related to treatment may occur in a patient 202 before, during, and after a dialysis session. Such health events can be dangerous for the patient 202 but can be predicted based on knowledge of past health events and the patient's current health data. Recommendation 208 may also include information about the nature of the predicted event as well as measurements of the patient's condition. It is expressly intended that this recommendation may be issued before the start of the dialysis session so that the treatment can be adjusted.

[0016] The recommendation can be based on various input data. This includes a static patient profile, such as age, sex, date of dialysis initiation, previous health events, etc. The information also includes dynamic data, such as dialysis measurements, which can be recorded at each dialysis session, including blood pressure, weight, venous pressure, blood test results, and the heart-thorax ratio (CTR). Blood tests can be performed regularly, for example, twice a month, and measure factors such as albumin, glucose, and platelet count. The CTR can also be measured regularly, for example, once a month. Dynamic information can also be recorded during the dialysis session, for example, using sensors in the 204 dialysis machine.The dynamic information can be modeled as time series over their respective frequencies.

[0017] Furthermore, the systems themselves can be monitored within a medical environment. For example, operating parameters of a dialysis machine 204 or another system in a hospital or other medical facility can be monitored along with a history of past events on the system in order to predict events such as those described below.

[0018] During treatment, the patient's condition can be continuously monitored, for example, by tracking their heart rate and other vital signs. If the patient's vital signs indicate an imminent or already occurring adverse health event, the treatment can be modified accordingly. For example, the treatment systems can automatically administer medication or terminate treatment in response to an adverse health event.

[0019] To train the skill-imitation system 108, the medical records 106 may contain a set of demonstration sequences of physicians, each of which is a sequence of condition-action pairs (s t , a t ) exhibits, where s t denotes the condition of a particular patient at a time t and a tThe treatment measure performed by a medical professional 102 is specified. The skill imitation system 108 learns a policy that can replicate the treatment measures taken by the medical professionals 102. Imitation learning can be based on step-by-step demonstrations without considering the patient's evolving symptoms and the corresponding treatments.

[0020] The imitation learning task can instead be formulated at the level of a sequence of states to leverage the sequential nature of the demonstrations. Each course can thus be divided into segments, and imitation learning can be performed for each segment. The skills in each segment can be representative and transferable to different courses. Each state in the segment is linked to a historical state segment from the same course, allowing the dynamics of the patient's state and the treatment demonstration to be utilized even at the segment level.

[0021] Starting from a series of segments {[(st(j),at(j))]t=1m}j=1n, By splitting the original data streams, where m is a fixed segment length and n is the number of segments, a series of prototypes can be determined that represent exemplary segments in the training data. These prototypes can be compiled as a skill embedding to facilitate imitation learning of a dynamic treatment regime at each step t.

[0022] With reference to Fig. Figure 3 shows an imitation learning model. The model's architecture can include a segment embedding layer 304, a prototype layer 306, and an imitation learning layer 308. The index j, which indicates the order of an instance, is omitted here for simplicity.

[0023] The input segment 302, which is assigned to the state of a patient in step t, can be described as [s t-m , ..., s t-1 , s t] ∈ ℝ{ m×dThe segment embedding layer 304 extracts an embedding of the input segment 302. This layer can comprise a multilayer perceptron (MLP) with input neurons that accept respective elements of the input segment 302, followed by a one-dimensional convolutional layer that encodes the segment information. By sharing the same MLP encoder, the state embedding can be performed using the following methods: t ∈ ℝ d' , in step t as f t = MLP(s t ) are generated. The segment of the coded state in the embedding space [f t-m , ..., f t-1 , f t ] ∈ ℝ m×d' is fed into the one-dimensional folding layer to enable segment embedding. t ∈ ℝ h×1 to generate: zt=CONCATi=0h−1(Wi⋆ft−m+1:t+bi) where “W i ∈ ℝm×d' " denotes the i-th convolution kernel consisting of a total of h kernels, "b i ∈ ℝ” denotes a corresponding bias term, the operator “*” provides the sum of the row-wise cross-correlations, “d'” is the dimensionality of the states after processing by the MLP encoder, and “CONCAT” provides the concatenation of all convolution results on h kernels.

[0024] The segment embedding layer can also be implemented in other ways, for example, using a long-short-term memory (LSTM) or gated recurrent unit (GRU) architecture, or using a transformer architecture. However, since long segments are relatively rare in dynamic treatment regimes, the one-dimensional convolution layer may be more efficient and effective for extracting the relevant embedding for short segments.

[0025] The prototype layer 306 can have k prototype vectors [p1, ..., p k ] ∈ ℝ k×huse parameters that can be considered trainable model parameters and have the same dimensionality as the segment embedding. t Learning from prototypes has advantages for interpretability, as the original data segments onto which the prototype vectors are projected can be retained for analysis.

[0026] Each prototype vector represents a class of example segments that reflect the patient's condition at a specific phase. The degree of similarity between the segment embedding z t and each prototype vector can be determined as follows: Sim(zt,pi)=e(−‖zt−pi‖22), where p i denotes the i-th prototype vector, ||·||2 denotes the L2 norm, and the exponential function confines the degree of similarity to a limited range for reasons of numerical stability.

[0027] The similarity values ​​of all prototype embedding pairs are scaled between 0 and 1 ([0,1]), and the resulting Skored value for the i-th prototype embedding pair is: Sim^(zt,pi)=e−‖zt−pi‖22∑i=1ke−‖zt−pi‖22

[0028] All scaled values ​​form a weighting vector W p ∈ ℝ k×1 as follows: Wp=[Sim^(zt,p1),...,Sim^(zt,pk)]. Based on W p can the skill embedding o t ∈ ℝ 1×h for the segment can be constructed by the weighted combination of the vectors p: ot=WpT⋅p where the operator · denotes the inner product operation. Instead of the original segment embedding z t All prototype vectors p will be incorporated into the final skill embedding o t included, so that the skill embedding can be interpreted based on the original state segments via the prototype segment mapping.

[0029] The imitation learning layer 308 modifies a flat policy network π. θ (a t |s t ), which is parameterized by θ, and learns mappings of s t to a t , by embedding skills o t includes, which serves as a high-level indicator that guides the agent to mimic expert demonstrations derived from the expert policy π E (a t |s t ) were taken. The patient's status s t is associated with skills embedding o t chained as conditional information and linked to the context-related policy π θ (a t |o t ,s t ) forwarded, from which the primitive initial action a t for the dynamic treatment task, the following can be derived: at←πθ(at|ot,st)

[0030] The context-related policy network can be built on the basis of a behavior cloning model that aims to mimic the physician's medication at each time step t by treating it as a supervised learning problem. The actual policy network can be implemented through a four-layer MLP. Interpretable skill learning explicitly models the segments of a patient's states with prototype vectors that are learned and regulated through behavior cloning and multiple interpretable learning objectives, based on which the user can obtain an explanation by tracing back to the training segments.

[0031] With reference to Fig. Section 4 now presents a procedure for training and using a skill imitation model. In Block 402, the model is trained using a series of historical training examples that include sequences of patient conditions and treatments performed by healthcare professionals, forming patient histories. Once the model has been trained, it can be deployed in a skill imitation system in a healthcare facility. There, it can be used to monitor and treat patients, coordinating the treatments of the treatment systems and healthcare professionals.

[0032] During training (402), the historical data is divided into fixed-size segments. Imitation learning is then performed to train the policy as described above, in order to minimize a specific objective function. L=LIM+λ1Lcluster+λ2Levidence+λ3Ldiversity where “λ1”, “λ2” and “λ3” are weighting coefficients between zero and one.

[0033] The objective function includes several terms, including an imitation learning term. "LIM" For a number of segments of size n, the context-related policy aims to mimic the physician's demonstration at the segment level in a supervised manner: LIM=∑j=1n∑t=1mπE(at(j)|st(j))log πθ(at(j)|ot(j),st(j)) where m is the length of the segment and “π E “refers to the expert policy from which the demonstrations are selected.”

[0034] To improve the interpretability of the skill learning model, three regularization components can be used for learning prototype vectors, including terms relating to the clustering structure of the segment embedding, the segment prototype evidence, and the variety of prototypes. The clustering structure regulation ensures that the segment embedding... t as close as possible to the next prototype: Lcluster=∑j=1nmini∈[k]‖zt(j)−pi‖22

[0035] Where [k] denotes the set of integers with the maximal element k representing all prototype vectors.

[0036] Regularizing the prototype-segment assignment imposes a dual optimization goal with respect to the segment embedding and the prototype vectors. Regularization encourages each prototype vector to be as similar as possible to a segment embedding. Levidence=∑i=1kminj∈[n]‖pi−zt(j)‖22

[0037] The clustering structure and the prototype-segment-proof terms interact with each other and together restrict the learning of both the segment embedding layer 308 and the prototype layer 306 towards a clear and interpretable structure.

[0038] The similarity between each pair of prototype vectors can also be penalized, as indistinguishable prototype vectors representing similar patients may be redundant, while encouraging prototype diversity improves generalizability when new segments and patient histories are encountered. A diversity regulation deadline can be imposed as follows: Ldiversity=∑i=1k∑i'≠ikmax(0,dmin−‖pi−pi'‖22)

[0039] Whereby d minThis is an approximate threshold that determines whether a given prototype pair is penalized. The restrictions mentioned above are applied to the last step t = m within each input segment, as this contains all the information for the segment without adding the patient state(s) from the previous segment in the same flow.

[0040] Training 402 will continue until the training loss L The process converges, with the prototype vectors being optimized to closely match the segment embeddings from the training data. However, the prototype vectors are not interpretable at this stage, as there is no correspondence between them and the actual segments. To establish the association between prototypes and segments in the training data, each prototype p can be i to be assigned to its nearest segment in latent space: pi←arg minzt(j)∈Z‖pi−zt(j)‖22 whereupon. "Z train A set of segment embeddings is created by feeding all segments of the training dataset into the segment embedding layer 304. Performance tends to increase with the number of prototypes, but then stabilizes, so that beyond a certain point the additional effort of using more prototypes is no longer justified. In some embodiments, "k = 25" prototypes can be used.

[0041] With reference to Fig. 5. Further details regarding monitoring and treatment 406 are shown. Block 502 measures an updated patient status, for example, using treatment systems 104, and adds the updated patient status information to the medical records 106. This information may include health indicators such as heart rate, blood pressure, laboratory results, as well as the patient's medication intake and treatment history. The collected data constitute the status t en of the patient.

[0042] The new patient status can be entered into the trained model at block 504, which generates a corresponding competency output. The trained model converts the status into an embedding, which is then used to identify the most suitable prototype. The model's imitation learning is activated and generates action plans based on both the input status and the identified prototype. The action plans can include recommendations or treatment instructions, such as specifying particular medication dosages. The output competency indicates a treatment to be performed, and block 506 automatically carries out the treatment, for example, by sending an instruction to the treatment system 502. The treatment can also be forwarded to healthcare professionals 102. This treatment information can assist healthcare professionals 102 in making decisions regarding patient care.

[0043] With reference to Fig. Figure 6 shows an exemplary computing device 600 according to an embodiment of the present invention. The computing device 600 is configured to perform competence recognition.

[0044] The Computing Unit 600 can be implemented as any type of computing or computer device capable of performing the functions described herein, including, but not limited to, a computer, server, rack-based server, blade server, workstation, desktop computer, laptop computer, notebook computer, tablet computer, mobile computing device, portable computing device, network device, web device, distributed computing system, processor-based system, and / or consumer electronics device. Additionally or alternatively, the Computing Unit 600 can be implemented as one or more compute sleds, storage sleds, or other racks, sleds, computer enclosures, or other components of a physically disaggregated computing device.

[0045] As in Fig. As shown in Figure 6, the computing device 600 includes, by way of example, the processor 610, an input / output subsystem 620, a memory 630, a data storage device 640, and a communication subsystem 650, and / or other components and devices commonly found in a server or similar computing device. In other embodiments, the computing device 600 may include further or additional components, such as those commonly found in a server computer (e.g., various input / output devices). Furthermore, in some embodiments, one or more of the illustrative components may be integrated into another component or otherwise form part of it. For example, in some embodiments, the memory 630, or parts thereof, may be integrated into the processor 610.

[0046] The 610 processor can be implemented as any type of processor capable of performing the functions described here. The 610 processor can be implemented as a single processor, as multiple processors, as one or more central processing units (CPUs), as one or more graphics processing units (GPUs), as single-core or multi-core processors, as digital signal processors, as microcontrollers, or as other processors or processing / control circuits.

[0047] The Memory 630 can be implemented as any type of volatile or non-volatile memory or data storage capable of performing the functions described herein. In operation, the Memory 630 can store various data and software used during the operation of the Computing Unit 600, such as operating systems, applications, programs, libraries, and drivers. The Memory 630 is connected to the Processor 610 via the I / O Subsystem 620, which can be implemented as circuitry and / or components to facilitate input / output operations with the Processor 610, the Memory 630, and other components of the Computing Unit 600. For example, the I / O Subsystem 620 can be implemented as or otherwise include memory controller hubs, I / O control hubs, platform controller hubs, integrated control circuits, firmware devices, communication links (e.g.,Point-to-point connections, bus connections, wires, cables, optical fibers, conductive traces on printed circuit boards, etc.) and / or other components and subsystems to facilitate I / O operations. In some embodiments, the I / O subsystem 620 can form part of a system-on-a-chip (SoC) and be integrated on a single integrated circuit chip together with the processor 610, the memory 630, and other components of the computing unit 600.

[0048] The data storage device 640 can be implemented as any type of device or devices configured for short- or long-term data storage, such as storage devices and circuits, memory cards, hard disk drives, solid-state drives, or other data storage devices. The data storage device 640 can store program code 640A for training a model, 640B for predicting an event, and / or 640C for executing a corrective action in response to the predicted event. Any or all of these blocks of program code can be contained within a given computer system. The communication subsystem 650 of the computing unit 600 can be implemented as any network interface controller or other communication circuit, device, or collection thereof that enables communication between the computing unit 600 and other remote devices over a network.The Communication Subsystem 650 can be configured to use one or more communication technologies (e.g., wired or wireless communication) and associated protocols (e.g., Ethernet, InfiniBand®, Bluetooth®, Wi-Fi®, WiMAX, etc.) to effect such communication.

[0049] As shown, the computing unit 600 can also include one or more peripheral devices 660. The peripheral devices 660 can include any number of additional input / output devices, interface devices, and / or other peripherals. For example, the peripheral devices 660 can include a display, a touchscreen, a graphics circuit, a keyboard, a mouse, a speaker system, a microphone, a network interface, and / or other input / output devices, interface devices, and / or peripherals.

[0050] Naturally, the computing unit 600 can also include other elements (not shown) that are readily conceivable to a person skilled in the art, as well as omit certain elements. For example, various other sensors, input devices, and / or output devices can be included in the computing unit 600, depending on its specific implementation, as is readily apparent to a person skilled in the art. For instance, different types of wireless and / or wired input and / or output devices can be used. Furthermore, additional processors, controllers, memory, etc., can also be used in various configurations. These and other variants of the processing system 600 are readily conceivable to a person skilled in the art of the present invention based on the teachings contained herein.

[0051] With reference to the Fig. 7 and Fig.Figure 8 presents exemplary neural network architectures that can be used to implement parts of the present models, such as the segment embedding layer 304. A neural network is a generalized system that improves its functionality and accuracy through the application of additional empirical data. The neural network is trained by this application of empirical data. During training, the neural network stores and adjusts a variety of weights that are applied to the incoming empirical data. By applying the adjusted weights to the data, the data can be identified as belonging to a specific predefined class from a set of classes, or a probability can be output indicating that the input data belongs to each of the classes.

[0052] The empirical data, also called training data, from a series of examples can be formatted as a sequence of values ​​and fed into the input of the neural network. Each example can be associated with a known result or output. Each example can be represented as a pair (x, y), where x represents the input data and y the known output. The input data can comprise a variety of different data types and contain multiple distinct values. The network can have an input node for each value that makes up the example's input data, and each input value can be assigned a separate weight. The input data can be formatted as a vector, array, or string, for example, depending on the architecture of the constructed and trained neural network.

[0053] The neural network "learns" by comparing the outputs generated from the input data with the known values ​​of the examples and adjusting the stored weights to minimize the differences between the output values ​​and the known values. These adjustments can be made to the stored weights by backpropagation, where the effect of the weights on the output values ​​can be determined by calculating the mathematical gradient and adjusting the weights in a way that shifts the output toward a minimal difference. This optimization, known as gradient descent, is a non-restrictive example of how training can be performed. A subset of examples with known values, not used for training, can be used to test and validate the accuracy of the neural network.

[0054] During operation, the trained neural network can be applied to new data not previously used for training or validation through generalization. The adjusted weights of the neural network can be applied to the new data, with the weights estimating a function derived from the training examples. The parameters of the estimated function, captured by the weights, are based on statistical inference.

[0055] In layered neural networks, nodes are arranged in layers. An example simple neural network has an input layer 720 consisting of source nodes 722 and a single compute layer 730 with one or more compute nodes 732 that also function as output nodes, with a single compute node 732 for each possible category into which the input sample could be classified. An input layer 720 can have a number of source nodes 722 corresponding to the number of data values ​​712 in the input data 710. The data values ​​712 in the input data 710 can be represented as a column vector. Each compute node 732 in the compute layer 730 generates a linear combination of weighted values ​​from the input data 710 fed into the input nodes 720 and applies a nonlinear activation function that is differentiable on the sum.The exemplary simple neural network can perform a classification of linearly separable examples (e.g., patterns).

[0056] A deep neural network, such as a multi-layered perceptron, can have an input layer 720 consisting of source nodes 722, one or more computation layers 730 with one or more computation nodes 732, and an output layer 740, with a single output node 742 for each possible category into which the input sample could be classified. An input layer 720 can have a number of source nodes 722 corresponding to the number of data values ​​712 in the input data 710. The computation nodes 732 in the computation layer(s) 730 can also be called hidden layers, as they are located between the source nodes 722 and the output nodes 742 and cannot be directly observed.Each node 732, 742 in a computation layer generates a linear combination of weighted values ​​from the values ​​output by the nodes of a previous layer and applies a nonlinear activation function that is differentiable over the domain of the linear combination. The weights applied to the value of each previous node can be represented, for example, by w1, w2, ... w. n-1 , w n The output layer provides the network's overall response to the input data. A deep neural network can be fully connected, with each node in a computation layer connected to every other node in the previous layer, or it can have other configurations of connections between layers. If connections between nodes are missing, the network is said to be partially connected.

[0057] Training a deep neural network can involve two phases: a forward phase in which the weights of each node are set and the input is passed through the network, and a backward phase in which an error value is passed backward through the network and the weight values ​​are updated.

[0058] The computation nodes 732 in one or more (hidden) layers 730 perform a nonlinear transformation of the input data 712, which generates a feature space. The classes or categories can be separated more easily in the feature space than in the original data space.

[0059] The embodiments described here can consist entirely of hardware, entirely of software, or of a combination of hardware and software elements. In a preferred embodiment, the present invention is implemented in software, which includes, among other things, firmware, resident software, microcode, etc.

[0060] Embodiments may include a computer program product accessible from a computer-readable or computer-transparent medium that provides program code for use by or in conjunction with a computer or any command-execution system. A computer-readable or computer-transparent medium may include any device that stores, transmits, disseminates, or transports the program for use by or in conjunction with the command-execution system, device, or apparatus. The medium may be a magnetic, optical, electronic, electromagnetic, infrared, or semiconductor system (or device or apparatus) or a dissemination medium.The medium can include a computer-readable storage medium such as semiconductor or solid-state memory, magnetic tape, removable computer disk, random access memory (RAM), read-only memory (ROM), rigid magnetic disk, optical disk, etc.

[0061] Any computer program can be stored in a machine-readable storage medium or device (e.g., a program memory or a magnetic disk) that can be read by a general-purpose or programmable special-purpose computer to configure and control the operation of a computer when the storage medium or device is read by the computer to perform the methods described herein. The system according to the invention can also be considered to be embodied in a computer-readable storage medium configured with a computer program, wherein the storage medium so configured causes a computer to operate in a specific and predefined manner to perform the functions described herein.

[0062] A data processing system capable of storing and / or executing program code may include at least one processor, which is directly or indirectly coupled to memory elements via a system bus. The memory elements may include local memory used during the actual execution of the program code, mass storage, and cache memory, which allows for the temporary storage of at least part of the program code to reduce the number of code retrievals from mass storage during execution. Input / output devices, or I / O devices (including but not limited to keyboards, displays, pointing devices, etc.), may be connected to the system either directly or via intermediary I / O controllers.

[0063] Network adapters can also be connected to the system to link the data processing system to other data processing systems, remote printers, or storage devices via intermediary private or public networks. Modems, cable modems, and Ethernet cards are just some of the currently available types of network adapters.

[0064] As used here, the term "hardware processor subsystem" or "hardware processor" can refer to a processor, memory, software, or combinations thereof, working together to perform one or more specific tasks. In useful embodiments, the hardware processor subsystem may include one or more data processing elements (e.g., logic circuits, processing circuits, instruction execution devices, etc.). The one or more data processing elements may be contained within a central processing unit, a graphics processing unit, and / or a separate processor- or computationally-element-based control unit (e.g., logic gates, etc.). The hardware processor subsystem may include one or more integrated memories (e.g., caches, dedicated memory arrays, read-only memories, etc.).In some embodiments, the hardware processor subsystem may include one or more memories, which may be located on or off the circuit board, or which are reserved for use by the hardware processor subsystem (e.g., ROM, RAM, basic input / output system (BIOS), etc.).

[0065] In some embodiments, the hardware processor subsystem can contain and execute one or more software elements. These software elements may include an operating system and / or one or more applications and / or specific code to achieve a particular result.

[0066] In other embodiments, the hardware processor subsystem may include dedicated, specialized circuits that perform one or more electronic processing functions to achieve a specific result. Such circuits may include one or more application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), and / or programmable logic arrays (PLAs).

[0067] These and other variants of a hardware processor subsystem are also provided according to embodiments of the present invention.

[0068] References in the description to "an embodiment" or "an embodiment" of the present invention, as well as to other variants thereof, mean that a particular feature, structure, property, etc., described in connection with the embodiment, is included in at least one embodiment of the present invention. Therefore, the expressions "in an embodiment" or "in an embodiment," as well as any other variants appearing at different points in the description, do not necessarily all refer to the same embodiment. It should be noted, however, that features of one or more embodiments can be combined, taking into account the teachings of the present invention contained herein.

[0069] It should be noted that the use of, for example, " / ", "and / or", and "at least one of" in the cases "A / B", "A and / or B", and "at least one of A and B" is intended to include the selection of the first listed option (A) alone, the selection of the second listed option (B) alone, or the selection of both options (A and B). As a further example, in the case of "A, B and / or C" and "at least one of A, B, and C", such wording is intended to include the selection of the first listed option (A) alone, the selection of the second listed option (B) alone, the selection of the third listed option (C) alone, the selection of the first and second listed options (A and B), the selection of the first and third listed options (A and C), the selection of the second and third listed options (B and C), or the selection of all three options (A, B, and C). This can be extended to any number of listed items.

[0070] The foregoing is to be understood in every respect as illustrative and exemplary, but not limiting, and the scope of the invention disclosed herein is not to be determined from the detailed description, but from the claims as interpreted in accordance with the full scope of patent law. It is understood that the embodiments shown and described herein serve only to illustrate the present invention and that those skilled in the art may make various modifications without departing from the scope and spirit of the invention. Skilled in the art could implement various other combinations of features without departing from the scope and spirit of the invention. Having thus described aspects of the invention with the details and specificities required by patent law, the claims to be protected by the patent are set forth in the accompanying claims. QUOTES INCLUDED IN THE DESCRIPTION

[0000] This list of documents cited by the applicant was automatically generated and is included solely for the reader's convenience. The list is not part of the German patent or utility model application. The DPMA accepts no liability for any errors or omissions. Cited patent literature

[0000] US 63 / 434,133

[0001] US 18 / 539,506

[0001]

Claims

[1] A computer-implemented method for training a machine learning model for healthcare, comprising: Segmentation (304) of a patient's medical history, comprising a sequence of patient status and treatment interventions; and Training (308) a machine learning model based on segments of the patient's history, including a prototype layer that learns prototype vectors representing respective classes of history segments, and an imitation learning layer that learns a policy for selecting a treatment action based on an input state and skill embedding. [2] The method of claim 1, further comprising embedding the segmented patient history using a segment embedding layer of the machine learning model. [3] Method according to claim 2, wherein the segment embedding layer comprises a multilayer perceptron and a one-dimensional folding layer. [4] The method of claim 1, further comprising: Measuring a patient's condition information; Selecting a treatment measure based on an ability predicted by the trained model, based on the measured condition information; and Notifying a medical professional about the treatment measure to assist the medical professional in making patient management decisions. [5] Method according to claim 1, wherein the skill embedding comprises a weighted combination of the prototype vectors based on how similar the prototype vectors are to the segmented patient trajectory. [6] Method according to claim 1, wherein the training of the machine learning model comprises minimizing a loss function comprising an imitation learning concept, a clustering structure regularization concept, a prototype segment document regularization concept and a diversity regularization concept. [7] Method according to claim 6, wherein the imitation learning concept is expressed as: LIM=∑j=1k∑t=1kπE(at(j)|st(j))log πθ(at(j)|ot(j),st(j)) where m is the length of a segment, n is the number of segments, π E an expert policy, at(j) an action performed in step t for segment j is, st(j) a patient state at step t for segment j, π θ a learned policy is and ot(j) a skill embedding in step t for segment j. [8] Method according to claim 6, wherein the term for regularizing the clustering structure is expressed as follows: Lcluster=∑j=1nmini∈[k]‖zt(j)−pi‖22 The prototype segment document regularization expression is expressed as follows: Levidence=∑i=1kminj∈[n]‖pi−zt(j)‖22 and the diversity regulation term is expressed as follows: Ldiversity=∑i=1k∑i'≠1kmax(0,dmin−‖pi−pi'‖22) where n is a number of segments, k is a number of prototype vectors, p i an i-th prototype vector is zt(j) a segment embedding in step t for segment j and d min is an approximate limit value. [9] Method according to claim 1, wherein the treatment measure comprises at least one of the following: a prescription plan, a meal plan, a rehabilitation plan and a discharge target plan. [10] Method according to claim 1, wherein the treatment measure comprises an instruction to a treatment system to automatically administer a treatment to a patient. [11] System for training a machine learning model for healthcare, comprising: a hardware processor (610); and a memory (640) that stores a computer program which, when executed by the hardware processor, causes the hardware processor to: Segmentation (304) of a patient's medical history, comprising a sequence of patient conditions and treatment interventions; and to train a machine learning model based on segments of the patient's history (308), including a prototype layer that learns prototype vectors representing respective classes of history segments, and an imitation learning layer that learns a policy to select a treatment action based on an input state and skill embedding. [12] System according to claim 11, wherein the computer program further instructs the hardware processor to embed the segmented patient history using a segment embedding layer of the machine learning model. [13] System according to claim 12, wherein the segment embedding layer comprises a multilayer perceptron and a one-dimensional folding layer. [14] System according to claim 11, wherein the prototype layer determines a similarity between the segments and the prototype vectors. [15] System according to claim 11, wherein the skill embedding comprises a weighted combination of the prototype vectors based on how similar the prototype vectors are to the segmented patient trajectory. [16] System according to claim 11, wherein the computer program further instructs the hardware processor to minimize a loss function comprising an imitation learning term, a clustering structure regularization term, a prototype segment document regularization term and a diversity regularization term. [17] System according to claim 16, wherein the imitation learning concept is expressed as follows: LIM=∑j=1n∑t=1mπE(at(j)|st(j))log πθ(at(j)|ot(j),st(j)) where m is the length of a segment, n is the number of segments, π E an expert policy, at(j) an action performed in step t for segment j is, st(j) a patient state at step t for segment j, π θ a learned policy is and ot(j) a skill embedding in step t for segment j. [18] The system according to claim 16, wherein the term for regularizing the clustering structure is expressed as follows: Lcluster=∑j=1nmini∈[k]‖zt(j)−pi‖22 The prototype segment document regularization expression is expressed as follows: Levidence=∑j=1kmini∈[k]‖pi−zt(j)‖22 and the diversity regulation term is expressed as follows: Ldiversity=∑i=1k∑i'≠ikmax(0,dmin−‖pi−pi'‖22) where n is a number of segments, k is a number of prototype vectors, p i an i-th prototype vector is zt(j) a segment embedding in step t for segment j and d min is an approximate limit value. [19] System according to claim 11, wherein the treatment measure comprises at least one of the following: a prescription plan, a meal plan, a rehabilitation plan and a discharge target plan. [20] System according to claim 11, wherein the treatment measure comprises an instruction to a treatment system to automatically administer a treatment to a patient.

Citation Information

Patent Citations

  • 18/539,506

  • 63/434,133