Deep learning architecture for analyzing unstructured data
A deep learning model architecture processes medical records into snippets using LSTM-based pipelines and regular expressions to effectively identify patient attributes in large volumes of unstructured data, addressing the inefficiencies of existing techniques.
Patent Information
- Application Number
- JP2022503903
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-05-18
- Filing Date
- 2020-07-23
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2040-07-23
AI Technical Summary
Existing machine learning techniques, such as LSTM neural networks, are ineffective for processing large volumes of unstructured medical data due to their inability to handle the vast amount of text in patient medical records, making it difficult to identify specific patient attributes.
A deep learning model architecture that parses medical records into snippets, processes them using an LSTM-based pipeline, and combines snippet representations for classification, utilizing regular expressions to extract relevant text and implement a specific model architecture to learn from these snippets.
Enables accurate and efficient identification of patient attributes by analyzing large medical records, improving the accuracy and efficiency of information extraction from unstructured data.
Smart Images

Figure 0007704731000001 
Figure 0007704731000002 
Figure 0007704731000003
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications
[0001] This application claims the benefit of priority of U.S. Provisional Patent Application No. 62 / 878,024, filed on Jul. 24, 2019, and U.S. Provisional Patent Application No. 63 / 026,418, filed on May 18, 2020. The content of the above applications is hereby incorporated by reference in its entirety into this specification.
[0002] Background Technical Field
[0002] This disclosure relates to identifying the representation of attributes within a large set of unstructured data, and more particularly, to the architecture of deep learning models configured to analyze data.
Background Art
[0003] Background Information
[0001] Information extraction is an increasingly important task that enables software applications to process information from unstructured documents. It has significant advantages for large - scale data processing in many industries, including the medical industry. For example, patient medical records, which can include hundreds of millions of unstructured text documents, often contain valuable insights relevant to a patient's treatment. However, when examining a large group of medical data, it can be difficult to identify specific attributes exhibited by patients. For example, this may require searching through thousands of medical documents, each of which may contain hundreds of pages of unstructured text. Additionally, due to the nature of the documents, information regarding patient attributes is often presented as handwritten notes or other text, which can make the automation of this process more difficult.
[0004]
[0002] Some solutions may involve developing a machine learning model to determine whether a patient is associated with a particular attribute. For example, the model can be trained based on a set of medical records where it is known whether the patient has been tested for a particular condition. However, many machine learning techniques are not prepared to process the large amounts of data required in the medical industry or other industries related to very large unstructured documents. Many of the information extraction techniques developed are effective for short documents (e.g., product reviews, social media posts, search engine queries) and often do not generalize well to longer documents. For example, long short-term models (LSTMs) or other recurrent neural networks may provide certain advantages when analyzing a series of medical records. However, due to the vast amount of unstructured text data that needs to be processed, traditional LSTM neural networks are not effective for this application.
[0005]
[0003] Therefore, an improved approach for identifying patients with specific medical attributes is needed. The solution should enable the development of a deep learning model architecture that allows for effective information extraction from long documents.
Summary of the Invention
Means for Solving the Problems
[0006] Overview
[0004] Embodiments consistent with the present disclosure include a system and method for determining probabilities associated with patient attributes. In an embodiment, the model assistance system may include at least one processor. The processor may be programmed to access a database storing at least one unstructured medical record associated with a patient, analyze the at least one unstructured medical record to identify a plurality of snippets of information within the at least one unstructured medical record associated with patient attributes. The processor may be programmed to generate a snippet vector including a plurality of snippet vector components based on each of the plurality of snippets, the plurality of snippet vector components including weighted values associated with at least one word included in the snippet, analyze the snippet vector to generate a summary vector including a plurality of summary vector components, each of the plurality of summary vector components being associated with a corresponding snippet vector component and further programmed to be determined based on an analysis of the corresponding snippet vector component. The processor may be further programmed to generate at least one output indicative of a probability associated with patient attributes based on the summary vector.
[0007]
[0005] In another embodiment, a computer-implemented method for determining a probability associated with a patient attribute. The method can include accessing a database storing at least one unstructured medical record and analyzing the at least one unstructured medical record to identify a plurality of snippets of information within the at least one unstructured medical record associated with the patient attribute. The method can further include generating a snippet vector including a plurality of snippet vector components based on each of the plurality of snippets, wherein the plurality of snippet vector components include weighted values associated with at least one word included in the snippet; analyzing the snippet vector to generate a summary vector including a plurality of summary vector components, wherein each of the plurality of summary vector components is associated with a corresponding snippet vector component and is determined based on an analysis of the corresponding snippet vector component; and generating at least one output indicative of a probability associated with the attribute based on the summary vector.
[0008]
[0006] Consistent with other disclosed embodiments, a non-transitory computer-readable storage medium may store program instructions that, when executed by at least one processing device, perform any of the methods described herein.
[0009] Brief Description of the Drawings
[0007] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate and, together with the description, serve to explain the principles of the various exemplary embodiments.
Brief Description of the Drawings
[0010]
Figure 1
[0008] A block diagram showing an exemplary system environment for implementing embodiments consistent with the present disclosure.
Figure 2
[0009] A block diagram showing an exemplary medical record for a patient consistent with the present disclosure.
Figure 3A
Figure 3B
[0011] Consistent with the disclosed embodiments, snippets are shown as examples that can be extracted from a document.
Figure 4A
[0012] Consistent with the disclosed embodiments, it is a block diagram showing a neural network as an example that operates on a single snippet.
Figure 4B
[0013] Consistent with the disclosed embodiments, a process is shown as an example for combining the hidden states of a neural network model using an attention mechanism.
Figure 5
[0014] Consistent with the disclosed embodiments, it is a block diagram showing a process as an example for generating a summary vector and probabilities based on a plurality of snippet vectors.
Figure 6
[0015] Consistent with the disclosed embodiments, it is a flowchart showing a process as an example for determining the probability associated with an attribute.
Best Mode for Carrying Out the Invention
[0011] Detailed Description
[0016] The following detailed description refers to the accompanying drawings. Whenever possible, the same reference numbers are used in the drawings and the following description to refer to the same or similar parts. Although several exemplary embodiments are described herein, modifications, adaptations, and other embodiments are possible. For example, substitutions, additions, or modifications may be made to the components shown in the drawings, and the exemplary methods described herein may be modified by substituting steps, changing the order, removing, or adding steps to the disclosed methods. Therefore, the following detailed description is not limited to the disclosed embodiments and examples. Instead, the appropriate scope is defined by the appended claims.
[0012]
[0017] Embodiments herein include a computer-implemented method, a tangible non-transitory computer-readable medium, and a system. The computer-implemented method can be executed, for example, by at least one processor (e.g., a processing device) that receives instructions from a non-transitory computer-readable storage medium. Similarly, a system consistent with the present disclosure may include at least one processor (e.g., a processing device) and a memory, which may be a non-transitory computer-readable storage medium. As used herein, a non-transitory computer-readable storage medium refers to any type of physical memory that can store information or data readable by at least one processor. Examples include random access memory (RAM), read-only memory (ROM), volatile memory, non-volatile memory, hard drives, CD ROMs, DVDs, flash drives, disks, and any other known physical storage media. Singular terms such as "memory" and "computer-readable storage medium" may further refer to multiple structures, such multiple memories and / or computer-readable storage media. As referenced herein, "memory" may include any type of computer-readable storage medium unless otherwise specified. A computer-readable storage medium may store instructions for execution by at least one processor that cause the processor to perform steps or stages consistent with the embodiments herein. Additionally, one or more computer-readable storage media may be used when implementing a computer-implemented method. The term "computer-readable storage medium" is to be understood to include tangible articles and to exclude carrier waves and transient signals.
[0013]
[0018] Embodiments of the present disclosure provide a system and method for determining probabilities associated with patient attributes. The users of the disclosed system and method can include any individual who may desire to access and / or analyze patient data. Thus, throughout the present disclosure, references to "users" of the disclosed system and method can include any individual, such as physicians, researchers, quality assurance departments of healthcare facilities, and / or any other individual.
[0014]
[0019] FIG. 1 shows an exemplary system environment 100 for implementing embodiments consistent with the present disclosure, which will be described in detail below. As shown in FIG. 1, the system environment 100 may include several components, including a client device 110, a data source 120, a system 130, and / or a network 140. The number and arrangement of these components are exemplary and are provided for illustrative purposes. It should be understood from the present disclosure that other arrangements and numbers of components may be used without departing from the teachings and embodiments of the present disclosure.
[0015]
[0020] As shown in FIG. 1, the exemplary system environment 100 may include a system 130. The system 130 may include one or more server systems, databases, and / or computing systems configured to receive information from entities via a network, process the information, store the information, and display / transmit the information to other entities via the network. Thus, in some embodiments, the network may facilitate cloud sharing, storage, and / or computing. In one embodiment, the system 130 may include a processing engine 131 and one or more databases 132 shown in the area delimited by the dashed line representing the system 130. The processing engine 140 may include at least one processing device, such as one or more general-purpose processors, such as a central processing unit (CPU), a graphics processing unit (GPU), and / or one or more dedicated processors, such as an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA).
[0016]
[0021] The various components of the system environment 100 may include hardware, software, and / or firmware, including memory, a central processing unit (CPU), and / or a user interface. The memory may include any type of RAM or ROM embodied in a physical storage medium, such as magnetic storage including floppy disks, hard disks, or magnetic tapes, semiconductor storage such as solid state disks (SSDs) or flash memory, optical disk storage, or magneto-optical disk storage. The CPU may include one or more processors for processing data according to a set of programmable instructions or software stored in the memory. The functionality of each processor may be provided by a single dedicated processor or multiple processors. Further, the processor may include, without limitation, digital signal processor (DSP) hardware, or any other hardware capable of executing software. The optional user interface may include any type or combination of input / output devices, such as a display monitor, keyboard, and / or mouse.
[0017]
[0022] Data transmitted and / or exchanged within the system environment 100 may occur via a data interface. As used herein, a data interface may include any boundary across which two or more components of the system environment 100 exchange data. For example, the environment 100 may exchange data between software, hardware, databases, devices, humans, or any combination of the foregoing. Further, it should be understood that any suitable configuration of software, processors, data storage devices, and networks may be selected to implement the components of the system environment 100 and the features of the related embodiments.
[0018]
[0023] The components of environment 100 (including system 130, client device 110, and data source 120) may communicate with each other or communicate with other components through network 140. Network 140 may include various types of networks, such as the Internet, a wired wide area network (WAN), a wired local area network (LAN), a wireless WAN (e.g., WiMAX), a wireless LAN (e.g., IEEE 802.11, etc.), a mesh network, a mobile / cellular network, an enterprise or private data network, a storage area network, a virtual private network using a public network, short-range wireless communication technologies (e.g., Bluetooth, infrared, etc.), or various other types of network communications. In some embodiments, the communication may occur across two or more of these types of networks and protocols.
[0019]
[0024] System 130 may be configured to receive and store data transmitted from various data sources including data source 120 via network 140, process the received data, and transmit the data and the results based on the processing to client device 110. For example, system 130 may be configured to receive unstructured data from data source 120 or other sources in network 140. In some embodiments, the unstructured data may include medical information stored in the form of one or more medical records. Each medical record may be associated with a particular patient. Data source 120 may be associated with various sources of medical information about the patient. For example, data source 120 may include the patient's healthcare providers such as physicians, nurses, specialists, consultants, hospitals, clinics, etc. Data source 120 may also be associated with laboratories such as radiology or other imaging laboratories, hematology laboratories, pathology laboratories, etc. Data source 120 may also be associated with insurance companies or any other source of patient data.
[0020]
[0025] System 130 can further communicate with one or more client devices 110 via network 140. For example, system 130 can provide results based on the analysis of information from data source 120 to client device 110. Client device 110 can include any entity or device capable of receiving or transmitting data via network 140. For example, client device 110 can include computing devices such as servers or desktop or laptop computers. Client device 110 can also include other devices such as mobile devices, tablets, wearable devices (i.e., smartwatches, embedded devices, fitness trackers, etc.), virtual machines, IoT devices, or other various technologies. In some embodiments, client device 110 can send queries about information regarding one or more patients, such as queries about patients having specific attributes or associated with specific attributes, or various other information about patients, to system 130 via network 140.
[0021]
[0026] In some embodiments, system 130 can be configured to analyze a patient's medical record (or other form of unstructured data) to identify the probability of a patient associated with a specific patient attribute. For example, system 130 can analyze a patient's medical record to determine whether the patient has undergone a test for a specific attribute and identify specific test results (tested positive, negative, etc.) or various other characteristics associated with the attribute. System 130 can be configured to identify these probabilities using one or more machine learning models. As described above, machine learning architectures have been developed for the analysis of relatively short documents. However, these techniques often do not translate well to longer documents such as a patient's medical record. For example, embodiments of long short-term memory (LSTM) models and other forms of recurrent neural networks may not be executable using conventional architectures due to the amount of text within the unstructured document.
[0022]
[0027] To overcome these and other limitations, the systems and methods of the present disclosure may parse a large data source into a series of snippet representations. These snippets may be individually processed using an LSTM-based pipeline or a similar neural network model to learn potential snippet representations. The potential snippet representations may be combined and used for classification. To improve the accuracy and efficiency of the model, it may be important to extract only relevant text from the data source for analysis. Thus, regular expressions associated with specific patient attributes may be used to extract relevant snippets from the data source. Additionally, a specific model architecture may be implemented to effectively learn from these extracted snippets. These processes are described in detail below.
[0023]
[0028] Throughout the present disclosure, patient medical records are used as exemplary embodiments, but it should be understood that in some embodiments, the disclosed systems, methods, and / or techniques may be similarly used to identify other types of individuals, objects, entities, etc. based on other forms of large unstructured data sources. Thus, the disclosed embodiments are not limited to the analysis of medical records. For example, similar techniques may be applied to legal documents, employee records, criminal or law enforcement databases, administrative (e.g., state, federal, or local) databases, transportation records (e.g., shipping records, etc.), educational institution records, public records, or various other data sources that may include vast amounts of unstructured data.
[0024]
[0029] Figure 2 shows an exemplary medical record 200 for a patient. The medical record 200 can be received from the data source 120 as described above and processed by the system 130 to identify whether the patient is associated with certain attributes. The records received from the data source 120 (or elsewhere) can include both structured data 210 and unstructured data 220, as shown in FIG. 2. The structured data 210 can include quantifiable or classifiable data about the patient, such as gender, age, race, weight, vital signs, test results, diagnosis date, diagnosis type, disease stage (e.g., billing code), treatment timing, procedures performed, clinic visit dates, type of medical treatment, insurance carrier and start date, medication instructions, medication management, or any other measurable data about the patient. The unstructured data can include information about the patient that is not quantifiable or not easily classifiable, such as a doctor's memo or a patient's laboratory report. The unstructured data 220 can include information such as an explanation of the doctor's treatment plan, a memo explaining what happened during a clinic visit, a patient's statement or explanation, a subjective assessment or explanation of the patient's health status, a radiology report, a pathology report, etc.
[0025]
[0030] In the data received from data source 120, each patient may be represented by one or more records generated by one or more healthcare professionals or patients. For example, a physician associated with the patient, a nurse associated with the patient, a physical therapist associated with the patient, etc. may each generate a medical record for the patient. In some embodiments, the one or more records may be collated and / or stored in the same database. In other embodiments, the one or more records may be distributed across multiple databases. In some embodiments, the records may store and / or provide multiple electronic data representations. For example, a patient record may be represented as one or more electronic files such as a text file, a Portable Document Format (PDF) file, an Extensible Markup Language (XML) file, etc. If the document is stored as a PDF file, an image, or other file without text, the electronic data representation may also include text associated with the document derived from an optical character recognition process. In some embodiments, unstructured data may be captured by an abstraction process, and structured data may be input by a healthcare professional or calculated using an algorithm.
[0026]
[0031] In some embodiments, the unstructured data may include data associated with specific patient attributes. As an illustrative example, the patient attributes may include the smoking status of the patient. In this example, the system 130 may analyze the patient's medical record to determine whether the patient is a smoker. For example, the unstructured data 211 may include that the patient smokes a certain number of boxes of cigarettes per week, a note (e.g., from a physician, nurse, technician, etc.) indicating that the patient uses an electronic cigarette, or a similar note. In another embodiment, the system 130 may analyze the patient's medical record to determine whether the patient has been tested for a specific biomarker such as the programmed cell death ligand 1 (PDL1) protein. For example, the unstructured data may include a note (e.g., from a physician, nurse, technician, etc.) discussing the PDL1 test results (e.g., whether the patient has been tested for PDL1, the results of the test, the analysis of the results, etc.). Throughout the present disclosure, patient identification based on PDL1 test status and / or smoking history is used, but this is an example. It should be understood that the disclosed systems, methods, and / or techniques may be used similarly for other means of identifying a patient (e.g., whether the patient has been prescribed a specific drug, whether the patient has received a specific treatment, etc.).
[0027]
[0032] As described above, system 130 can analyze unstructured medical records to extract text snippets from the unstructured data of the medical records. As used herein, a snippet can refer to a relatively small portion of text or other data contained within a larger document. A snippet can include text surrounding and including a portion of text that is relevant to a particular patient attribute. To identify snippets, system 130 may perform a keyword search to find the locations within the document where the relevant attributes are discussed. FIG. 3A shows keyword 312 as an example that can be used to search for patient attributes, consistent with the disclosed embodiments. In the example shown in FIG. 3A, system 130 can be configured to determine whether a patient associated with a set of medical records has been tested for the PDL1 protein and / or the results of that test. Accordingly, search term 310 can include the text "PDL1".
[0028]
[0033] In some embodiments, system 130 can perform a keyword search for the term "PDL1". However, in some instances, PDL1 tests and test results may be discussed using alternative notations. For example, in some instances, the term may include a dash and may be represented as "PD-L1". To avoid missing text snippets that include these alternative representations, the keyword search may be performed using a regular expression, or "regex" 312. A regular expression can include any sequence of characters that defines a search pattern. For the search for PDL1 tests, regular expression 312 may include the term "\b(pd-?l1)\b", where "-?" is a variable element for including instances with and without a dash. The term "\b" may represent a word boundary, thereby enabling system 130 to search for an exact match of the entire word of the term.
[0029]
[0034] In some embodiments, more complex regular expressions may be used. For example, regular expression 312 may include a more permissive regex such as "\b(p\W{0,2}d\W{0,2}[1lit]\W{0,2}[1lit])\b", which can account for additional characters and potential errors due to optical character recognition (OCR) from the scanned document. Regular expression 312 can be automatically generated by system 130, for example, by adding word boundary words to the search term, including variable elements at various positions associated with the search term, etc. In other embodiments, regular expression 312 may be deployed and input into system 130 by the user. It should be understood that the above-described search terms and regular expressions are provided as examples. Various other search terms, regular expressions, and / or regular expression formats may be used.
[0030]
[0035] In addition to the regular expression 312, the system 130 may search for snippets using other target words that may be associated with patient attributes. For example, if the patient attribute includes a PDL1 test, target words such as "high expression", "low expression", "tumor proportion score", "tps", "staining", and "insufficient" may generally be associated with the PDL1 test and may also be used to perform searches on unstructured documents. Similarly, if the patient attribute is the patient's smoking status, the target words may include, for example, "cigarette", "packet", "cigar", "smoker / smoke / smoked", "chew", "smoking", "ppd", "nicotine", "pipe", "tobacco", "hookah", "marijuana", "smokeless", "chewing", and "smoker". Regular expressions based on these target words similar to the regular expression 312 may also be used. Since these terms are broader, they may be used in connection with other characteristics other than the specific patient attribute being searched. For example, the term "staining" may be used in many other contents in addition to the PDL1 test. To avoid returning irrelevant snippets, additional target words may be used to extract snippets only from documents related to the patient attribute. For example, the system 130 may first perform a search using the regular expression 312 to find documents containing discussions of the PDL1 test and then extract snippets based only on the additional target words from those documents. By using these target words, it can be guaranteed that relevant snippets not including the regular expression 312 will also be identified and analyzed by the system 130.
[0031]
[0036] The above search process can be performed on each of the unstructured documents to extract snippets associated with patient attributes. FIG. 3B shows snippet 330 as an example that can be extracted from a document, consistent with the disclosed embodiments. Based on regular expression 312, system 130 can identify a document 320 that includes a target token 322 representing an instance of a search term within the text. System 130 can then extract a snippet of the text surrounding target token 322, as shown by snippet 330 in FIG. 3B. In some embodiments, snippet 330 can be defined based on a predefined window. For example, the snippet may be defined based on a predetermined number of characters before and after target token 322 within the text (e.g., 20 characters, 50 characters, 60 characters, or any suitable number of characters to capture the context for the use of the term). For example, the window may also be defined to take into account word boundaries so that no partial words are included at the end of the snippet by expanding or shrinking the window to the end at the word boundary. In some embodiments, the window may be defined based on a predefined number of words or other variables.
[0032]
[0037] In some embodiments, system 130 may replace target token 322 with a synonym 332. This can ensure that patient attributes are represented using the same specialized terms in each of the extracted snippets. For example, a document containing "PDL1" and a document containing "PD-L1" may both result in an extracted snippet containing the term "[_pdl1_]", as shown in FIG. 3B. The use of synonyms can also improve the performance of the machine learning model by reducing feature sparsity, accelerating training time, and enabling the model to converge with a more limited set of labeled data.
[0033]
[0038] Snippet 330 may then be sanitized to remove non-noun text from the snippet. Non-noun text may include, for example, HTML tags, dates, page numbers, or other data not relevant to the discussion of patient attributes. Non-noun text may be identified using a custom set of regular expression filters configured to identify common formats of non-noun text. For example, one or more regular expression filters may be designed to search for text in the MM / DD / YYYY format (or other variations) and other common date formats and remove this text from the snippet. Many punctuation characters may also be removed, although system 130 may be configured to retain any punctuation (e.g., "+", "-", etc.) that may be relevant to patient attributes. A list of potentially relevant punctuation symbols may be maintained in a database (e.g., database 132). The list may be a universal list applicable to many patient attributes or may be developed in relation to the specific attributes being examined.
[0034]
[0039] System 130 may also tokenize snippet 330 to distribute raw text to a plurality of tokens such as token 340 shown in FIG. 3B. Tokens may be distributed according to word boundaries identified within the text such that each token contains a word within the snippet. For example, starting with substitution term 332, system 130 may extract tokens “[_pdl1_]”, “high”, and “expression” from snippet 330. Tokens may be extracted throughout snippet 330 in both directions from substitution term 332. In some embodiments, a token may include a single word, as shown in FIG. 3B. In other embodiments, a token may be configured to include multiple words. For example, tokens associated with the term “BRAF negative” may be generated as “negative”, “BRAF negative”, and “BRAF”. The present disclosure is not limited to the format of tokens extracted from any particular form or snippet. In addition to tokenization, system 130 may also extract document category 350 associated with document 320. For example, document category 350 may indicate whether document 320 is a clinical note, a pathology report, or another common document type. Document category 350 may be identified within the document itself (e.g., within metadata or tags associated with document 320, the file name of document 320, etc.) or may be determined through analysis of the text of document 320 (e.g., based on document format, keywords included in the document, etc.).
[0035]
[0040] The process described above in connection with FIG. 3B may be repeated for each instance of regular expression 312 or additional target words identified within the text to extract multiple snippets from the unstructured document. Each of the snippets may be tokenized as described above. The extracted snippets may then be provided to a deep learning model architecture to identify probabilities for a patient associated with patient attributes.
[0036]
[0041] In some embodiments, two or more of the generated snippets may be identical or very similar due to text that is repeated within the unstructured data. For example, in a clinic note or other long-term patient data, text from a previous visit may be copied and pasted and thus may appear multiple times within the same record. To remove this redundancy, system 130 may remove duplicate snippets. In some instances, some but not all of the text may be duplicated within the record, and thus even if a snippet does not exactly match another snippet, it may be redundant. To account for this, system 130 may implement an overlap-based metric to measure snippet similarity. For example, a greedy algorithm may be employed. In a greedy algorithm, system 130 loops through the snippets and adds a snippet only if, based on a predefined percentage, its words are not subsumed by another snippet. The amount of subsumption may be defined as the amount of word overlap between two snippets divided by the length of the snippet being analyzed. For example, a candidate snippet may be included only if at least 80% of its words are not already included in another snippet. Various other inclusion percentages may be used.
[0037]
[0042] The model architecture may first operate in parallel on each snippet before integrating this information to generate an overall prediction about the patient. FIG. 4A is a block diagram showing a neural network as an example of operating on a single snippet, consistent with the disclosed embodiments. The snippet may include a plurality of tokens 401, 402, and 403, which may be identified through the tokenization process described above. For example, tokens 401, 402, and 403 may correspond to token 340 shown in FIG. 3B. Each of the tokens may be converted into a word embedding before passing through the neural network. For example, token 401 may be converted into word embedding 411. Word embedding 411 may be a representation of token 401 that is mapped to a real-valued vector having a predefined dimension. For example, a dimension of 128 values may be used, but word embedding 411 may have any suitable dimension. Word embedding 411 may be determined based on a training set of data. System 130 may construct a vocabulary that includes all of the tokens represented in the extracted snippets within the training data. These tokens may be indexed and projected into an embedding space. Token 411 may then be converted into word embedding 411 defined by the learned word embeddings.
[0038]
[0043] Next, the word embedding 411 can pass through a recurrent neural network such as the LSTM 420. In some embodiments, the LSTM can include a bidirectional LSTM. The LSTM may have hidden dimensions corresponding to the word embedding, which may consistently include 128 hidden dimensions as in the above examples. The LSTM 420 can be trained to generate a final hidden state that includes weighted values based on the input tokens. For example, the LSTM 420 may be trained based on a training dataset of snippet tokens having known outcomes (such as whether a patient is associated with a patient attribute). The final hidden state 421 can be generated as a result of the forward and reverse passes of the bidirectional LSTM. The same process may be executed across all snippet tokens 401 - 403, and these final hidden states can be combined to form a snippet vector 430.
[0039]
[0044] In some embodiments, the snippet vector 430 can be a concatenation of the final hidden states. For example, the snippet vector 430 can be the hidden states h 00 , h 01 , and h 02may include concatenation. Various other means may be used to combine the hidden states to form the snippet vector 430. FIG. 4B shows, in accordance with the disclosed embodiments, a process as an example for combining the hidden states of a neural network model using an attention mechanism. At each time stamp of the LSTM 420, the system 130 may take a weighted average of the hidden states. The weights may be calculated by taking the dot product of each intermediate hidden state vector 441 with the learned attention weight vector, as shown in operation 440. The attention weight vector may be learned as part of the training process for the LSTM 420. The softmax operation 450 may be used to convert the dot product output for each hidden state vector 441 to a ratio such as ratio 451. The snippet vector 430 may be determined based on all weighted combinations of the hidden states by the ratio. In particular, this process may be performed for all intermediate hidden states generated by the LSTM 420. Thus, the LSTM 420 may directly pass information from any intermediate hidden state to the complete snippet vector representation.
[0040]
[0045] In some embodiments, the initial hidden state of the LSTM 420 may be encoded with snippet metadata to improve the model. For example, the LSTM 420 may be hot encoded with the category of the snippet (e.g., as indicated by the document category 350) and the target word (e.g., PDL1, etc.) based on which the snippet was extracted. In other words, instead of initializing the LSTM with a vector of zeros before proceeding to the first (or last token), the LSTM model may be initialized with a one-hot encoding of the snippet metadata. By providing the context of the snippet in the initial state, it may be treated differently by the LSTM and may improve the results of the model.
[0041]
[0046] The processes shown in FIGS. 4A and 4B are provided as examples. It should be understood that various other suitable methods may be used to compile the resulting snippet vectors from the hidden states generated in the LSTM. Further, LSTM 420 is provided as an example. For example, LSTM 420 may be single-layer or multi-layer, and may be unidirectional or bidirectional, etc. Other forms of recurrent neural networks may also be used to generate snippet vector 430.
[0042]
[0047] The processes described above in connection with FIGS. 4A and 4B may be repeated for each snippet extracted from the unstructured data, thereby yielding a plurality of snippet vectors. It may be necessary to combine the sequence of snippet vectors into a single summary vector prior to classification in order to determine the probability associated with the patient attributes.
[0043]
[0048] FIG. 5 is a block diagram showing, in accordance with the disclosed embodiments, an example of a process for generating a summary vector 510 and a probability 530 based on a plurality of snippet vectors. One or more snippet vectors 501, 502, and 503 can be generated based on the relevant input snippets using a trained neural network as described above. The snippet vectors 501, 502, and 503 can be combined into a single summary vector 510. As an example, each of the snippet vectors 501, 502, and 503 may include 128 components (or any suitable number of components defined by the neural network model), and the summary vector 510 may similarly include 128 components. In some embodiments, the summary vector 510 can be determined based on an element-wise function performed on the snippet vectors 501, 502, and 503. For example, the summary vector 510 may be determined using an element-wise maximum operation performed across the snippet vectors such that each component of the summary vector 510 includes the maximum value of the corresponding components within the snippet vectors 501, 502, and 503. For example, the first component of the summary vector 510 may be the maximum value among the first components of the snippet vector 501, the first component of the snippet vector 502, and the first component of the snippet vector 503. Similarly, the second component of the summary vector 510 may be the maximum value among the second components of the snippet vector 501, the second component of the snippet vector 502, and the second component of the snippet vector 503. This can be repeated for each component position such that the summary vector 510 can be defined. Various other operations, including an element-wise minimum operation, an element-wise average operation, etc., can be used to define the summary vector 510.
[0044]
[0049] System 130 can be trained to project summary vector 510 onto output space 520 in the feed-forward layer. Finally, a softmax layer can be used to generate predicted probabilities 530 for each output class. The predicted probabilities 530 can be converted to predicted class labels. The number and type of probabilities determined using summary vector 510 can depend on the type of patient attributes being analyzed. For example, if the PDL1 status is used as a patient attribute, the probabilities can include the probability that the patient tests positive for PDL1, the probability that the patient tests negative for PDL1, the probability that the patient has not been tested, and the probability that the result is inconclusive. Similarly, if the patient attribute is the patient's smoking status, the probabilities can include the probability that the patient has a smoking history, the probability that the patient has no smoking history, and the probability that the result is inconclusive. Depending on the type of patient attributes being analyzed, various other probabilities can be included. Each probability can be represented in various formats. For example, the probability can be represented as a percentage, a pre-defined scale (e.g., 1-10, 1-5, etc.), a list of pre-defined classifications (e.g., "high probability", "low probability", etc.), or any other suitable form.
[0045]
[0050] The resulting probabilities can indicate whether the patient is associated with the patient attribute. For example, the probabilities can indicate whether the patient has been tested for PDL1 and the result of that test, along with the associated confidence level. Accordingly, system 130 can be used to classify patients based on unstructured medical data in the patient's medical record. Since only relevant snippets of each document are analyzed, system 130 can advantageously use an LSTM model to determine probabilities associated with patient attributes despite relatively large documents commonly included in the patient's medical record.
[0046]
[0051] FIG. 6 is a flowchart showing a process 600 as an example for determining probabilities associated with attributes, consistent with the disclosed embodiments. The process 600 can be executed by at least one processing device, such as the processing engine 131, as described above. Throughout the present disclosure, it should be understood that the term "processor" is used as an abbreviation for "at least one processor". In other words, a processor can include one or more structures that perform logical operations, regardless of whether such structures are co-located, connected, or distributed. In some embodiments, the non-transitory computer-readable medium can include instructions that cause the processor to execute process 600 when executed by the processor. Further, process 600 is not necessarily limited to the steps shown in FIG. 6, and any steps or processes of the various embodiments described throughout the present disclosure, including those described above with respect to FIGS. 3A-5, can also be included in process 600.
[0047]
[0052] In step 610, process 600 can include accessing a database storing at least one unstructured medical record. For example, system 130 can access a patient's medical record from an external data source such as local database 132 or data source 120. The medical record can include one or more electronic files such as text files, image files, PDF files, XLM files, YAML files, etc. The at least one unstructured medical record can correspond to the medical record 210 described above. For example, the unstructured medical record can include at least some unstructured data 211. The unstructured information can include text written by a healthcare provider, radiology reports, pathology reports, or various other forms of text associated with the patient. In some embodiments, the medical record can further include additional structured data 212.
[0048]
[0053] In step 620, process 600 may include analyzing at least one unstructured medical record to identify multiple snippets of information within at least one unstructured medical record associated with patient attributes. In some embodiments, identifying the snippets may include searching at least one unstructured medical record for keywords associated with the patient attributes. For example, the patient attributes may include whether the patient has been tested for PDL1, and the keyword may include the text "PDL1". In some embodiments, the keyword may include at least one variable element. For example, the keyword may be represented using a regular expression such as regular expression 312. Accordingly, the keyword may take into account alternative spellings of the patient attributes, additional or unwanted characters that appear in the text, errors resulting from OCR processing of the scanned document, word boundaries, and other variables that may affect snippet extraction.
[0049]
[0054] In some embodiments, additional snippets may be identified based on target words related to the keyword. For example, when the patient attribute is the patient's smoking history, the target word may include "cigarette", "pack", "vaping", or other terms related to smoking. To avoid identifying irrelevant snippets, the snippets based on these target words may be extracted only from the documents that contain the keyword in the initial search. Further, in some embodiments, when the number of words (or percentage of words) included by another snippet exceeds a predetermined threshold, one or more redundant snippets may be removed. Although step 620 is described based on a single snippet, it should be understood that the same process may be performed on multiple snippets extracted from the unstructured medical record.
[0050]
[0055] In step 630, process 600 may include generating a snippet vector that includes a plurality of snippet vector components based on each of the plurality of snippets. The plurality of snippet vector components may include weighted values associated with at least one word included in the snippet. In some embodiments, the snippet vector may be generated using a neural network such as a long short-term memory network, or other forms of recurrent neural networks. For example, a snippet including tokens 401, 402, and 403 may pass through LSTM 420 to generate a snippet vector 430. Accordingly, step 630 may include combining a plurality of hidden states to form a snippet vector 430. This may include concatenation, an attention mechanism, or various other means for generating a snippet vector, as described above with respect to FIGS. 4A and 4B.
[0051]
[0056] In step 640, process 600 may include analyzing the snippet vector to generate a summary vector that includes a plurality of summary vector components. For example, snippet vectors 501, 502, and 503 may be combined into a single snippet vector 510. Each of the plurality of summary vector components may be associated with a corresponding snippet vector component. For example, the snippet vector and the summary vector may each include the same number of components such that there is a corresponding component in the snippet vector for each component in the summary vector. Further, each of the plurality of summary vector components may be determined based on an analysis of the corresponding snippet vector components. For example, each summary vector component may include the maximum value of the corresponding snippet vector components in the plurality of snippet vectors (e.g., using an element-wise maximum operation), as described above.
[0052]
[0057] In step 650, process 600 may include generating at least one output indicating a probability associated with an attribute based on the summary vector. In some embodiments, the probability may include the probability of whether an examination is being performed on a patient associated with the patient attribute. For example, the probability may include the probability of whether a patient is being tested for PDL1. Additionally, or alternatively, the probability may include the probability of whether a patient has been tested positive (or negative) for the patient attribute. For example, the probability may include the probability of whether a patient has been tested positive for PDL1. In other embodiments, the probability may include the probability that a patient exhibits a particular health-related characteristic. For example, the probability may include the probability of whether a patient has a smoking history. In some embodiments, the output may include an indication that the association between the patient and the patient attribute is uncertain. For example, the output may include the probability that the correlation between the patient and the patient attribute cannot be determined based on the unstructured medical record.
[0053]
[0058] The foregoing description has been presented for purposes of illustration. It is not exhaustive and is not limited to the exact forms disclosed. Modifications and adaptations will be apparent to those skilled in the art from the consideration of the specification and the practice of the disclosed embodiments. Additionally, although aspects of the disclosed embodiments are described as being stored in memory, those skilled in the art will understand that these aspects can also be stored on other types of computer-readable media, such as secondary storage devices, e.g., hard disks or CD ROMs, or other forms of RAM or ROM, USB media, DVDs, Blu-rays, 4K Ultra HD Blu-rays, or other optical drive media.
[0054]
[0059] Computer programs based on the written description and disclosed methods are within the scope of the skills of experienced developers. Various programs or program modules can be generated using any of the techniques known to those skilled in the art or designed in relation to existing software. For example, program sections or program modules may be designed in or using.Net Framework,.Net Compact Framework (and related languages such as Visual Basic, C), Java, Python, R, C++, Objective-C, HTML, combinations of HTML / AJAX, XML, or Java applets, or a combination thereof.
[0055]
[0060] Furthermore, while exemplary embodiments are described herein, the scope of any and all embodiments having related elements, modifications, omissions, combinations (such as aspects across various embodiments), adaptations, and / or alterations is to be understood by those skilled in the art based on this disclosure. Limitations within the claims are to be broadly construed based on the language employed in the claims and should not be limited to the examples described herein or during the prosecution of this application. Examples are to be construed as non-exclusive. Additionally, the steps of the disclosed methods can be modified in any way including changing the order of the steps and / or inserting or deleting steps. Accordingly, this specification and the examples are intended to be considered only as examples, using the true scope and spirit indicated by the following claims and their equivalents in their entirety.
Claims
1. A model assistance system for determining probabilities associated with patient attributes, comprising: at least one processor, accessing a database storing at least one unstructured medical record associated with a patient, analyzing the at least one unstructured medical record to identify a plurality of snippets of information within the at least one unstructured medical record, the plurality of snippets including text indicative of the patient attributes, determining an overlap metric indicative of an amount of text repeated within a first snippet and a second snippet of the plurality of snippets, when the overlap metric is greater than a threshold value, removing one of the first snippet or the second snippet from the plurality of snippets, after removing one of the first snippet or the second snippet from the plurality of snippets, generating a snippet vector including a plurality of snippet vector components based on each snippet of the plurality of snippets, the plurality of snippet vector components including weighted values associated with at least one word included in the snippet, analyzing the snippet vector to generate a summary vector including a plurality of summary vector components, each of the plurality of summary vector components being associated with a corresponding snippet vector component and determined based on an analysis of the corresponding snippet vector component, generating at least one output indicative of a probability associated with the patient attributes based on the summary vector, The model assistance system comprising the at least one processor programmed as such.
2. The model assistance system according to claim 1, wherein identifying the plurality of snippets includes searching the at least one unstructured medical record for keywords associated with the patient attributes.
3. The model assistance system according to claim 2, wherein the keyword includes at least one variable element.
4. The model assistance system according to claim 3, wherein the variable element is represented as a regular expression.
5. The model assistance system according to claim 1, wherein the snippet vector is generated using a neural network.
6. The model assistance system according to claim 5, wherein the neural network includes a long short-term memory network.
7. The model support system according to claim 1, wherein each summary vector component includes a maximum value of corresponding snippet vector components within a plurality of snippet vectors.
8. The model support system according to claim 1, wherein the probability includes a probability of whether the patient has been tested for PDL1.
9. The model support system according to claim 1, wherein the probability includes a probability of whether the patient has been tested positive for PDL1.
10. The model support system according to claim 1, wherein the probability includes a probability of whether the patient has a smoking history.
11. The model support system according to claim 1, wherein the output includes a label indicating that the association between the patient and the patient attribute is uncertain.
12. The probability is whether an examination has been performed on the patient associated with the patient attribute The model support system according to claim 1, including the probability of.
13. The model support system according to claim 1, wherein the probability includes a probability of whether the patient has been tested positive for the patient attribute.
14. A computer-assisted method for determining a probability associated with a patient attribute, comprising: Accessing a database storing at least one unstructured medical record associated with a patient; Analyzing the at least one unstructured medical record to identify a plurality of snippets of information within the at least one unstructured medical record, wherein the plurality of snippets includes text indicating the patient attribute; Determining an overlap metric indicating an amount of text repeated within a first snippet and a second snippet of the plurality of snippets; When the overlap metric is greater than a threshold value, removing one of the first snippet or the second snippet from the plurality of snippets; After removing one of the first snippet or the second snippet from the plurality of snippets, generating a snippet vector including a plurality of snippet vector components based on each snippet of the plurality of snippets, wherein the plurality of snippet vector components include weighted values associated with at least one word included in the snippet. Analyzing the snippet vector to generate a summary vector including a plurality of summary vector components, wherein each of the plurality of summary vector components is associated with a corresponding snippet vector component and is determined based on the analysis of the corresponding snippet vector component; Generating at least one output indicating a probability associated with the patient attribute based on the summary vector; A computer-assisted method comprising the above.
15. The computer-assisted method according to claim 14, wherein identifying the plurality of snippets includes searching the at least one unstructured medical record for keywords associated with the patient attribute.
16. The computer-assisted method according to claim 15, wherein the keyword includes at least one variable element.
17. The computer-assisted method according to claim 16, wherein the variable element is represented as a regular expression.
18. The computer-assisted method according to claim 14, wherein the snippet vector is generated using a neural network.
19. The computer-assisted method according to claim 18, wherein the neural network includes a long short-term memory network.
20. The computer-assisted method according to claim 14, wherein each summary vector component includes the maximum value of the corresponding snippet vector component within a plurality of snippet vectors.
21. The computer-assisted method according to claim 14, wherein the probability includes the probability that the patient has been tested for PDL1.
22. The computer-assisted method according to claim 14, wherein the probability includes the probability that the patient has been tested positive for PDL1.
23. The computer-assisted method according to claim 14, wherein the probability includes the probability that the patient has a smoking history.
24. The computer-assisted method according to claim 14, wherein the output includes a label indicating that the association between the patient and the patient attribute is uncertain.
25. The model-assisted system according to claim 14, wherein the probability includes the probability that the patient has been tested for the patient attribute associated with the patient.
26. The model-assisted system according to claim 14, wherein the probability includes the probability that the patient has been tested positive for the patient attribute.
Citation Information
Patent Citations
Document retrieval method and apparatus, and computer program therefor
JP2008287394A
Method of monitoring electronic media
US20090119275A1
Systems and methods for model-assisted cohort selection
US20180300640A1