System and method for state identification and classification of text data

A method for processing non-standardized pet insurance claims using machine learning models enhances claim classification efficiency and accuracy by converting text data into modelable features and aggregating states, addressing the limitations of current text analysis systems.

JP2026035571APending Publication Date: 2026-03-04TRUPANION INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025176332
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-05-13
Filing Date
2025-10-20
Publication Date
2026-03-04

AI Technical Summary

Technical Problem

Current automatic text analysis models struggle with non-standardized text data containing unfamiliar words or phrases, particularly in pet insurance claims processing, leading to inefficiencies and delays due to the need for technical expertise.

Method used

A computer-implemented method that extracts text data, converts it into transformed features using machine learning models, processes these features to identify multiple states, and aggregates them to generate an output indicating the status of the event, enabling efficient and accurate processing of insurance claims.

Benefits of technology

The method enables rapid and precise classification of insurance claims by handling non-standardized text data, improving processing speed and accuracy in pet insurance claim automation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026035571000001_ABST
    Figure 2026035571000001_ABST
Patent Text Reader

Abstract

A system and method for state identification and classification of text data is provided. The present disclosure provides a system and method for identifying one or more states of a text string describing an event and classifying the event based on the one or more identified states. The disclosed method includes receiving the text string describing the event, converting the text string into modelable data, analyzing word structures in the converted data to identify one or more states of the event, and classifying the event based on the identified states.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 024299, filed May 13, 2020, which is incorporated herein by reference in its entirety. [Background technology]

[0002] background

[0002] Automatic text analysis is important for extracting information from text data. However, current automatic text analysis models are limited in that they cannot handle text data that is in unexpected formats or contains unfamiliar words or phrases. This can be particularly problematic in processing claims in pet insurance programs. For example, adjusters may need to review veterinary records that contain non-standard pet hygiene protocols, which can require technical knowledge and expertise in animal science and slow down the process of processing insurance claims. Summary of the Invention [Problem to be solved by the invention]

[0003] overview

[0003] There is a need for a system and method for processing text data that is not in a standardized format or that contains non-standard language or phrases. Additionally, there is recognized herein a need for a system and method for automating claims processing in the pet insurance industry. The systems and methods provided herein can efficiently process insurance claims with improved speed and accuracy. [Means for solving the problem]

[0004] In an aspect of the present disclosure, a computer-implemented method for classifying events is provided, the method including: (a) extracting text data from input data, the text data describing an event; (b) converting the text data into transformed input features for processing by a number of machine learning algorithm-trained models; (c) processing the transformed input features using a number of machine learning algorithm-trained models to output a number of states of the event; and (d) aggregating the multiple states to generate an output indicating a status of the event.

[0005] In a related but separate aspect, a non-transitory computer-readable medium is provided that includes instructions that, when executed by a processor, cause the processor to perform a method for classifying events. The method includes (a) extracting text data from input data, the text data describing an event, (b) converting the text data into transformed input features for processing by multiple machine learning algorithm-trained models, (c) processing the transformed input features using multiple machine learning algorithm-trained models to output multiple states of the event, and (d) aggregating the multiple states to generate an output indicating a status of the event.

[0006] In some embodiments, the input data includes unstructured text data. In some embodiments, extracting the text data includes identifying word combinations from the input data. In some cases, the method further includes determining boundaries for locations of anchor words based at least in part on locations of the anchor words. In some examples, the method further includes recognizing subsets of the text data within the boundaries. For example, the method further includes grouping at least a portion of the subsets of the text data based on coordinates of the subsets of the text data. In some cases, the anchor words are predetermined based on a format of the input data. In some cases, the anchor words are identified by using a machine learning algorithm trained model to predict the presence of line item words.

[0007] In some embodiments, extracting the text data includes (i) identifying a word that is outside a data distribution range of the multiple machine learning algorithm trained models, and (ii) replacing the word with a replacement word that is within the data distribution range of the multiple machine learning algorithm trained models. In some embodiments, the transformed input features include numeric values.

[0008] In some embodiments, the multiple conditions are different types of conditions. In some embodiments, the multiple conditions include medical conditions, medical procedures, dental treatments, preventative treatments, diets, health checks, medications, treatment sites, costs, discounts, chronic conditions, illnesses, or diseases. In some embodiments, the multiple conditions are aggregated using a trained model. In some cases, the output includes a probability of the status.

[0009] In some embodiments, the output includes insights inferred from aggregating multiple states. In some embodiments, the status of the event includes approval, rejection, or a request for further validation action. In some embodiments, the method further includes providing two different machine learning algorithm-trained models corresponding to the same state. In some cases, the method further includes selecting a model from the two different machine learning algorithm-trained models to process the transformed input features based on the features of the event. In some embodiments, the input data includes transcribed data.

[0010]

[0010] In an aspect of the present disclosure, there is provided a computer-implemented method for classifying an event, the method including receiving a transformed text string describing the event, identifying words present in the transformed text string, identifying combinations of words present in the transformed text string, and classifying the event based on (i) the words, (ii) the combinations of words, or (iii) the combinations thereof.

[0011] In some embodiments, the classifying includes identifying a state of the event. In some cases, the state is selected from at least 100, at least 500, at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, or at least 10,000 possible states. In some embodiments, the classifying includes identifying two or more states. In some cases, the two or more states are determined from two or more processes. In some examples, the two or more processes run in parallel.

[0012] In some embodiments, identifying the words includes identifying the words from a database of words identified in past text strings. In some instances, the database of words includes at least 100, at least 500, at least 1000, at least 5000, at least 10,000, at least 20,000, or at least 30,000 known words. In some embodiments, identifying the words includes assigning numeric identifiers to the words. In some instances, the numeric identifiers correspond to words identified in past text strings. In some instances, the numeric identifiers do not correspond to words identified in past text strings. In some embodiments, identifying word combinations includes identifying significant word combinations. In some instances, significant word combinations are identified from a database of significant word combinations. In some instances, the database of significant word combinations includes word combinations identified from past text strings as indicative of a condition. In some cases, the database of significant word combinations includes at least 100, at least 500, at least 1000, at least 5000, or at least 10,000 significant word combinations.

[0013] In some embodiments, the condition is a medical condition. In some embodiments, the condition is a medical procedure. In some embodiments, the condition is a dental treatment, a preventative treatment, a diet, a medical checkup, a medication, a treatment site, a cost, a discount, a chronic condition, a disease, or an illness.

[0014] In some cases, the classifying includes identifying multiple states. In some cases, the states of the multiple states are identified independently. In some cases, the classifying further includes aggregating the multiple states to determine a result. In some cases, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, or at least 17 states are identified.

[0015] In some embodiments, the state is a standardized state. In some embodiments, the converted text data includes data that has been converted from non-standardized text data. In some embodiments, classifying includes applying a trained machine learning model to determine the possible states. In some cases, the trained machine learning model includes a neural network. In some examples, identifying the word includes activating input neurons. In some cases, the trained machine learning model is trained using a training set that includes past text strings.

[0016]

[0016] Another aspect of the present disclosure provides a non-transitory computer-readable medium containing machine-executable code that, when executed by one or more computer processors, performs any of the methods described above or elsewhere in this specification.

[0017] Another aspect of the present disclosure provides a system including one or more computer processors and a computer memory coupled to the one or more computer processors, the computer memory including machine-executable code that, when executed by the one or more computer processors, performs any of the methods described above or elsewhere herein.

[0018]

[0018] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description. In the following detailed description, only exemplary embodiments of the present disclosure are shown and described. As will be appreciated, the present disclosure is capable of other different embodiments and its several details can be modified in various obvious respects, all of which can be done without departing from the present disclosure. Accordingly, the drawings and descriptions are to be regarded as illustrative in nature, and not restrictive. The present invention provides, for example, the following. (Item 1) 1. A computer-implemented method for classifying events, comprising: (a) extracting text data from input data, said text data describing said events; (b) converting the text data into transformed input features for processing by a number of machine learning algorithm trained models; (c) processing the transformed input features using the multiple machine learning algorithm trained models to output multiple states of the event; and (d) aggregating said multiple states to generate an output indicative of the status of said event; 11. A computer-implemented method comprising: (Item 2) Item 10. The computer-implemented method of item 1, wherein the input data comprises unstructured text data or transcribed data. (Item 3) Item 10. The computer-implemented method of item 1, wherein extracting the text data includes identifying word combinations from the input data. (Item 4) Item 10. The computer-implemented method of item 1, wherein extracting the text data includes identifying anchor words from the input data. (Item 5) Item 5. The computer-implemented method of item 4, further comprising determining a boundary for the location of the anchor word based at least in part on the location of the anchor word. (Item 6) Item 6. The computer-implemented method of item 5, further comprising recognizing a subset of the text data within the boundary. (Item 7) 7. The computer-implemented method of claim 6, further comprising grouping at least a portion of the subset of the text data based on coordinates of the subset of the text data. (Item 8) Item 5. The computer-implemented method of item 4, wherein the anchor words are predetermined based on the format of the input data. (Item 9) Item 5. The computer-implemented method of item 4, wherein the anchor words are identified by using a machine learning algorithm trained model to predict the presence of line item words. (Item 10) Item 1. The computer-implemented method of item 1, wherein extracting the text data includes (i) identifying words outside a data distribution range of the multiple machine learning algorithm trained models; and (ii) replacing the words with replacement words within the data distribution range of the multiple machine learning algorithm trained models. (Item 11) Item 10. The computer-implemented method of item 1, wherein the transformed input features include numerical values. (Item 12) Item 10. The computer-implemented method of item 1, wherein the multiple states are states of different types. (Item 13) Item 10. The computer-implemented method of item 1, wherein the multiple conditions include a medical condition, a medical procedure, a dental treatment, a preventative treatment, a diet, a health check, a medication, a treatment area, a cost, a discount, a chronic condition, a disease, or an illness. (Item 14) Item 10. The computer-implemented method of item 1, wherein the multiple states are aggregated using a trained model. (Item 15) Item 15. The computer-implemented method of item 14, wherein the output comprises a probability of the status. (Item 16) Item 10. The computer-implemented method of item 1, wherein the output comprises insights inferred from aggregating the multiple states. (Item 17) Item 10. The computer-implemented method of item 1, wherein the status of the event includes approval, rejection, or a request for further validation action. (Item 18) Item 10. The computer-implemented method of item 1, further comprising providing two different machine learning algorithm-trained models corresponding to the same state. (Item 19) 20. The computer-implemented method of claim 18, further comprising selecting a model from the two different machine learning algorithm trained models to process the transformed input features based on the event features. (Item 20) 20. The computer-implemented method of claim 19, wherein the characteristics of the event include a latency to classify the event. (Item 21) 1. A non-transitory computer-readable medium containing instructions that, when executed by a processor, cause the processor to perform a method for classifying events, the method comprising: (a) extracting text data from input data, said text data describing said events; (b) converting the text data into transformed input features for processing by a number of machine learning algorithm trained models; (c) processing the transformed input features using the multiple machine learning algorithm trained models to output multiple states of the event; and (d) aggregating said multiple states to generate an output indicative of the status of said event; 1. A non-transitory computer-readable medium, comprising: (Item 22) Item 22. The non-transitory computer-readable medium of item 21, wherein the input data comprises unstructured text data. (Item 23) 22. The non-transitory computer-readable medium of claim 21, wherein extracting the text data includes identifying word combinations from the input data. (Item 24) 22. The non-transitory computer-readable medium of claim 21, wherein extracting the text data includes identifying anchor words from the input data. (Item 25) 25. The non-transitory computer-readable medium of claim 24, wherein the method further comprises determining a boundary for a location of the anchor word based at least in part on the location of the anchor word. (Item 26) 26. The non-transitory computer-readable medium of claim 25, wherein the method further comprises recognizing a subset of the text data within the boundary. (Item 27) 27. The non-transitory computer-readable medium of claim 26, wherein the method further comprises grouping at least a portion of the subsets of the text data based on coordinates of the subsets of the text data. (Item 28) Item 25. The non-transitory computer-readable medium of item 24, wherein the anchor words are predetermined based on a format of the input data. (Item 29) Item 25. The non-transitory computer-readable medium of item 24, wherein the anchor words are identified by using a machine learning algorithm trained model to predict the presence of line item words. (Item 30) Item 22. The non-transitory computer-readable medium of item 21, wherein extracting the text data includes (i) identifying words outside a data distribution range of the multiple machine learning algorithm trained models; and (ii) replacing the words with replacement words within the data distribution range of the multiple machine learning algorithm trained models. (Item 31) Item 22. The non-transitory computer-readable medium of item 21, wherein the transformed input features include numerical values. (Item 32) 22. The non-transitory computer-readable medium of claim 21, wherein the multiple states are different types of states. (Item 33) 22. The non-transitory computer-readable medium of item 21, wherein the multiple conditions include a medical condition, a medical procedure, a dental treatment, a preventative treatment, a diet, a health check, a medication, a treatment area, a cost, a discount, a chronic condition, a disease, or an illness. (Item 34) Item 22. The non-transitory computer-readable medium of item 21, wherein the multiple states are aggregated using a trained model. (Item 35) Item 35. The non-transitory computer-readable medium of item 34, wherein the output includes a probability of the status. (Item 36) 22. The non-transitory computer-readable medium of claim 21, wherein the output comprises insights inferred from aggregating the multiple states. (Item 37) 22. The non-transitory computer-readable medium of claim 21, wherein the status of the event includes approval, rejection, or a request for further validation action. (Item 38) 22. The non-transitory computer-readable medium of claim 21, wherein the method further comprises providing two different machine learning algorithm trained models corresponding to the same state. (Item 39) Item 39. The non-transitory computer-readable medium of Item 38, wherein the method further comprises selecting a model from the two different machine learning algorithm trained models to process the transformed input features based on features of the event. (Item 40) 22. The non-transitory computer-readable medium of claim 21, wherein the input data comprises transcribed data. (Item 41) 1. A computer-implemented method for classifying events, comprising: a. receiving a converted text string describing the event; b. identifying words present in said transformed text string; c. identifying word combinations present in said transformed text string; d. classifying said events based on (i) said words, (ii) combinations of said words, or (iii) combinations thereof; 11. A computer-implemented method comprising: (Item 42) Item 42. The computer-implemented method of item 41, wherein the classifying includes identifying a state of the event. (Item 43) Item 43. The computer-implemented method of item 42, wherein the states are selected from at least 100, at least 500, at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, 10,000 possible states or at least 10,000 possible states. (Item 44) Item 42. The computer-implemented method of item 41, wherein the classifying comprises distinguishing between two or more states. (Item 45) Item 45. The computer-implemented method of item 44, wherein the two or more states are determined from two or more processes. (Item 46) Item 46. The computer-implemented method of item 45, wherein the two or more processes run in parallel. (Item 47) Item 42. The computer-implemented method of item 41, wherein identifying the word includes identifying the word from a database of words identified in past text strings. (Item 48) Item 48. The computer-implemented method of item 47, wherein the database of words comprises at least 100, at least 500, at least 1000, at least 5000, at least 10,000, at least 20,000, or at least 30,000 known words. (Item 49) Item 42. The computer-implemented method of item 41, wherein identifying the word includes assigning a numeric identifier to the word. (Item 50) Item 50. The computer-implemented method of item 49, wherein the numeric identifier corresponds to a word identified in a past text string. (Item 51) Item 50. The computer-implemented method of item 49, wherein the numeric identifier does not correspond to a word identified in a past text string. (Item 52) Item 42. The computer-implemented method of item 41, wherein identifying word combinations includes identifying significant word combinations. (Item 53) Item 53. The computer-implemented method of item 52, wherein the significant word combinations are identified from a database of significant word combinations. (Item 54) Item 54. The computer-implemented method of item 53, wherein the database of significant word combinations includes word combinations identified from past text strings as indicative of conditions. (Item 55) Item 54. The computer-implemented method of item 53, wherein the database of significant word combinations comprises at least 100, at least 500, at least 1000, at least 5000, or at least 10,000 significant word combinations. (Item 56) Item 53. The computer-implemented method of item 52, wherein the condition is a medical condition. (Item 57) Item 53. The computer-implemented method of item 52, wherein the condition is a medical procedure. (Item 58) Item 53. The computer-implemented method of item 52, wherein the condition is dental treatment, preventative treatment, diet, medical checkup, medication, treatment area, cost, discount, chronic condition, disease or illness. (Item 59) Item 52. The computer-implemented method of item 51, wherein the classifying includes identifying multiple states. (Item 60) Item 60. The computer-implemented method of item 59, wherein the states of the multiple states are identified independently. (Item 61) 60. The computer-implemented method of claim 59, wherein the classifying further comprises aggregating the multiple states to determine an outcome. (Item 62) Item 60. The computer-implemented method of item 59, wherein at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, or at least 17 states are identified. (Item 63) Item 53. The computer-implemented method of item 52, wherein the state is a standardized state. (Item 64) Item 42. The computer-implemented method of item 41, wherein the converted text data comprises data that has been converted from non-standardized text data. (Item 65) Item 42. The computer-implemented method of item 41, wherein the classifying comprises applying a trained machine learning model to determine the likely states. (Item 66) Item 66. The computer-implemented method of item 65, wherein the trained machine learning model comprises a neural network. (Item 67) 67. The computer-implemented method of claim 66, wherein identifying the word comprises activating an input neuron. (Item 68) Item 66. The computer-implemented method of item 65, wherein the trained machine learning model is trained using a training set that includes past text strings. (Item 69) 69. A non-transitory computer-readable medium containing instructions that, when executed by a processor, cause the processor to perform the method of any one of items 41 to 68.

[0019] Incorporation by Reference

[0019] All publications, patents, and patent applications mentioned in this specification are incorporated by reference into this specification to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference.

[0020] BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The novel features of the present disclosure are set forth with particularity in the appended claims. The features and advantages of the present disclosure will be better understood by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the disclosure are utilized, and the accompanying drawings. [Brief explanation of the drawings]

[0021] [Figure 1]

[0021] A method for transforming and categorizing event description text data according to one or more embodiments of the present disclosure is depicted. [Figure 2]

[0022] 1 depicts a method for classifying event description text data based on word structure, according to one or more embodiments of the present disclosure. [Figure 3]

[0023] 1 illustrates a neural network for classifying event description text data, in accordance with one or more embodiments of the present disclosure. [Figure 4]

[0024] 1 depicts a method for classifying event description text data based on word structure using a trained neural network, according to one or more embodiments herein. [Figure 5]

[0025] 1 illustrates a system for identifying and classifying one or more conditions in accordance with one or more embodiments of the present disclosure. [Figure 6]

[0026] 1 illustrates a system for identifying word structures for training and using neural networks, in accordance with one or more embodiments of the present disclosure. [Figure 7]

[0027] 1 depicts a method of operation of a system for identifying and classifying one or more conditions, according to one or more embodiments herein. [Figure 8]

[0028] 1 illustrates a schematic of an insurance claims processing system according to some embodiments of the present invention. [Figure 8A]

[0029] 1 illustrates a schematic diagram of another example of an insurance claims processing system, according to some embodiments of the present invention. [Figure 8B]

[0030] 1 shows an example of an image processed by an OCR algorithm. [Figure 8C]

[0031] 1 shows an example of anchors identified from image input. [Figure 8D]

[0032] Here is an example of separated line item text grouped by line number. [Figure 9]

[0033] 1 illustrates a workflow of a method for determining a probable outcome based on a plurality of states identified in a plurality of processes. [Figure 10]

[0034] 1 illustrates generally a platform on which methods and systems for automated insurance claims processing may be implemented. [Figure 11]

[0035] 1 illustrates a schematic representation of a predictive model creation and management system according to some embodiments of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0022] Detailed Description

[0036] The present disclosure provides systems and methods for processing and classifying text data related to descriptions of events. Specifically, the present disclosure provides systems and methods for automating pet insurance claim processing. As described herein, the systems and methods of the present disclosure can process text data, such as insurance claims or pet insurance claims in non-standardized formats, by converting the text data into modelable data, identifying one or more states of the events described in the text, and classifying the events based on the one or more states.

[0023]

[0037] In some embodiments of the present disclosure, the text data may include claims data obtained from a claims database and / or various bills and documents associated with pet insurance claims. Raw input data may be associated with an insurance claim, such as structured claims data obtained from a claims data store or insurance system. For example, the structured claims data may be submitted by a veterinary clinic or pet owner in a customized claim form. In some cases, the structured claims data may include text data such as insurance policy ID / number, illness / injury, or other fields for the pet or treatment. In some cases, the text data may include structured data such as JavaScript Object Notation (JSON) data. In optional cases, the raw input data may include unstructured data related to a claim, such as an insurance claim, an invoice image, a medical report, email, or web-based content. The text data may be received as an online form submission, email text, a word processing document, a portable document format (PDF), an image of text, or various other formats. Unstructured input data, such as an email or invoice image, may be preprocessed to extract text data before processing.

[0024]

[0038] As explained above, due to non-standard pet hygiene codes or the lack of other uniform standards or regulations, the text data may be in a variety of non-standardized formats. The non-standardized text data may describe an event in prose without adhering to standardized terminology, phrasing, or formatting. The non-standardized text data may include a description of an event that does not conform to a standard description of an event. In some embodiments, the event description is prepared by a user or by a member of the general public. In some embodiments, the event description may be prepared by an observer of the event. In some embodiments, the event description is prepared by a skilled medical practitioner. For example, the event description may include a description of a medical procedure performed on a subject. A medical event description may be prepared by a medical professional and provided to the system of this disclosure.

[0025]

[0039] In some cases, the patient may also be referred to as a pet. As used herein, the term "veterinary clinic" may refer to a hospital, clinic, or the like where services are provided to animals.

[0026]

[0040] As used herein, "medicine" may include human medicine, veterinary medicine, dental medicine, natural medicine, alternative medicine, or the like. A subject may be a human subject or an animal subject. "Medical professional," as used herein, may include a physician, veterinarian, medical technician, veterinary nurse, medical researcher, veterinary researcher, naturopath, homeopath, therapist, or the like. A medical procedure may include a medical procedure, veterinary procedure, dental procedure, naturopathic procedure, or the like performed on a human. A medical event may include a medical event, veterinary event, dental event, naturopathic event, or the like involving a human subject. In some cases, a description of a medical event may include one or more line items corresponding to, for example, a procedure, a product, a reagent, a result, a condition, or a diagnosis.

[0027]

[0041] FIG. 1 illustrates a workflow of a method 100 described herein. The method includes receiving 110 a text string describing an event. The text string may be received, for example, through an online submission form or email, or may be obtained in a variety of electronic ways, including from a PDF, a word processing document, an image of the text, or screen scraping. The text string may be in a non-standardized format. The text string may be converted 120 into modelable data. Converting the text string to modelable data may include converting the text string to numerical data. For example, the text string may be converted into a series of numerical identifiers, which correspond to and identify words. In some embodiments, converting the text string to modelable data may further include removing common words (such as pronouns, prepositions, articles, or conjunctions) from the text string. Word structures in the converted data may be analyzed 130 to determine one or more states indicated by the word structures. Analyzing the word structures may include determining the presence or absence of words in the text string. In some embodiments, determining the presence or absence of a word in the text string may include determining whether a numeric identifier corresponding to the word is present in the transformed data, and determining that the word is present in the text string if the numeric identifier corresponding to the word is present in the transformed data.

[0028]

[0042] Analyzing the word structure may further include identifying word combinations present in the text string. In some embodiments, the word combinations may include two or more words that indicate a condition. One or more conditions can be identified based on word structures (e.g., words or word combinations) present in the text string. In some embodiments, the conditions may correspond to elements of an event description, such as a line item. For example, a condition may correspond to a procedure, product, reagent, result, health condition, or diagnosis selected from a finite number of possible conditions. The event described in the text string or the one or more conditions identified in 130 can be classified (140) based on the one or more identified conditions. In some embodiments, the classification may be based on previous classifications of the condition. The condition may be a standardized condition (e.g., a medical billing code associated with a health condition or treatment).

[0029]

[0043] An exemplary implementation of method 100 described with respect to FIG. 1 may be for identifying and classifying text strings describing medical events. In some embodiments, the text strings describing medical events may be descriptions of treatments, conditions, or diagnoses prepared by a medical professional. The descriptions of treatments, results, conditions, or diagnoses may further include products or reagents used in the medical event. The descriptions need not be in a standardized format or use standardized terminology. For example, a test measuring kidney function may be interchangeably described as a "kidney function panel," "renal function test," or "renal panel." As shown in step 110, text strings describing medical events may be submitted to the system of the present disclosure by a medical professional, a patient, a customer, a pet owner, or any other individual. As shown in step 120, the text strings describing medical events may be converted into modelable data including a numeric identifier identifying each word present in the text string. As shown in 130, the word structure of the text string may be analyzed to determine one or more conditions of the medical event. For example, a word configuration including the word "kidney" or "renal" combined with the word "test" or "panel" may identify a test measuring kidney function as a condition of the medical event. In some embodiments, the condition may be associated with a medical billing code, such as a Physician Profession Terminology (CPT) code. As shown at 140, the medical event or condition of the medical event may be further classified. For example, the procedure identified at 130 may be classified as a routine procedure, a preventative procedure, or a procedure associated with a chronic condition.

[0030]

[0044] 2 illustrates a workflow of a first method 200 for analyzing the word structure of a text string describing an event and classifying the event based on the word structure of the text string. Transformed text data (e.g., modelable data 120 described with respect to FIG. 1) can be received by a system of the present disclosure (210). The transformed data can include a series of numeric identifiers corresponding to individual words in the text string. In some embodiments, the numeric identifiers corresponding to individual words can be assigned based on words identified in past text strings (e.g., text strings previously received by the system). Words in the list of words (e.g., including words previously identified in past text strings or training text strings) can be identified (220) as either present in the transformed data or absent from the transformed data. The list of words may include up to 100, up to 200, up to 300, up to 400, up to 500, up to 600, up to 700, up to 800, up to 900, up to 1000, up to 5000, up to 10,000, up to 20,000, up to 30,000, up to 40,000, up to 50,000, up to 100,000, up to 125,000, up to 150,000, up to 175,000, or up to 200,000 previously identified words. The list of words may include at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 5000, at least 10,000, at least 20,000, at least 30,000, at least 40,000, at least 50,000, at least 100,000, at least 125,000, at least 150,000, at least 175,000, or at least 200,000 previously identified words. In some embodiments, new words present in the text string that correspond to numeric identifiers may be identified. In such cases, numeric identifiers may be assigned to the new words.In some embodiments, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% of the words present in the text string are assigned a numeric identifier. In some embodiments, up to 50%, up to 55%, up to 60%, up to 65%, up to 70%, up to 75%, up to 80%, up to 85%, up to 90%, up to 91%, up to 92%, up to 93%, up to 94%, up to 95%, up to 96%, up to 97%, up to 98%, up to 99%, or 100% of the words present in the text string are assigned a numeric identifier. In an exemplary implementation, a matrix containing numeric identifiers for all previously identified words can be populated with ones and zeros to indicate the presence or absence of a word in the text string, respectively. When a new word is identified, a new element containing the numeric identifier of the new word can be added to the matrix. Word combinations present in the transformed data can then be identified (230). Significant word combinations that may indicate a particular condition can be determined using machine learning. For example, a machine learning model can be trained using the transformed text data associated with one or more conditions. In some embodiments, words that frequently occur in combination in the text string that correspond to the same condition can be identified as significant word combinations. In some embodiments, word combinations can include at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 words. In some embodiments, word combinations can include at most 2, at most 3, at most 4, at most 5, at most 6, at most 7, at most 8, at most 9, or at most 10 words or more. If significant word combinations are identified in the transformed data, the text string can be identified as corresponding to a state.In some embodiments, a word combination can be identified as a significant word combination if the word combination indicates a state. Significant word combinations can be identified from at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 5000, at least 10,000, at least 20,000, at least 30,000, at least 40,000, at least 50,000, or at least 100,000 significant word combinations. Significant word combinations may be identified from up to 100, up to 200, up to 300, up to 400, up to 500, up to 600, up to 700, up to 800, up to 900, up to 1000, up to 5000, up to 10,000, up to 20,000, up to 30,000, up to 40,000, up to 50,000, or up to 100,000 significant word combinations. In some embodiments, the states corresponding to word combinations may differ from the states corresponding to individual words of the word combinations. The text data may be classified (240) based on word configurations (e.g., identified words or word combinations) or based on the identified states. In some embodiments, the text data may be classified using a machine learning model trained on previous classified text data corresponding to one or more states.

[0031]

[0045] Classifying 240 the data may include identifying one or more states using one or more independent processes. The independent processes may determine a state independently from the determination of the second state. For example, the determination of a state identified by an independent process may be unaffected by the identification of the second state. Methods of the present disclosure may include at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, or at least 50 independent processes. The disclosed methods may include up to 1, up to 2, up to 3, up to 4, up to 5, up to 6, up to 7, up to 8, up to 9, up to 10, up to 11, up to 12, up to 13, up to 14, up to 15, up to 16, up to 17, up to 18, up to 19, up to 20, up to 25, up to 30, up to 35, up to 40, up to 45, or up to 50 independent processes. The independent processes may identify a condition from a type of condition. For example, the type of condition may be a medical condition, a medical procedure, a medication, a treatment, a diagnosis, or a cost. The processes (e.g., the independent processes) may process a text string. In some embodiments, the process processes the entire text string. In some embodiments, the process may identify relevant portions of the text string. Determining the multiple conditions identified by the independent processes is described in further detail with respect to FIG. 9.

[0032]

[0046] FIG. 3 shows an exemplary schematic diagram of a neural network that can be implemented in the methods of the present disclosure. The neural network can include an input layer 310 including a number of input neurons 311, one or more hidden layers 320 including a number of hidden neurons 321, and an output layer 330 including a number of output neurons 331. The input neurons can be connected to one or more hidden neurons by input parameters 315, and the hidden layer neurons can be connected to one or more output neurons by output parameters 325. The hidden layer neurons can be connected to one or more input layer neurons. The output layer neurons can be connected to one or more hidden layer neurons. The input parameters can include weights based on the frequency, occurrence, or probability of connections or interactions. The output parameters can include weights based on the frequency, occurrence, or probability of connections or interactions. The hidden parameters can include weights based on the frequency, occurrence, or probability of connections or interactions. The input layer neurons can be activated or deactivated based on the presence or absence of the input parameters, respectively.

[0033]

[0047] The input layer may include up to 100, up to 200, up to 300, up to 400, up to 500, up to 600, up to 700, up to 800, up to 900, up to 1000, up to 5000, up to 10,000, up to 20,000, up to 30,000, up to 40,000, up to 50,000, up to 100,000, up to 125,000, up to 150,000, up to 175,000, or up to 200,000 input neurons. The input layer may include at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 5000, at least 10,000, at least 20,000, at least 30,000, at least 40,000, at least 50,000, at least 100,000, at least 125,000, at least 150,000, at least 175,000, at least 200,000, or at least 1 million input neurons. For example, the input layer neurons may correspond to words. In some embodiments, the input layer may include an input neuron for each word identified in the training text dataset. Input neurons corresponding to words present in the test dataset may be activated, and input neurons corresponding to words not present in the test dataset may be deactivated. A hidden layer may include up to 10, up to 20, up to 30, up to 40, up to 50, up to 60, up to 70, up to 80, up to 90, up to 100, up to 200, up to 300, up to 400, up to 500, up to 1000, up to 2000, up to 3000, up to 4000, or up to 5000 hidden neurons. For example, a hidden layer may include a hidden neuron for each word identified in the training text dataset. Neural networks of the present disclosure may be trained using text data corresponding to one or more states or one or more classifications. Input parameters connecting input neurons to hidden neurons may include weights representing the frequency with which the word corresponding to the input neuron occurs in combination with the word corresponding to the hidden neuron in the training text dataset.A larger weight may indicate a higher frequency of occurrence. The output parameters connecting hidden neurons to output neurons may include weights that represent the frequency with which the word combination corresponding to the hidden neuron is associated with a state or classification in the training text dataset. A larger weight may indicate a higher association frequency. The output layer may include up to 100, up to 500, up to 1000, up to 2000, up to 3000, up to 4000, up to 5000, up to 6000, up to 7000, up to 8000, up to 9000, up to 10,000, up to 11,000, up to 12,000, up to 13,000, up to 14,000, or up to 15,000 output neurons. The output layer may include at least 100, at least 500, at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, at least 6000, at least 7000, at least 8000, at least 9000, at least 10,000, at least 11,000, at least 12,000, at least 13,000, at least 14,000, or at least 15,000 output neurons. For example, the output layer may include an output neuron for each health state, condition, or diagnostic classification that can be identified based on the input text dataset. The output layer neurons may include a probability corresponding to the probability that the input text dataset will be classified as the health state, condition, or diagnosis corresponding to the output layer neuron. In some embodiments, the probabilities of the output layer neurons sum to 1.

[0034]

[0048] In some embodiments, the neural network of the present disclosure may be a convolutional neural network (CNN) including an input layer, an output layer, and multiple hidden layers. The convolutional neural network may include 1, 2, 3, 4, 5, 6, 7, 8, 9, or at least 10 hidden layers. In some embodiments, the convolutional neural network may include up to 2, up to 3, up to 4, up to 5, up to 6, up to 7, up to 8, up to 9, or at least 10 hidden layers. In some embodiments, the convolutional neural network may include at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 hidden layers. Input neurons may be connected to one or more hidden neurons by input parameters. Hidden layer neurons may be connected to one or more output neurons by output parameters. Hidden layer neurons in a first hidden layer may be connected to one or more hidden layer neurons in a second hidden layer by hidden parameters. The input parameters may include weights based on the frequency, occurrence, or probability of connections or interactions. Output parameters may include weights based on the frequency, occurrence, or probability of connections or interactions. Hidden parameters may include weights based on the frequency, occurrence, or probability of connections or interactions. Input layer neurons may be activated or deactivated based on the presence or absence of input parameters, respectively.

[0035]

[0049] 4 illustrates a workflow of a second method 400 for analyzing the word structure of a text string describing an event, assigning one or more states to the event, and classifying the event based on the word structure or one or more states of the text string using a neural network. In some embodiments, method 400 may implement the neural network described with respect to FIG. 3. Modelable data that has been converted from the text string (e.g., the converted data 120 described with respect to FIG. 1) may be received by a system of the present disclosure (410). The presence of words may be identified in the text string based on the presence of numeric identifiers in the converted data (420). Neurons of the trained neural network that correspond to words present in the text string may be activated (430). In some embodiments, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% of the words present in the text string correspond to neurons, in some embodiments, up to 50%, up to 55%, up to 60%, up to 65%, up to 70%, up to 75%, up to 80%, up to 85%, up to 90%, up to 91%, up to 92%, up to 93%, up to 94%, up to 95%, up to 96%, up to 97%, up to 98%, up to 99%, or 100% of the words present in the text string correspond to neurons. Hidden layer neurons can be activated 440 based on word combinations present in the text string, which can include at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 words.In some embodiments, the word combinations may include up to 10, up to 20, up to 30, up to 40, up to 50, up to 60, up to 70, up to 80, up to 90, up to 100, up to 200, up to 300, up to 400, up to 500, up to 1000, up to 2000, up to 3000, up to 4000, or up to 5000 or more words. In some embodiments, all possible word combinations in the text string are identified. In some embodiments, all possible word combinations in the text string that may indicate a condition are identified. Word combinations can be identified from at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 5000, at least 10,000, at least 20,000, at least 30,000, at least 40,000, at least 50,000 or at least 100,000 word combinations. Word combinations may be identified from at most 100, at most 200, at most 300, at most 400, at most 500, at most 600, at most 700, at most 800, at most 900, at most 1000, at most 5000, at most 10,000, at most 20,000, at most 30,000, at most 40,000, at most 50,000, or at most 100,000 word combinations. In some embodiments, a word combination may be identified as a significant word combination if the word combination indicates a state. In some embodiments, word combinations in the text string that do not indicate a state are not identified. If the weights of the input parameters connecting the neurons corresponding to a first word or a first combination of words are similar to the weights of the input parameters connecting the second word or a second combination of words, then the first word or first combination of words may correspond to the same state as the second word or second combination of words.For example, if the weights of the input parameters connecting neurons associated with the word "kidney" are similar to the weights of the input parameters connecting neurons associated with the word "renal," then the word "kidney" may be identified as being synonymous with the word "renal." One or more states corresponding to the text data may be identified 450 based on word configurations (e.g., words or word combinations) present in the text string. The states may correspond to output neurons. The output neurons may correspond to possible states. A trained neural network of the present disclosure may include up to 100, up to 500, up to 1000, up to 2000, up to 3000, up to 4000, up to 5000, up to 6000, up to 7000, up to 8000, up to 9000, up to 10,000, up to 11,000, up to 12,000, up to 13,000, up to 14,000, up to 15,000, up to 16,000, up to 17,000, up to 18,000, up to 19,000, or up to 20,000 or more states. A trained neural network of the present disclosure may include at least 100, at least 500, at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, at least 6000, at least 7000, at least 8000, at least 9000, at least 10,000, at least 11,000, at least 12,000, at least 13,000, at least 14,000, at least 15,000, at least 16,000, at least 17,000, at least 18,000, at least 19,000, or at least 20,000 states. One or more states can be identified using the trained neural network based on the frequency of associations between words or word combinations and states in a training dataset. Related states can be identified (460) based on states frequently associated with the state identified in the test text string in the training dataset.

[0036]

[0050] Identifying 450 the possible states may include identifying one or more states using one or more independent processes. The independent processes may determine the state independently from the determination of the second state. For example, the determination of the state identified by the independent processes may be unaffected by the identification of the second state. Methods of the present disclosure may include at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, or at least 50 independent processes. The disclosed methods may include up to 1, up to 2, up to 3, up to 4, up to 5, up to 6, up to 7, up to 8, up to 9, up to 10, up to 11, up to 12, up to 13, up to 14, up to 15, up to 16, up to 17, up to 18, up to 19, up to 20, up to 25, up to 30, up to 35, up to 40, up to 45, or up to 50 or more independent processes. The independent processes may identify a condition from a type of condition. For example, the type of condition may be a medical condition, a medical procedure, a medication, a treatment, a diagnosis, or a cost. A process (e.g., an independent process) may process a text string. In some embodiments, the process processes the entire text string. In some embodiments, the process may identify relevant portions of the text string. Determining the multiple conditions identified by the independent processes is described in further detail with respect to FIG. 9.

[0037]

[0051] The text string or one or more conditions can be classified based on the identified conditions (470). In some embodiments, classifying the text string can include determining an outcome based on the one or more conditions. Determining an outcome can include determining a probability of the outcome. The outcome can be determined using an aggregator to identify a most likely outcome based on multiple conditions. Determining a likely outcome based on multiple conditions is described in more detail with respect to FIG. 9. The outcome can be a binary outcome. For example, a binary outcome can include yes, no, approve, reject, support, reject, and the like. The outcome can be a non-binary outcome. For example, a non-binary outcome can include cost, prognosis, or success rate. The outcome can be reported to the user in a report. In some embodiments, the report can include an outcome based on one or more identified conditions and a reason for the outcome.

[0038]

[0052] Automated Claims Processing Engine

[0039]

[0053] In one aspect of the present disclosure, a claims processing engine is provided for automatically processing pet invoice data and generating claims processing results. The claims processing engine may employ machine learning techniques, as described elsewhere herein, to improve the speed and accuracy of claims processing with little or no human intervention.

[0040]

[0054] The provided claims processing engine can employ a parallel processing architecture to reduce prediction latency. For example, the claims processing engine can include multiple state inference engines, each of which includes a trained classifier or predictive model. The multiple state inference engines can operate in parallel to process input claim data, and the output of the multiple state inference engines can be aggregated to generate a claims processing output. Utilizing multiple trained classifiers operating in parallel instead of a single classifier can beneficially reduce total prediction latency. Moreover, the multiple state inference engines can operate independently, thereby providing flexibility in retraining, updating, or managing individual predictive models without affecting the performance of other predictive models.

[0041]

[0055] In some cases, an insurance claims processing engine can employ an optimized parallel data processing mechanism that distributes load based on insurance products. For example, incoming claim data related to different insurance products can be routed to different models corresponding to the same state. The selection of different models and the routing of incoming claim data can depend on differences between insurance products. For example, when two insurance products are the same except for a time constraint of the insurance products, such as latency. The latency can be roughly the latency to process the claim or classify the event. The insurance claims processing engine can spin up two separate and independent latency models (both for predicting latency states) and route traffic to the appropriate model while still utilizing any other models. For example, the insurance claims processing engine can provide two different machine learning algorithm-trained models corresponding to the same state and select a model from the two different machine learning algorithm-trained models to process input features based on the characteristics of the insurance product / event. The optimized load balancing mechanism can beneficially improve the efficiency of claims processing by dynamically routing data streams to different models (for predicting the same state) corresponding to different features of the insurance products.

[0042]

[0056] FIG. 8 schematically illustrates an insurance claims processing system 800 according to some embodiments of the present invention. The insurance claims processing system 800 may include a claims processing engine 810, which includes multiple state inference engines 813-1, 813-2, ... 813-n, each configured to receive input features generated by a corresponding transformation engine 811-1, 811-2, ... 811-n. The insurance claims processing system may include multiple parallel pipelines, each including a transformation engine and a state inference engine. The output of the multiple state inference engines is aggregated by an aggregator 815 to generate output data 809. The output data 809 may be related to a claim processing result. In some examples, the output data may be further validated or processed by an agent to generate a claim processing result.

[0043]

[0057] In some embodiments of the present disclosure, the claims processing system 800 may include a data input module 803 configured to receive and pre-process input data. In some instances, the data input module 803 may receive request data 801 indicating the submission of a claim. The request data 801 may be submitted by a user (e.g., a pet owner) via a client application or by a veterinary clinic via a veterinary client application.

[0044]

[0058] In some cases, the request data may include billing data received as an online form submission, email text, a word processing document, a portable document format (PDF), an image of text (e.g., an invoice), or other format. The data input module 803 may utilize any suitable technique for extracting billing data, such as optical character recognition (OCR) or transcription. More details about OCR and transcription methods are described with respect to Figures 8A-8D.

[0045]

[0059] In some cases, the input data received by the data input module 803 may include claims data obtained from a claims database and / or various bills and documents associated with the insurance claim. As described above, the input data may be related to the insurance claim, such as structured claims data obtained from the claims data store 805 or an insurance system. For example, the structured claims data may be submitted by a veterinary clinic or pet owner in a customized claim form, electronically or otherwise, via a veterinary practice management system. In some cases, the structured claims data may include text data such as insurance policy ID / number, illness / injury, or other fields for the pet or treatment. In some cases, the input data may include structured text data such as JavaScript Object Notation (JSON) data. In optional cases, the input data may include unstructured data related to the claim, such as an insurance claim form, invoice images, medical reports, police investigation reports, emails, or web-based content. Unstructured input data, such as email or invoice images, may be processed by the data input module 803 to extract claim data prior to processing by the claims processing engine 810.

[0046]

[0060] In some instances, the data input module 803 may include a data integration agent that provides connectivity between the data input module and one or more databases. The data integration agent may include an abstraction engine that allows communication with various management systems and also has the ability to integrate with additional ones in an ad-hoc manner in the future. For example, the data abstraction engine may provide a data abstraction layer over any database, storage system, and / or stored data stored or persisted by the system. The data abstraction layer may include various components, subsystems, and logic for substitution standards and mapping to translate various incoming database access requests into appropriate queries of the underlying database. For example, the data abstraction layer sits between the claims processing engine / application and the underlying physical data. The data abstraction layer may define a collection of logical fields that are loosely coupled to the underlying physical mechanism (e.g., database) that stores the data. The logical fields can be used to create queries to search, retrieve, add, and modify data stored in the underlying database. This advantageously allows the claims processing system to communicate with various databases or storage systems through a uniform interface.

[0047]

[0061] In some embodiments, the data input module 803 can be in communication with one or more data resources 809, as shown in FIG. 8A. For example, the data input module can receive input data from one or more systems, platforms, or applications, such as via an application programming interface (API). In some cases, the one or more data sources can include an optical character recognition (OCR) engine or a transcription engine for processing raw input data. Alternatively, the OCR engine or transcription engine can be part of the data input module that processes input data received from one or more data sources.

[0048]

[0062] The OCR engine 809-1 may be capable of recognizing text data from image files, PDF files, scanned documents, photographs, or various other types of files, as described above. The OCR engine may utilize any suitable technique or method for processing images to recognize text data. For example, the OCR engine may include preprocessing techniques such as deskewing, despeckling, binarization, zoning, character segmentation, or normalization; text recognition techniques such as pattern matching or pattern recognition; computer vision techniques for feature extraction or neural networks; and postprocessing techniques such as neighborhood analysis or application of vocabulary constraints. In some cases, the OCR engine may include a neural network trained to recognize entire lines of text instead of focusing on single characters. The OCR output may include the location of the identified text, the predicted text, and the confidence of the prediction.

[0049]

[0063] The OCR engine of the present disclosure can improve the accuracy or success rate of text recognition by employing a proprietary algorithm that allows the OCR to accurately extract text relevant to a claim transaction while ignoring irrelevant text. For example, the OCR algorithm can process an image and extract claim-related information such as invoice number, pet name, treatment line items, price, sales tax, subtotal, discounts, and various other claim data.

[0050]

[0064] In some embodiments, an OCR algorithm can be implemented to (i) identify one or more anchors (i.e., anchor words) in an image, (ii) determine boundaries based on the anchors, and (iii) extract text data within the boundaries. In some cases, the OCR algorithm can further determine word combinations by grouping subsets of the text data based at least in part on properties identified for the text data. Figures 8B-8D show examples of input data processed by an OCR algorithm.

[0051]

[0065] FIG. 8B shows an example of an image processed by an OCR algorithm. The raw input data may be an image of an invoice. The image may include one or more anchor words 821. In some cases, the anchors may be predetermined text data based on the known format of the document. For example, if the document is an invoice, the anchors may be date, breakdown, quantity, unit price, discount, sales tax, amount, etc. The anchors may be items related to the billing process. In some cases, the items may be line items, and the item values, such as "3 / 29 / 2021" for the "Date" item or "1.00" for the "Quantity" item, may be located at known locations for the items. The location of the item values ​​(e.g., image coordinates or x-y coordinates) may be determined based on the detected location of the corresponding item (e.g., breakdown coordinates) and the known format of the document.

[0052]

[0066] An OCR algorithm can begin by identifying one or more anchors from an image document. FIG. 8C shows an example of anchors identified from an image input. The output 831 of the image processing may include the coordinates (x, y) of the identified anchor (e.g., breakdown, quantity, subtotal) and properties of the anchor 833, such as a predicted confidence (e.g., 95). In some cases, the coordinates may be image coordinates. Other user-defined coordinates may be used. The output may also include other properties of the identified anchor, such as the anchor's level, page number, block number, paragraph number, word number, width, height, and predicted text.

[0053]

[0067] The OCR algorithm can then determine boundaries for the anchor locations to separate the anchor's value items (e.g., line item text). For example, based on the known format where the item value for "Item Breakdown" is left-justified along the [0,0] to [0,100] axis and the unit price is left-justified along the [100,0] to [100,100] axis, the boundary locations are determined once the anchor "Item Breakdown" is identified at [0,0] in [x,y] coordinates, "Unit Price" is identified at [100,0] coordinates, and "Subtotal" is identified at [100,100] coordinates. In the example 835 shown, the item value text for "Breakdown" is filtered within the boundaries and identified using the OCR engine's neural network. The output 835 may include various properties of the recognized item values, such as coordinates (e.g., image coordinates, x, y coordinates) and confidence level, as well as various other properties, such as level, text width, height, predicted text, page number, block number, paragraph number, line number, and the like. In some cases, padding (e.g., + / - 5) may be used to adjust boundaries to ensure all text is identified.

[0054]

[0068] In some cases, the location of the boundaries can be determined based on the known format of the document. For example, the location of a line item value relative to a corresponding anchor can be known based on the invoice format or branding of the practice management software. The format can vary depending on the practice management software utilized by the veterinary clinic. In some cases, the system can pre-store various formats for claims or documents to be processed, and the algorithm can call each format to determine the boundaries.

[0055]

[0069] The OCR algorithm can determine word combinations by grouping subsets of the text data based at least in part on properties identified for the text data. For example, the OCR algorithm can further process the identified line item text / words to form grouped line items that correspond to the original word combinations. In some cases, the properties identified for the text data can be locations or coordinates associated with words. FIG. 8D shows an example of separated line item text grouped by line number. Groups of text or word combinations can correspond to line items (e.g., patient intent exams / consultations). The grouped line items or words can be word combinations, as described elsewhere herein.

[0056]

[0070] Alternatively, instead of predetermining anchors, the OCR algorithm may have a trained model that can identify text that is likely to be a line item or that is likely to be an anchor. For example, anchor words are identified by predicting the presence of line item words using a machine learning algorithm trained model. In some cases, the model may be a trained neural network that can process raw input images and predict text that is likely to be an anchor. This can beneficially identify anchors from documents of unknown format. The model may be trained using training data that includes labels indicating whether text is a line item or not. In some cases, the boundaries of each line item value may also be predicted using the trained model.

[0057]

[0071] Returning to Figure 8A, transcription engine 809-2 may be capable of transcribing an audio file into text. For example, a user may read an invoice or a portion of an invoice and submit the audio file via a user application. The transcription engine may then process the audio file to transcribe the invoice. The transcribed invoice data may be received by a data input module for further extraction of structured text data.

[0058]

[0072] 8 , in some instances, the data input module 803 can communicate with one or more databases 807 to retrieve relevant data upon receiving the request data 801. For example, the request data 801 may include information such as the pet's name, illness, insurance policy ID, and the like, and the data input module 803 can retrieve historical data (e.g., the pet's treatment history from any veterinary clinic, claims history, data from other insurance providers, etc.) from a history database based on the pet's name, policyholder name, and the like. In some examples, the data input module 803 can retrieve the insurance coverage plan, insurance policy, or other relevant data (e.g., pre-authorization verification rules) based on the policy ID to verify the validity of the submitted claim.

[0059]

[0073] In some instances, the data input module 803 may preprocess input data to extract and / or generate claim data to be processed by the claims processing engine. In some instances, the data input module 803 may employ predictive models or natural language processing (NLP) techniques to extract data points from claim data to extract claim data. The data input module may employ any suitable NLP technique, such as a parser, to perform syntactic analysis of the input text. A parser may include instructions for syntactically, semantically, and lexically analyzing the text content of an input document and identifying relationships between text fragments of the document. A parser utilizes syntactic and morphological information about individual words found in a dictionary or "vocabulary" or derived through morphological processing (organized in a lexical analysis stage). In an example, the input data analysis process may include multiple stages, including item creation, segmentation, lexical analysis, and syntactic analysis.

[0060]

[0074] In some cases, the data input module 803 may perform data cleansing (e.g., removing noise such as spelling mistakes, punctuation errors, and grammatical errors present in the text data, or correcting technical terms to standard language) or other processes to obtain a claims dataset. In some cases, the data input module 803 may aggregate data received or collected from various data sources and send the aggregated claims dataset to multiple transformation engines for further processing.

[0061]

[0075] The multiple conversion engines 811-1, 811-2, ... 811-n can be configured to generate input features that are fed to corresponding state inference engines. As described elsewhere herein, the conversion engines can convert text data into numerical values ​​(e.g., one-dimensional arrays, two-dimensional arrays, etc.). In some cases, the data received by the multiple conversion engines 811-1, 811-2, ... 811-n can be the same text data, and each conversion engine can be configured to convert a particular word / word combination from the input data. Alternatively or additionally, the data received by the multiple conversion engines can be different. For example, the data input module can partition the data sent to the multiple conversion engines based on a state or event.

[0062]

[0076] In some cases, the transformation engine or data input module may further include a translation layer. The translation layer may be capable of (i) identifying words outside the data distribution range of the multiple machine learning algorithm trained models, transformation engine, or state inference engine, and (ii) replacing the words with replacement words within the data distribution range of the multiple machine learning algorithm trained models, transformation engine, or state inference engine. The translation layer may be capable of replacing previously unseen text with text within the data distribution range of the model. This may advantageously avoid retraining a model or training a new model for the unseen text. For example, if a first veterinary market (e.g., Country A) uses unfamiliar treatments or medications, the claims processing engine may identify the unfamiliar text and replace them with similar treatments or medications used in a second market (e.g., Country B). The identification and translation of unfamiliar text may be performed based on the frequency of occurrence of the text. For example, the frequency of occurrence of all medications and treatments may be measured. If medication "A" occurs in 10% of claims in Country A and 0% of claims in Country B, and medication "B" occurs in 0% of claims in Country A and 10% of claims in Country B, then "A" and "B" can be determined to be a candidate language pair, or "B" can be suggested as a replacement for "A." In some cases, the language pair or replacement can be validated by experts in the field. In some cases, the translation layer can include trained models for identifying unfamiliar text / words and replacing them with familiar text or replacement words.

[0063]

[0077] It should be noted that the conversion engine and input data module are for illustrative purposes. The system may include additional optional components, subcomponents, or fewer components. For example, the input data module may be part of the conversion engine such that the conversion engine performs at least a portion of the functionality of the input data module. Similarly, an OCR engine or transcription engine may be part of the data input module. The data input module may implement an OCR algorithm or transcription algorithm to perform one or more operations of the OCR method or transcription method as described above.

[0064]

[0078] The input features generated by the multiple transformation engines 811-1, 811-2, ..., 811-n can be provided to corresponding state inference engines 813-1, 813-2, ..., 813-n. The state inference engines can include trained classifiers or predictive models for identifying specific states. The state inference engines can employ deep learning techniques, as described elsewhere herein, to process the input features and generate outputs 814-1, 814-2, ..., 814-n. For example, the state inference engines can process the input features generated by the corresponding transformation engines using a predictive model to output a specific medical condition associated with the insurance claim. The predictive model can be the same as that described in FIG. 3. The predictive model may be of any suitable type, including, but not limited to, unsupervised clustering methods (e.g., K-nearest neighbors), support vector machines (SVMs), naive Bayes classification, random forests, tree-based ensemble models, convolutional neural networks (CNNs), feed-forward neural networks, radial basis function networks, recurrent neural networks (RNNs), deep residual learning networks, and the like, as described elsewhere herein.

[0065]

[0079] The outputs 814-1, 814-2, ... 814-n of the condition inference engine may include a condition type. A condition type may include a medical category or description, such as dental care, preventive care, medical procedure, diet, health check, medication, terminal care, or treatment site. A condition type may be a billing category, such as cost or discount. A condition type may be a subject's health condition, such as a chronic condition, illness, or disease. The outputs 814-1, 814-2, ... 814-n may indicate the presence of one or more condition types or the likelihood of the presence of a condition. For example, the first output 814-1 may be a medical description, and the second output 814-2 may be a cost. The aggregator 815 may combine the outputs 814-1, 814-2, ... 814-n to generate output data 809 as the final result of the claims processing engine 810.

[0066]

[0080] The output data 809 may be a result of the claim processing. The output data may indicate a decision or status of the processed claim. For example, the output data may include the status of the claim, such as approved, denied, upheld, rejected, and the like. In some cases, the output data 809 may include a probability of a status / decision, such as a confidence level of claim approval or likelihood of fraud. In some cases, the aggregator 815 or one or more of the state inference engines may generate the probability of the status / decision based on business rules.

[0067]

[0081] In some cases, output data 809, such as a probability of a decision, can be determined based on the individual outputs of multiple state inference engines. For example, an aggregator 815 can aggregate outputs 814-1, 814-2, ... 814-n from each of the state inference engines 813-1, 813-2, ... 813-n to generate a probability. In some cases, the outputs 814-1, 814-2, ... 814-n from each of the state inference engines can be a probability of a type of state. The aggregator 815 can combine the outputs 814-1, 814-2, ... 814-n using any suitable method (e.g., linear combination, non-linear combination). In optional cases, the aggregator can include a predictive model to generate the output data based at least in part on business rules.

[0068]

[0082] In some cases, the output data 809 may include an explanation, such as, for example, the reason for the claim denial. The explanation may be determined based on one or more identified states as the output of a state inference engine and / or business rules. In some cases, the explanation may be an implicit insight (e.g., potential fraud) generated based on one or more identified states. The output may include insight (e.g., potential fraud) inferred by aggregating multiple states or at least a portion of the states. In some cases, the explanation may include one or more of the identified states to assist the agent in further validating the claim.

[0069]

[0083] In some cases, the event status or final output may include approval, rejection, or a request for further validation actions. In some examples, human intervention may be required to further validate / verify the claim based on a probability or confidence level. For example, if the confidence level of the claim's approval falls below a predetermined confidence threshold (e.g., 80%, 90%, or 99%), the output data 809 and associated claim may be sent to a user interface module for further review / processing by an agent. In some cases, feedback or input provided by the agent may be collected by the system for training / retraining the state inference engine. In some examples, human intervention may be required based on the payment amount. For example, when the identified state indicates that the payment will exceed a predetermined threshold (e.g., $500), the output data 809 (e.g., the payment amount) may be sent to a user interface module along with the claim for review by an agent.

[0070]

[0084] In some cases, the output data 809 may include information to assist the agent in validating or further processing the claim. For example, the output data 809 may include health conditions identified by one or more of multiple condition inference engines, highlight questionable health conditions or conditions, generate recommendations for the agent based on business rules, or include other identified conditions translated into terms that are easy for the agent to understand.

[0071]

[0085] The claims processing system may be a standalone system or self-contained component capable of operating and working independently and may communicate with other systems or entities (e.g., predictive modeling and management systems, insurance systems, third-party healthcare systems, etc.). Alternatively, the claims processing system may be a component or subsystem of another system. In some cases, the claims processing system provided herein may be a Platform as a Service (PaaS) and / or Software as a Service (SaaS) application configured to offer a suite of pre-built cross-industry applications developed on the platform to facilitate the automation of claims processing by various entities. In some cases, the claims processing system may be an on-premise platform where applications and / or software are hosted locally.

[0072]

[0086] The claims processing system or one or more components of the claims processing system may be implemented using software, hardware, or a combination of both. For example, the claims processing system may be implemented using one or more processors. The processor may be a hardware processor (which may be a single-core or multi-core processor), such as a central processing unit (CPU), a graphics processing unit (GPU), a general-purpose processing unit, or multiple processors for parallel processing. The processor may be any suitable integrated circuit, such as a computing platform or microprocessor, a logic device, and the like. While this disclosure is described with reference to a processor, other types of integrated circuits and logic devices may also be applicable. The processor or machine is not limited by its data manipulation capabilities. The processor or machine may perform 512-bit, 256-bit, 128-bit, 64-bit, 32-bit, or 16-bit data operations.

[0073]

[0087] FIG. 9 illustrates a workflow of a method 900 for determining a probable outcome based on multiple states identified in multiple processes. The method 900 can be implemented by an insurance claims processing system such as that described in FIG. 8. Text data (e.g., structured text data) can be provided to each process of the multiple processes (910). The text data can include transformed text data, such as the transformed data 120 described with respect to FIG. 1. The text data can be structured. In some embodiments, the text can be structured to indicate a type of information. The text can be structured to distinguish between subject information, event information, supporting information, or a combination thereof. For example, the text can be structured to indicate an item breakdown, treatment, procedure, diagnosis, subject name, historical data, insurance coverage, or a combination thereof. In some embodiments, the structured text data can include JavaScript Object Notation (JSON) data. A state process 920 (e.g., a first state process, a second state process, a third state process, or an nth state process) can determine a state based on the text data.

[0074]

[0088] The state process can identify a state from a type of state. The process can determine a state from a type of state. In some embodiments, the type of state can be dental treatment, preventative treatment, medical procedure, diet, health check, medication, terminal care, treatment site, cost, discount, chronic condition, disease, or illness. In some embodiments, the state process can be an independent state process. In some embodiments, the process can verify the identity of the subject. An independent state process can determine a state without being influenced by a second state process. For example, a first independent state process can determine a first state independently from one or more of a second state process, a third state process, or an nth state process. The independent process can function independently from the second state process, such that an error in the second state process does not interrupt the functionality of the state process. In some embodiments, two or more independent processes can be implemented in parallel. Implementing two or more independent processes in parallel can improve computer functionality by increasing the speed at which the processes are implemented. For example, a first state process may execute on a first central processing unit (CPU), CPU core, or graphics processing unit (GPU), a second state process may execute on a second CPU, CPU core, or GPU, a third state process may execute on a third CPU, CPU core, or GPU, and an nth state process may execute on an nth CPU, CPU core, or GPU. In some embodiments, the state processes may be dependent state processes. A dependent state process may depend on a second state process to determine a state. For example, a first dependent state process may determine a first state based on one or more of the second state process, the third state process, or the nth state process.

[0075]

[0089] The states identified from the multiple state process may be aggregated (930) to determine a probable outcome based on the states (940). The outcome may be a binary outcome. For example, binary outcomes may include yes, no, approve, reject, support, reject, and the like. The outcome may be a non-binary outcome. For example, non-binary outcomes may include cost, diagnosis, prognosis, or success rate. The probability of the outcome may be determined based on the individual probabilities of the association between each state and the outcome. In some embodiments, the probability of the outcome may be determined using machine learning. In some embodiments, the probability of the outcome may be determined by mathematically combining the individual probabilities of the association between each state and the outcome. The probable outcome may be the outcome with the highest probability determined by the aggregator. The probable outcome may include a confidence level describing the confidence in the result determined by the aggregator. The confidence level may be determined from one or more probabilities from one or more states. The confidence level may be determined from one or more types of information from the structured text data. In some embodiments, one or more types of information from the structured text data can be ignored when determining the confidence level. A high probability result can include an explanation, such as why the high probability result was identified. The explanation can be determined from one or more states.

[0076]

[0090] 10 schematically illustrates a platform 1000 upon which methods and systems for automated insurance claims processing may be implemented. The platform 1000 may include one or more user devices 1001, 1028, an insurance system 1020, one or more third party entities / systems 1030, and databases 1031, 1033. Each of the components may be operatively connected to one another via a network 1050 or via any type of communication link that allows for the transmission of data from one component to another.

[0077]

[0091] The insurance system 1020 may include one or more components, such as a predictive modeling and management system 1021, a claims processing system 1023, an insurance application 1027, or other components. The insurance system 1020 may be implemented as one or more computing resources or hardware devices. The insurance system 1020 may be implemented in one or more server computers, one or more cloud computing resources, and the like, each resource having one or more processors, memory, persistent storage, and the like. For example, the insurance system 1020 may include a web server, online services, pet insurance management components, and the like, for providing the insurance application 1027 to pet owners 1003 and / or veterinary clinics 1030. For example, the web server may be implemented as a hardware web server or a software-implemented web server, and may generate and exchange web pages with each computing device 1001, 1028 using a browser.

[0078]

[0092] The insurance application 1027 may include a software application (i.e., client software) for the veterinary clinic 1030 that enables information exchange between the hospital and the insurance system. For example, an application running on the hospital / veterinary clinic device (e.g., client / browser) may enable claims submission, issuing insurance service offers, searching PIMS data for clients, booking appointments, mapping clients between systems, and displaying information for all of these activities in a manner digestible by veterinary clinic employees, improving patient care. The application may be a cloud-powered application or a local application. The insurance application 1027 may also provide a software application (i.e., client software) for pet owners. The client application may enable pet owners 1003 to enroll in pet insurance, submit claims / invoices, track the status of submitted claims and the outcomes and payments of those claims, and the like.

[0079]

[0093] The insurance application 1027 or predictive model creation and management system may employ any suitable technology, such as containers and / or microservices. For example, the insurance application may be a containerized application. The insurance system may deploy a microservices-based architecture in its software infrastructure, such as implementing the insurance application or services within containers. In another example, the cloud application and / or predictive model creation and management system may provide a model management console underpinned by microservices.

[0080]

[0094] In some embodiments, a user (e.g., a pet owner 1003, a veterinary clinic 1030) can utilize a user device to interact with the insurance system 1020 through one or more software applications (i.e., client software) running on and / or accessed by the user device 1001, and the user device and the insurance system 1020 can form a client / server relationship.

[0081]

[0095] In some embodiments, the client software (i.e., the software application installed on the user device 1001) may be available as any of a variety of downloadable mobile applications for various types of mobile devices. Alternatively, the client software may be implemented in one or more combinations of programming and markup languages ​​for execution by various web browsers. For example, the client software may run in web browsers that support JavaScript and HTML rendering (e.g., Chrome, Mozilla Firefox, Internet Explorer, Safari, etc.), as well as any other compatible web browser. Various embodiments of the client software application may be compiled for various devices across multiple platforms and optimized for their respective native platforms. In some cases, the client software may allow a user to submit an insurance claim by capturing an image of an invoice. For example, the user may be permitted to submit the claim through a user interface (e.g., a mobile application) running on the user's mobile device, the user may be prompted to scan the insurance form with the mobile device's camera, and the user may receive claim processing results generated by the claims processing system 1023. The provided insurance claims processing system and method can process claims with reduced processing time, thereby improving the user claims processing experience.

[0082]

[0096] User devices 1001 associated with pet owners or veterinary clinics, and user devices 1028 associated with agents for claim processing or predictive model management, may be computing devices configured to perform one or more operations (e.g., rendering a user interface for claim submission, reviewing claim status, reviewing final output of the claims processing system, validating the claim, processing the claim, etc.). Examples of user devices may include, but are not limited to, mobile devices, smartphones / cell phones, wearable devices (e.g., smart watches), tablets, personal digital assistants (PDAs), laptops or notebook computers, desktop computers, media content players, televisions, video game stations / systems, virtual reality systems, augmented reality systems, microphones, or any electronic device capable of analyzing, receiving (e.g., receiving an image of an invoice or claim form, modifying fields on a claim form, agent-entered data, etc.), providing, or displaying to a user certain types of data (e.g., system-generated claims processing results, etc.). User devices may be handheld objects. User devices may be portable. User devices may be carried by a human user. In some cases, a user device is located remotely from a human user, and the user can control the user device using wireless and / or wired communications. A user device can be any electronic device with a display.

[0083]

[0097] The user devices 1001, 1028 may include a display. The display may be a screen. The display may or may not be a touchscreen. The display may be a light-emitting diode (LED) screen, an OLED screen, a liquid crystal display (LCD) screen, a plasma screen, or any other type of screen. The display may be configured to present a user interface (UI) or graphical user interface (GUI) rendered through an application (e.g., via an application programming interface (API) running on the user device). The GUI may present claims processing requests, the status of submitted claims, and interactive elements related to submitting a claims request (e.g., editable fields, claim forms, etc.). The user devices may also be configured to display web pages and / or websites on the Internet. One or more of the web pages / websites may be hosted by the server 1020 and / or rendered by the insurance system, as described above.

[0084]

[0098] The user device 1001 can be associated with one or more users (e.g., pet owners). In some embodiments, a user can be associated with a unique user device. Alternatively, a user can be associated with multiple user devices. A user (e.g., a pet owner) can register with the insurance platform. In some cases, for a registered user, user profile data can be stored in a database (e.g., database 1033) along with a user ID uniquely associated with the user. The user profile data can include, for example, pet name, pet owner name, geographic location, contacts, historical data, and various other information as described elsewhere herein. In some cases, a registered user can be required to log in to an insurance account using credentials. For example, to perform activities such as filing an insurance claim or reviewing the status of a claim, a user can be required to log in to the application via the user device 1001 by providing a passcode, scanning a QR code, biometric verification (e.g., fingerprint, face scan, retina scan, voice recognition, etc.), or various other verification methods.

[0085]

[0099] The predictive model creation and management system 1021 can be configured to train and develop predictive models. In some cases, trained predictive models can be deployed to claims processing systems 1023 or edge infrastructure through a predictive model update module. The predictive model update module can monitor the performance of the trained predictive models (e.g., state inference engines) after deployment and can retrain the models if performance drops below a predetermined threshold. In some cases, the predictive model creation and management system 1021 can also support ingesting data sent from user devices 1028 (e.g., agent feedback data) or other data sources 1031 into one or more databases or cloud storage 1033 for the ongoing training of one or more predictive models.

[0086]

[0100] The predictive model creation and management system 1021 may include applications that enable operational and management integration (including monitoring or storing data in the cloud or a private data center). In some embodiments, the predictive model creation and management system 1021 may include a user interface (UI) module for monitoring predictive model performance and / or configuring predictive models. For example, the UI module may render a graphical user interface on the computing device 1028 to allow the manager / agent 1029 to view model performance or provide user feedback. In some cases, data collected from the agent user device 1028, such as validating outputs generated by a claims processing system or confirming health conditions generated by a state inference engine, may be used by the predictive model creation and management system 1021 to train / retrain one or more predictive models.

[0087]

[0101] It should be noted that while the predictive modeling and management system is shown as a component of insurance system 1020, the predictive modeling and management system may be a standalone system. More details about the predictive modeling and management system are described with respect to FIG. 11.

[0088]

[0102] The claims processing system 1023 can be configured to perform one or more operations consistent with the disclosed methods described herein. The claims processing system 1023 can be the same as the claims processing system as illustrated in FIG.

[0089]

[0103] In one configuration, the insurance system 1020 may be software stored in memory accessible by a server (e.g., memory locally connected to the server or remote memory accessible over a communications link such as a network). Thus, in some aspects, the insurance system may be implemented as one or more computers, as software stored on memory devices accessible by the server, or as a combination thereof.

[0090]

[0104] However, the claims processing system 1023 is shown hosted on a server. The claims processing system 1023 can be implemented as a hardware accelerator, processor-executable software, and various other implementations. In some cases, the insurance system 1020 can employ an edge intelligence paradigm in which data processing and predictions are performed at the edge or edge gateway. For example, one or more of the predictive models can be built, developed, and trained on the cloud and run for inference on a user device and / or other device (e.g., a hardware accelerator) locally connected to the user or hospital. In some cases, the predictive models can undergo continuous training as new claims data and feedback data are collected. The continuous training can be performed on the cloud or on a server. In some cases, new claims data and agent feedback data can be sent to a remote server and used to update the model, and the updated model (e.g., updated model parameters) can be downloaded to the physical system (e.g., the claims processing system 1023) for implementation.

[0091]

[0105] The various functions performed by the insurance system, such as data processing, predictive model training, trained model execution, continuous training / retraining of predictive models, model monitoring, and the like, may be implemented in software, hardware, firmware, embedded hardware, standalone hardware, application-specific hardware, or any combination thereof. The predictive model creation and management system 1021, the claims processing system 1023, and the techniques described herein may be realized in digital electronic circuitry, integrated circuits, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, a central processing unit (CPU), a graphics processing unit (GPU), a general-purpose processing unit (which may be a single-core or multi-core processor), or multiple processors for parallel processing, and / or combinations thereof.

[0092]

[0106] In some cases, the insurance system 1020 may also be configured to store data and information, and to search, retrieve, and / or analyze data and information stored in one or more of the databases 1033, 1031. The data and information may include, for example, veterinary clinic information for the system, information about each of the insurance service offerings, information about each pet enrolled in the pet insurance system, historical data such as past pet insurance claims, data about predictive models (e.g., parameters, model architecture, training data sets, performance metrics, thresholds, etc.), data generated by the predictive models such as status or claims processing results, feedback data, and the like.

[0093]

[0107] Network 1050 may be a network configured to provide communication between the various components shown in FIG. 10 . The network, in some embodiments, may be implemented as one or more networks connecting devices and / or components in a network layout to enable communication therebetween. Direct communication may be provided between two or more of the above components. Direct communication may occur without the need for an intermediate device or network. Indirect communication may be provided between two or more of the above components. Indirect communication may occur using one or more intermediate devices or networks. For example, indirect communication may utilize a telecommunications network. Indirect communication may be performed using one or more routers, communication towers, satellites, or any other intermediate devices or networks. Examples of types of communication may include, but are not limited to, communication via the Internet, a local area network (LAN), a wide area network (WAN), Bluetooth, near field communication (NFC) technology, a network based on a mobile data protocol (such as General Packet Radio Service (GPRS), GSM, Enhanced Data GSM Environment (EDGE), 3G, 4G, 5G, or Long Term Evolution (LTE) protocol), infrared (IR) communication technology, and / or Wi-Fi, and may be wireless, wired, or a combination thereof. In some embodiments, the network may be implemented using cellular and / or pager networks, satellite, licensed radio, or a combination of licensed and unlicensed radio. The network may be wireless, wired, or a combination thereof.

[0094]

[0108] The user devices 1001, 1028, the veterinary clinic computer system 1030, or the insurance system 1020 can be connected or interconnected to one or more databases 1033, 1031. A database can be one or more memory devices configured to store data. Additionally, a database, in some embodiments, can be implemented as a computer system with storage devices. In one aspect, a database can be used by components of a network layout to perform one or more operations consistent with disclosed embodiments. The one or more local databases and the platform's cloud database can utilize any suitable database technique. For example, Structured Query Language (SQL) or "NoSQL" databases can be utilized to store claims data, pet / user profile data, historical data, predictive models, training datasets, or algorithms. Some of the databases can be implemented using various standard data structures, such as arrays, hashes, (linked) lists, structures, structured text files (e.g., XML), tables, JavaScript Object Notation (JSON), NOSQL, and / or the like. Such data structures can be stored in memory and / or (structured) files. In another alternative, an object-oriented database can be used. An object database may contain many collections of objects grouped together and / or linked together by common attributes, and many collections of objects may be related to other collections of objects by some common attribute. An object-oriented database performs similarly to a relational database, except that objects are not simply pieces of data but may have other types of functionality encapsulated within a given object. In some embodiments, the database may include a graph database that uses a graph structure for semantic-based queries with nodes, edges, and properties for representing and storing data.If the database of the present invention is implemented as a data structure, the use of the database of the present invention can be incorporated into another component, such as a component of the present invention. The database can also be implemented as a mixture of data structures, object and relational structures. The database can be consolidated and / or distributed in many variations through standard data processing techniques. Portions of the database (e.g., tables) can be exported and / or imported, thus allowing for decentralization and / or consolidation.

[0095]

[0109] In some embodiments, the insurance system 1020 can build databases for fast and efficient data retrieval, querying, and distribution. For example, the predictive modeling and management system 1021 or the claims processing system 1023 can provide customized algorithms to extract, transform, and load (ETL) data.

[0096]

[0110] In some cases, the database 1033 can store data related to the predictive models. For example, the database can store data about trained predictive models (e.g., parameters, hyperparameters, model architecture, performance metrics, thresholds, rules, etc.), data generated by the predictive models (e.g., intermediate results, model outputs, latent features, inputs and outputs of components of the model system, etc.), training datasets (e.g., labeled data, user feedback data, etc.), predictive models, algorithms, and the like. The database can store algorithms or rule sets utilized by one or more methods disclosed herein. For example, a predefined rule set used by an aggregator in combination with a machine learning trained model can be stored in the database. In some embodiments, one or more of the databases can be co-located with a server, co-located with each other over a network, or remotely located from other devices. One skilled in the art will recognize that the disclosed embodiments are not limited to the configuration and / or arrangement of the databases.

[0097]

[0111] In some cases, data stored in database 1033 is available or accessible by various applications through an application programming interface (API). Access to the database may be authorized per API level, per data level (e.g., type of data), per application level, or according to other authorization policies.

[0098]

[0112] While particular computing devices are shown and networks are described, it will be appreciated and understood that other computing devices and networks can be utilized without departing from the spirit and scope of the embodiments described herein. Additionally, as will be appreciated by those skilled in the art, one or more components of a network layout can be interconnected in a variety of ways, and in some embodiments can be directly connected to one another, co-located, or remotely located.

[0099]

[0113] 11 illustrates a schematic of a predictive modeling and management system 1100 according to some embodiments of the present invention. In some cases, the predictive modeling and management system 1100 may include services or applications running in a cloud or on-premise environment for remotely configuring and managing the insurance claims processing system. This environment may run on one or more public clouds (e.g., Amazon Web Services (AWS), Azure, etc.) and / or in a hybrid cloud configuration (where one or more parts of the system run on a private cloud and other parts run on one or more public clouds).

[0100]

[0114] In some embodiments of the present disclosure, the predictive model creation and management system 1100 may include a model training module 1101 configured to train, develop, or test predictive models using data from a cloud data lake and a metadata database. The model training process may further include operations such as model pruning and compression to improve inference speed. Model pruning may include removing nodes of a trained neural network that do not affect the network output. Model compression may include using lower precision network weights, such as using 16 floating point instead of 32. This may beneficially enable real-time inference (e.g., at high inference speeds) while ensuring model performance.

[0101]

[0115] In some cases, the predictive model creation and management system 1100 may include a model monitor system that monitors data drift or performance of models in different phases (e.g., development, deployment, prediction, validation, etc.) The model monitor system may also perform data integrity checks for models being deployed in development, test, or production environments.

[0102]

[0116] The model monitor system can be configured to perform data / model integrity checks to detect data drift and accuracy degradation. The process can begin by detecting data drift in the training and prediction data. During training and prediction, the model monitor system can monitor differences in the distributions of the training, testing, validation, and prediction data, changes in the distributions of the training, testing, validation, and prediction data over time, covariates causing changes in the prediction output, and a variety of other things.

[0103]

[0117] In some cases, the model monitor system may include an integrity engine that runs one or more integrity tests on the model, the results of which may be displayed on the model management console. For example, the integrity test results may indicate the number of failed predictions, the percentage of row entries that failed the test, the test execution time, and details of each entry. Such results may be displayed to a user (e.g., a developer, a manager, etc.) via the model management console.

[0104]

[0118] Data monitored by the model monitor system may include data involved during model training and production. Data in model training may include, for example, training, testing, and validation data, predictions, or statistics characterizing the datasets (e.g., the mean, variance, and higher moments of the dataset). Data involved in production time may include time, input data, predictions made, and confidence limits for the predictions made. In some embodiments, ground truth data may also be monitored. Ground truth data may be monitored to assess model accuracy and / or to trigger model retraining. In some cases, a user may provide ground truth data (e.g., proxy feedback) to the predictive model creation and management system 1100 after the model enters the deployment phase. The model monitor system may monitor changes in the data, such as changes in ground truth data, or monitor new training or prediction data as they become available.

[0105]

[0119] As explained above, multiple state inference engines can be individually monitored or retrained as soon as model performance is detected to fall below a threshold. During prediction time, predictions can be associated with the model to track data drift or incorporate feedback from new ground truth data.

[0106]

[0120] In some cases, the predictive model creation and management system 1100 may also be configured to manage data flow between various components (e.g., cloud data lake, metadata database, claims processing engine, model training module), and provide precise, complex, and fast queries (e.g., model queries, training data queries), model deployment, maintenance, monitoring, model updates, model versioning, model sharing, and various others.

[0107]

[0121] The methods of the present disclosure (e.g., the methods illustrated in Figures 1, 2, 4, 9, or combinations thereof) can be implemented on a system (e.g., the system illustrated in any one of Figures 5-8) as described herein. The method can classify events based on a text string describing the event. The method can identify at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, or at least 50 states of the event. The method may identify up to 1, up to 2, up to 3, up to 4, up to 5, up to 6, up to 7, up to 8, up to 9, up to 10, up to 11, up to 12, up to 13, up to 14, up to 15, up to 16, up to 17, up to 18, up to 19, up to 20, up to 25, up to 30, up to 35, up to 40, up to 45, or up to 50 or more event states. The events may be classified based on the identified states. For example, events may be classified as one or more of: up to 100, up to 500, up to 1000, up to 2000, up to 3000, up to 4000, up to 5000, up to 6000, up to 7000, up to 8000, up to 9000, up to 10,000, up to 11,000, up to 12,000, up to 13,000, up to 14,000 or up to 15,000 or more classifications. Events can be classified as one or more of at least 100, at least 500, at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, at least 6000, at least 7000, at least 8000, at least 9000, at least 10,000, at least 11,000, at least 12,000, at least 13,000, at least 14,000 or at least 15,000 classifications.In some embodiments, the method can classify an event in less than about 1 second, less than about 2 seconds, less than about 3 seconds, less than about 4 seconds, less than about 5 seconds, less than about 6 seconds, less than about 7 seconds, less than about 8 seconds, less than about 9 seconds, less than about 10 seconds, less than about 15 seconds, less than about 20 seconds, less than about 25 seconds, less than about 30 seconds, less than about 35 seconds, less than about 40 seconds, less than about 45 seconds, less than about 50 seconds, less than about 55 seconds, less than about 60 seconds, less than about 70 seconds, less than about 80 seconds, less than about 90 seconds, less than about 100 seconds, less than about 110 seconds, or less than about 120 seconds.

[0108]

[0122] FIG. 5 illustrates a system 500 of the present disclosure for training and performing a method for identifying and classifying one or more conditions (e.g., method 200 described with respect to FIG. 2 or method 400 described with respect to FIG. 4 ). The system may include a condition classification module 510 capable of identifying one or more conditions of a text string. The condition classification system may include a non-transitory computer-readable medium 515. The non-transitory computer-readable medium may include read-only memory, random-access memory, flash memory, a hard disk, a semiconductor memory, a tape drive, a disk drive, or any combination thereof. The non-transitory computer-readable medium may further include a data area capable of storing data including text strings 516, a training data set 517, a trained model 518, and classification or condition data 519. In some embodiments, the condition classification system may include a user interface 511, a conversion process 512, a training set generator process 513, and a machine learning process 514. The user interface 511 may allow a user to interact with the system of the present disclosure to perform the methods of the present disclosure. The conversion process 512 can be configured to convert the text string data into modelable data (e.g., data including numeric identifiers corresponding to words in the text string data). The text strings or the converted data, or both, can be stored in the text string data area 516. The training set generator process 513 can be configured to generate training set data from the text string data associated with one or more classifications or one or more states. The training set can be stored in the data area 517. A trained model can be prepared based on the training data set and stored in the data area 518. The machine learning process 514 can implement the trained model to identify one or more states or one or more classifications of the text string and store in the data area 519.

[0109]

[0123] The state classification system 510 may be operatively connected to input users 530 or output users 540, or both, through a communication network 520. The input users may interact with the communication network through an input data interface 531. The output users may interact with the communication network through a classification interface 541. The communication network may be configured to receive event description information 535 from the input users and provide the event description information to the state classification system 510. The event description information may be stored as a text string in the text string data area 516. The communication network may be configured to receive state or classification information from the state classification system. The state or classification information may be stored in the state data area 519. The state or classification information may be provided to the output users through the classification interface. In some embodiments, the input users and the output users may be the same.

[0110]

[0124] FIG. 6 illustrates a system 600 of the present disclosure for training and implementing a method for identifying and classifying one or more states using a neural network (e.g., method 400 described with respect to FIG. 4 ). A transformation engine 630 can receive text string data from a network 610 or a data store 620, or both. In some embodiments, the transformation engine can convert the text string data into modelable data. For example, the modelable data can include numerical identifiers corresponding to words present in the text string data. The converted data can be stored in a data store or provided to a user over a network. A word construction engine 640 can identify one or more words present in a modelable dataset prepared by the transformation engine. A state identification engine 650 can identify one or more states of the dataset based on the words identified by the word construction engine. The state identification engine can include a related state identification engine 651 to identify relationships between two or more states. The state identification engine can include a state likelihood engine 652 that can determine the likelihood that a state is associated with the dataset or the likelihood that a first state is related to a second state, or both. A training engine 660 can train the neural network using the training data set. The training engine can interact with the associated state identification and state likelihood engine to adjust the associated state identification and state likelihood based on the training data. A classification engine 670 can use the trained state identification engine to identify one or more states or classifications for the transformed text string. The classifications can be stored in a data store or communicated to a user over a network.

[0111]

[0125] FIG. 7 illustrates a method 700 of operating a system for identifying and classifying one or more conditions. Beginning at step 711, the system may receive text data including a description of an event. At step 712, the system may receive condition data and classification data corresponding to the event description text received at step 711. At step 713, the event description data may be converted into modelable data, which may be used to generate a training set at step 714. At step 715, a model may be generated based on the training set. The trained model at step 711 may be provided to the system for iteratively training the model. The trained model may be used to implement the method beginning at step 721. At step 721, a user may provide an event description. The event description may be unclassified. At step 731, the system may receive the event description. At step 732, the event text data may be converted into modelable data. At step 733, words may be identified in the converted text data. At step 734, the model generated at step 715 may be used to identify one or more conditions associated with the text data. In step 735, associated states associated with the state identified in step 734 may be identified. In step 736, the text data may be classified based on the states identified in steps 734 and 735. In step 722, the state data and classification data may be reported to the user.

[0112]

[0126] Unless otherwise defined, all technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. As used in this specification and the appended claims, singular forms such as "a," "an," and "the" include plural references unless the context clearly dictates otherwise. References herein to "or" are intended to include "and / or" unless expressly stated otherwise.

[0113]

[0127] Whenever the terms "at least," "greater than," or "greater than or equal to" precede the first number in a series of two or more numbers, the terms "at least," "greater than," or "greater than or equal to" apply to each number in the series. For example, greater than or equal to 1, 2, or 3 is equivalent to greater than or equal to 1, greater than or equal to 2, or greater than or equal to 3.

[0114]

[0128] Whenever the terms "not greater than," "less than," "less than or equal to," or "up to" precede the first number in a series of two or more numerical values, the terms "not greater than," "less than," "less than or equal to," or "up to" apply to each number in the series. For example, less than or equal to 3, 2, or 1 is equivalent to less than or equal to 3, less than or equal to 2, or less than or equal to 1.

[0115]

[0129] Where values ​​are described as ranges, such disclosure will be understood to include disclosure of all possible subranges within such ranges as well as specific numerical values ​​falling within such ranges, whether or not a specific numerical value or specific subrange is specified.

[0116]

[0130] While preferred embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. The present invention is not intended to be limited by the specific examples provided within the specification. While the present invention has been described with reference to the foregoing specification, the description and illustration of the embodiments herein are not intended to be construed in a limiting sense. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. Moreover, all aspects of the present invention are not limited to the specific depictions, configurations, or relative proportions set forth herein, but are to be understood as depending upon a variety of conditions and variables. It should be understood that various alternative forms of the embodiments of the invention described herein may be employed in the practice of the present invention. Accordingly, it is contemplated that the present invention shall cover any such alternatives, modifications, variations, or equivalents. It is intended that the following claims define the scope of the invention, and that methods and structures within the scope of these claims and their equivalents be covered thereby.

Claims

[Claim 1] The invention described in this specification.