Wood wide models for improved machine learning pipeline

The machine learning pipeline addresses the limitations of existing models by iteratively extracting neural vectors and applying structural causal models, enhancing scalability and interpretability for numeric data, particularly in high-stakes domains.

WO2026080714A1PCT designated stage Publication Date: 2026-04-16CARNEGIE MELLON UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/050252
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-10-09
Filing Date
2025-10-09
Publication Date
2026-04-16

AI Technical Summary

Technical Problem

Current generative AI foundation models are ill-suited for numeric tabular or time-series data, requiring time and resource-intensive manual techniques and lacking scalability and transferability, which is critical for business operations and decision-making.

Method used

A computer-implemented machine learning pipeline that iteratively extracts identifiable neural vectors, conceptual interpretable factors, and applies a structural causal model with shared probability distributions to address causal associations, enabling scalable and interpretable models.

Benefits of technology

The solution simplifies and speeds up the workflow, ensuring accuracy, reliability, and interpretability, allowing for human-interpretable domain knowledge and interventions, particularly in high-stakes settings like healthcare and finance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025050252_16042026_PF_FP_ABST
    Figure US2025050252_16042026_PF_FP_ABST
Patent Text Reader

Abstract

The invention is systems and methods directed to a method of training a machine learning model that is particularly adept at handling numeric data and is able to learn from prior models trained using the same method.
Need to check novelty before this filing date? Find Prior Art

Description

WOOD WIDE MODELS FOR IMPROVED MACHINE LEARNING PIPELINECROSS-REFERENCE TO RELATED APPLICATION

[0001] This application claims priority to U.S. Provisional Application Serial No. 63 / 705,274, filed on October 9, 2024, which application is incorporated by reference herein in its entirety.BACKGROUND OF THE INVENTIONField of the Invention

[0002] The field of the invention is methods and systems for improved machine learning pipelines.Federal Sponsorship

[0003] This invention was made with United States government support under 2211907 awarded by the National Science Foundation and FA8750-23-2-1015 awarded by the U.S. Air Force. The U.S. government has certain rights in the invention.Background of the Invention

[0004] Current generative artificial intelligence (“Al”) foundation models are general-purpose monolithic models trained on a broad set of data that are then fine-tuned to specific domains. This is natural for text data, where what a word means and its relationship with other words is largely preserved across text datasets. Other domains where this makes sense include images and audio. However, such monolithic foundation models are ill-suited to numeric tabular or time-series data: time-series data about stock prices has little relevance to a time series on the vital signs of an intensive care unit patient. This is in large part responsible for the poor numeric reasoning abilities of modem generative AL

[0005] Numeric data, rather than just text data, plays a crucial role in business operations and decision-making. With traditional approaches, machine learning (“ML”) engineers often need to combine techniques and extensive experimentation to find what works for each unique tabular dataset. These approaches are time and resource-intensive to implement, costly to maintain, and are not transferable or scalable. Corporations have also invested significant amounts of money in testing out foundation model-based solutions, but these are ill-suited to context-specific numeric data.SUMMARY OF THE INVENTION

[0006] A first embodiment of the present invention is a computer-implemented machine learning model training method for a machine learning pipeline, as applied to a computer system having a memory coupled to one or more processors and storing a plurality of programs to be executed by the processors. The computer system also has an input device and an output device. The computer system is programed to execute the method comprising iterative training, by the programmed computer system, wherein the iterative training comprises the following steps: (i) obtaining domain data; (ii) extracting identifiable neural vectors from the domain data; (iii) extracting conceptual interpretable factors from the identifiable vectors; (iv) applying a structural causal model for extraction of causal associations among the conceptual features to output causal connectors; and (v) employing a shared prior probability distribution over the connectors.

[0007] A second embodiment of the present invention that builds upon the first embodiment and also comprises that the step of extracting identifiable neural vectors from the domain data requires that the identifiable neural vectors having the same underlying model parameters are the same for vectors that are the same and that vector construction is automatic.

[0008] A third embodiment builds upon the first embodiment and also comprises that the conceptual interpretable factors have stable functions and support causal interventions.

[0009] A fourth embodiment builds upon the first embodiment, wherein applying a structural causal model further comprises extracting causal and statistical associations among the conceptual features in a manner that is scalable to settings comprising high-dimensional raw inputs and conceptual feature spaces. A fifth embodiment builds upon the fourth embodiment, wherein the structural causal model is represented as a plurality of local mechanisms, wherein each of the mechanisms comprises a conditional distribution of the conceptual feature(s) given a subset of related features. A sixth embodiment builds upon the fifth embodiment, further comprising a prior distribution defined over the conditional distributions. A seventh embodiment builds upon the sixth embodiment and further comprises: (i) training the structural causal model and regularizing via a Kullback-Leibler divergence between the structural causal model and a model prior; and (ii) updating the model prior using Bayesian inference conditioned on a structural casual model obtained from a different computer-implemented machine learning model.

[0010] An eighth embodiment is a computing device comprising: (i) one or more processors; (ii) memory coupled to the one or more processors; and (iii) a plurality of programs stored in the memory that, when executed by the one or more processors, cause the computing device to perform a plurality of iterative training operations comprising: (a) obtaining domain data; (b) extracting identifiable neural vectors from the domain data; (c) extracting conceptual interpretable factors from the identifiable vectors; (d) applying a structural causal model for extraction of causal associations among the conceptual features to output causal connectors; and (e) employing a shared prior probability distribution over the connectors.

[0011] A ninth embodiment builds upon the eight embodiment, wherein the step of extracting identifiable neural vectors from the domain data requires that the identifiable neural vectors having the same underlying model parameters are the same for vectors that are the same and that vector construction is automatic.

[0012] A tenth embodiment builds upon the eighth embodiment, wherein the conceptual interpretable factors have stable functions and support causal interventions.

[0013] An eleventh embodiment builds upon the eighth embodiment, wherein applying a structural causal model further comprises extracting causal and statistical associations among the conceptual features in a manner that is scalable to settings comprising high-dimensional raw inputs and conceptual feature spaces. A twelfth embodiment builds upon the eleventh embodiment, wherein the structural causal model is represented as a plurality of local mechanisms, wherein each of the mechanisms comprises a conditional distribution of the conceptual feature given a subset of related features. A thirteenth embodiment builds upon the twelfth embodiment, further comprising a prior distribution defined over the conditional distributions. A fourteenth embodiment builds upon the thirteenth embodiment, further comprising: (i) training the structural causal model and regularizing via a Kullback-Leibler divergence between the structural causal model and a model prior; and (ii) updating the model prior using Bayesian inference conditioned on a structural casual model obtained from a different computer-implemented machine learning model.

[0014] A fifteenth embodiment is a non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of an electronic device having a display, cause the electronic device to implement the iterative training of a machine learning pipeline by: (i) obtaining domaindata; (ii) extracting identifiable neural vectors from the domain data; (iii) extracting conceptual interpretable factors from the identifiable vectors; (iv) applying a structural causal model for extraction of causal associations among the conceptual features to output causal connectors; and (v) employing a shared prior probability distribution over the connectors.

[0015] A sixteenth embodiment builds upon the fifteenth embodiment, wherein the step of extracting identifiable neural vectors from the domain data requires that the identifiable neural vectors having the same underlying model parameters are the same for vectors that are the same and that vector construction is automatic.

[0016] A seventeenth embodiment builds upon the fifteenth embodiment, wherein the conceptual interpretable factors have stable functions and support causal interventions.

[0017] An eighteenth embodiment builds upon the fifteenth embodiment, wherein applying a structural causal model further comprises extracting causal and statistical associations among the conceptual features in a manner that is scalable to settings comprising highdimensional raw inputs and conceptual feature spaces.

[0018] A nineteenth embodiment builds upon the eighteenth embodiment, wherein the structural causal model is represented as a plurality of local mechanisms, wherein each of the mechanisms comprises a conditional distribution of the conceptual feature given a subset of related features. A twentieth embodiment builds upon the eighteenth embodiment, further comprising a prior distribution defined over the conditional distributions.BRIEF DESCRIPTION OF THE SEVERAL VIEWS OF THE DRAWINGS

[0019] To facilitate understanding of the invention, the drawings included within this application and associated descriptions illustrate certain preferred embodiments thereof, fromwhich the invention, various embodiments of its structures, construction and method of operation, and many advantages may be understood and appreciated. The drawings are incorporated by reference.

[0020] Figure 1 illustrates a traditional machine learning pipeline;

[0021] Figure 2 illustrates a deep neural ML pipeline;

[0022] Figure 3 illustrates a foundational model ML pipeline;

[0023] Figure 4 illustrates one embodiment of a Wood Wide ML pipeline of the present invention;

[0024] Figure 5 illustrates one embodiment of a first rung;

[0025] Figure 6 illustrates one embodiment of a second rung;

[0026] Figure 7 illustrates one embodiment of a third rung;

[0027] Figure 8 illustrates one embodiment of a fourth rung; and

[0028] Figure 9 shows one embodiment of a computer system or computing device implementing the method of the present invention.DETAILED DESCRIPTION OF THE INVENTION

[0029] The following describes example embodiments in which the present invention may be practiced. This invention, however, may be embodied in many ways, and the description provided herein should not be construed as limiting. Among other things, the following invention may be embodied as systems, methods, or devices. The following detailed descriptions should not be taken in a limiting sense. The accompanying drawings are hereby incorporated by reference.

[0030] The phrases “in some embodiments,” “in one embodiment,” “in various embodiments,” “according to various embodiments,” “in the embodiments shown,” “in other embodiments,” and the like generally mean the particular feature, structure, or characteristicfollowing the phrase is included in at least one embodiment of the present invention, and may be included in more than one embodiment of the present invention. In addition, such phrases do not necessarily refer to the same or different embodiments.

[0031] If the specification states a component, element, part, or feature “may,” “can,” “could,” or “might” be included or have a characteristic, that particular component or feature is not required to be included or have the characteristic.

[0032] In this document, the terms “a” or “an” are used, as is common in patent documents, to include one or more than one. In this document, the term “or” is used to refer to a nonexclusive “or ” such that “A or B” includes “A but not B,” “B but not A,” and “A and B,” unless otherwise indicated. Furthermore, all publications, patents, and patent documents referred to in this document are incorporated by reference herein, as though individually incorporated by reference. In the event of inconsistent usages between this document and those so incorporated by reference, the usage in the incorporated reference(s) should be considered supplementary to that of this document; for irreconcilable inconsistencies, the usage in this document controls. Before explaining at least one embodiment of the disclosure in detail, it is to be understood that the disclosure is not limited in its application to the details of construction, experiments, exemplary data, and / or the arrangement of the components set forth in the following description or illustrated in the drawings unless otherwise noted. The disclosure is capable of other embodiments or of being practiced or carried out in various ways. Also, it is to be understood that the phraseology and terminology employed herein is for purposes of description and should not be regarded as limiting.

[0033] As used in the description herein, the terms “comprises,” “comprising,” “includes,” “including,” “has,” “having,” or any other variations thereof, are intended to cover a nonexclusive inclusion. For example, unless otherwise noted, a process, method, article, or apparatus that comprisesa list of elements is not necessarily limited to only those elements but may also include other elements not expressly listed or inherent to such process, method, article, or apparatus.

[0034] Further, unless expressly stated to the contrary, “or” refers to an inclusive and not to an exclusive “or”. For example, a condition A or B is satisfied by one of the following: A is true (or present) and B is false (or not present), A is false (or not present) and B is true (or present), and both A and B are true (or present).

[0035] In addition, use of the “a” or “an” are employed to describe elements and components of the embodiments herein. This is done merely for convenience and to give a general sense of the inventive concept. This description should be read to include one or more, and the singular also includes the plural unless it is obvious that it is meant otherwise. Further, use of the term “plurality” is meant to convey “more than one” unless expressly stated to the contrary.

[0036] As used herein, qualifiers like “substantially,” “about,” “approximately,” and combinations and variations thereof, are intended to include not only the exact amount or value that they qualify, but also some slight deviations therefrom, which may be due to computing tolerances, computing error, manufacturing tolerances, measurement error, wear and tear, stresses exerted on various parts, and combinations thereof, for example.

[0037] As used herein, any reference to “one embodiment,” “an embodiment,” “some embodiments,”“one example,” “for example,” “various embodiments,” or “an example” means that a particular element, feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment and may be used in conjunction with other embodiments. The appearance of the phrase “in some embodiments” or “one example” in various places in the specification is not necessarily all referring to the same embodiment, for example.

[0038] The use of ordinal number terminology (i.e., “first”, “second”, “third”, “fourth”, etc.) is solely for the purpose of differentiating between two or more items and, unless explicitly stated otherwise, is not meant to imply any sequence or order of importance to one item over another.

[0039] The use of the term “at least one” or “one or more” will be understood to include one as well as any quantity more than one. In addition, the use of the phrase “at least one of X, Y, and Z” will be understood to include X alone, Y alone, and Z alone, as well as any combination of X, Y, and Z.

[0040] Where a range of numerical values is recited or established herein, the range includes the endpoints thereof and all the individual integers and fractions within the range, and also includes each of the narrower ranges therein formed by all the various possible combinations of those endpoints and internal integers and fractions to form subgroups of the larger group of values within the stated range to the same extent as if each of those narrower ranges was explicitly recited. Where a range of numerical values is stated herein as being greater than a stated value, the range is nevertheless finite and is bounded on its upper end by a value that is operable within the context of the invention as described herein. Where a range of numerical values is stated herein as being less than a stated value, the range is nevertheless bounded on its lower end by a non-zero value. It is not intended that the scope of the invention be limited to the specific values recited when defining a range. All ranges are inclusive and combinable.

[0041] The following description has set forth aspects of computer system 2000 or computer- implemented devices and / or processes via the use of block diagrams, flowcharts, and / or examples, which may contain one or more functions and / or operations. As used herein, the term or graphic of a “block” in the block diagrams and flowcharts refers to a step of a computer-implemented process executed by a computer system 2000, which may be implemented as a machine learning system or an assembly of machine learning systems. Each block can be implemented as either a machine learning system or as a nonmachine learning system, according to the function describedin association with each particular block. Furthermore, each block can refer to one of multiple steps of a process embodied by computer-implemented instructions or programs executed by a computer system 2000 (which may include, in whole or in part, a machine learning system) or an individual computer system 2000 (which may include, e.g., a machine learning system) executing the described step, which is in turn connected with other computer systems (which may include, e.g. , additional machine learning systems) for executing the overarching process described in connection with each figure or figures.

[0042] Circuitry, as needed herein to connect components (as will be known to one skilled in the art), may be analog and / or digital components, or one or more suitably programmed processors (e.g., microprocessors) and associated hardware and software, or hardwired logic. The term “processor 2100” as used herein means a single processor 2100 or multiple processors 2100 working independently or together to collectively perform a task or a functional unit that interprets and executes instruction data. Also, “components” may perform one or more functions. The term “processing component,” refers to a central processing unit that can include hardware, such as a processor 2100 (e.g., microprocessor), an application specific integrated circuit (“ASIC”), a field programmable gate array (“FPGA”), a combination of hardware and software, software, and / or the like. A processing component comprises the hardware and software configured to perform or execute the models, methods, and process of the present invention including performing systematic operations upon data or information exemplified by functions such as data or information transferring, merging, sorting, and computing (e.g., arithmetic operations or logical operations).

[0043] Software may include one or more computer readable instruction that when executed by one or more component, e.g., a processor 2100, causes the component to perform a specified function, ft should be understood that the algorithms described herein may be stored on one or more non-transitorycomputer-readable medium. Exemplary non-transitory computer-readable media (all examples of memory 2200) may include a non-volatile memory, a random access memory (“RAM”), a read only memory (“ROM”), a CD-ROM, a hard drive, a solid-state drive, a flash drive, a memory card, a DVD- ROM, a Blu-ray Disk, a laser disk, a magnetic disk, an optical drive, combinations thereof, and / or the like. Such non-transitory computer-readable media may be electrically based, optically based, magnetically based, resistive based, and / or the like.

[0044] As used herein, the terms “network-based,” “cloud-based,” and any variations thereof, are intended to include the provision of configurable computational resources on demand via interfacing with a computer and / or computer network, with software and / or data at least partially located on a computer and / or computer network.

[0045] The various embodiments of the present invention may include one or more input device (hereinafter “input device”), one or more output device (hereinafter “output device 2400”), one or more processors 2100 or processing component 2100 (used interchangeably), one or more communication device (hereinafter “communication device”) capable of interfacing with a network, and one or more memory (hereinafter “memory 2200”) storing processor-executable code and / or application(s) (hereinafter “application ”).

[0046] An input device, the output device, the processing component, the communication device, and the memory may be connected via a path such as a data bus that permits communication among the components of the classification device.

[0047] The input device may be capable of receiving information input from a user and / or the processing component and transmitting such information to other components and / or a network. The input device may include, but is not limited to, implementation as a keyboard, a touchscreen, a mouse, a trackball, a microphone, a camera, a fingerprint reader, an infrared port, an optical port, a cell phone,a smart phone, a PDA, a remote control, a fax machine, a wearable communication device, a network interface, combinations thereof, and / or the like, for example.

[0048] The output device 2400 may be capable of outputting information in a form perceivable by the processing component. Implementations of the output device 2400 may include, but are not limited to , a computer monitor, a screen, a touchscreen, a speaker, a website, a television set, a smart phone, a PDA, a cell phone, a fax machine, a printer, a laptop computer, a haptic feedback generator, an olfactory generator, combinations thereof, and the like, for example. It is to be understood that in some exemplary embodiments, the input device and the output device may be implemented as a single device, such as, for example, a touchscreen of a computer, a tablet, or a smartphone. It is to be further understood that as used herein the term user e.g., the user) is not limited to a human being, and may comprise a computer, a server, a website, a processor, a network interface, a user terminal, a virtual computer, combinations thereof, and / or the like, for example. The output device may display a user interface. The output device may feed into another computer system or processor.

[0049] The processing component or processor 2100 may be implemented as a single processor or multiple processors working together, or independently, to execute the application as described herein. It is to be understood, that in certain embodiments using more than one processing component, the processing components may be located remotely from one another, located in the same location, or comprising a unitary multi-core processor, or a combination thereof. The processing component may be capable of reading and / or executing processor-executable code and / or capable of creating, manipulating, retrieving, altering, and / or storing data structures into the memory such as in a database. The processing component may be capable of communicating with the memory via the path (e.g., the data bus). The processing component may be capable of communicating with the input device and / orthe output device communicably coupled, or otherwise connected, to the classification device of the classification system.

[0050] The processing component or processor 2100 may be further capable of interfacing and / or communicating with a server system via the network using the communication device. For example, the processing component may be capable of communicating via the network by exchanging signals (e.g., analog, digital, optical, and / or the like) via one or more port (e.g., physical ports or virtual ports) using a network protocol to provide updated information to the application or the user interface. In one embodiment, the server system is another embodiment of the classification device, however, the server system may be constructed, for example, as one or more server having a plurality of CPUs, GPUs, NPUs, TPUs, and / or the like, or a combination thereof. The server system may thus have a processing power available to both execute, or run, an Al model, as well as train, fine-tune, pre-train, instruction-tune, and / or align the Al model. The server system may be specially designed to handle large-scale datasets efficiently.

[0051] In one implementation, the processing component may be operable to receive the electrical signals from an artificial intelligence (“Al”) processor (a type of processor 2100.) The Al processor may be constructed in accordance with the processing component, for example, and, in some embodiments, may be incorporated into the processing component. In some embodiments, the Al processor may be separate from the processing component but may work together with the processing component to execute the application and / or access the memory. In one embodiment, the Al processor may operate at the request of, or be instructed to execute code by, the processing component.

[0052] Exemplary implementations of the processing component or processor 2100 also may include, but are not limited to, a digital signal processor (“DSP”), a central processing unit (“CPU”), a graphical processing unit (“GPU”), a neural processing unit (“NPU”), a tensor processing unit (“TPU”),a field programmable gate array (“FPGA”), a microprocessor, a multi-core processor, an application specific integrated circuit (“ASIC”), combinations thereof, and / or the like, for example. The processing component may include one or more processing component, having the same or different implementations, working together, or independently, and located locally, or remotely, e.g., accessible via the network such as located in the server system, and may include a multi-core, multi-processor component. As such, the application may be considered a cloud-based application, enabling access to powerful computing resources of the server system and simplifying user experience via the user interface. This implementation as a cloud-based application also drastically reduces processing time, as CUDA and tensor cores (e.g., the Al processors) allow the processing component to perform matrix multiplication at much faster rates.

[0053] In one implementation, the memory 2200 may be one or more non-transitory processor- readable medium. The memory may store processor-executable instructions, such as the application, that, when executed by the processing component, causes the processing component of the classification device to perform an action such as communicate with or control one or more component of the classification device and the classification system and / or to perform one or more process such as the classification system. The memory may be one or more memory working together, or independently, to store processorexecutable code and may be located locally or remotely, e.g., accessible via the network.

[0054] The memory 2200 may be physical memory or implemented as a “cloud” non-transitory processor-readable medium i.e., the one or more memory may be partially or completely based on or accessed using the network). The memory may store processor-executable code and / or information comprising the database and the application. In some embodiments, the application may be stored as a compiled application file, such as an executable file, for example, or in a structure (or unstructured)format, such as, e.g., in a non-compiled file. The application may be stored in a computer-readable format, and may, in some embodiments, further be stored in a human-readable format.

[0055] In some implementations, a database may be a time-series database, a relational database, a vector database, or a non-relational database. Examples of such databases include DB2®, Microsoft® Access, Microsoft® SQL Server, Oracle®, mySQL, PostgreSQL, MongoDB, Apache Cassandra, InfluxDB, Prometheus, Redis, Elasticsearch, TimescaleDB, Chroma, Pinecone, Weaviate, and / or the like. It should be understood that these examples have been provided for the purposes of illustration only and should not be construed as limiting the presently disclosed inventive concepts. The database 30 may be centralized or distributed across multiple systems.

[0056] In one embodiment, the database may be a centralized database with a distributed backup database, a distributed database with a centralized backup database, a distributed database with a distributed backup database, or a centralized database with a centralized backup database. In one embodiment, the database abides by, or exceeds, the 3-2-1 backup best practices. In one embodiment, each backup database is maintained as a real-time backup database, e.g., the backup database may be a mirror of the database.

[0057] In some implementations, the various embodiments of the present invention may include, but is not limited to, implementations as a personal computer, a cellular telephone, a smart phone, a network-capable television set, a tablet, a laptop computer, a desktop computer, a network-capable handheld device, a server, a digital video recorder, a wearable network-capable device, a virtual reality / augmented reality device, and / or the like.

[0058] In one implementation, the network may permit bi-directional communication of information and / or data between the server system and / or the classification device of the classification system. The network may interface with the classification device and / or the server system in a varietyof ways. For example, in some embodiments, the network may interface by optical and / or electronic interfaces, and / or may use a plurality of network topographies and / or protocols including, but not limited to, Ethernet, TCP / IP, circuit switched path, combinations thereof, and / or the like, as described above.

[0059] In some embodiments, the network may be the Internet and / or other network. For example, if the network is the Internet, the classification device may interact with the server system via the user interface implemented on the output device and / or the input device, such as a series of web pages or private internal web pages of a company or corporation, which may be written in hypertext markup language (HTML / PHP) and may utilize one or more suitable framework (such as JavaScript, Python, Flask, Django, and / or the like), for example. It should be noted that the user interface of the classification device may be another type of interface including, but not limited to, a Windows®- based application, a tablet-based application, a mobile web interface, an application running on a mobile device, a virtual-reality interface, an augmented-reality interface, and / or the like.

[0060] The network may be almost any type of network. For example, in some embodiments, the network may be a version of an Internet network e.g., exist in a TCP / IP-based network). In one embodiment, the network is the Internet. It should be noted, however, that the network may be almost any type of wireless network and may be implemented as the World Wide Web (or Internet), a local area network (“LAN”), a wide area network (“WAN”), a low power wide area network “LPWAN”, a LoRa network (e.g., “LoRaWAN”), a metropolitan network, a wireless network, wireless networking technology a “WiFi network”, a cellular network, a Bluetooth network, a Global System for Mobile Communications (“GSM”) network, a code division multiple access (“CDMA”) network, a 3G network, a 4G network, a long term evolution (“LTE”) network, a 5G network, a satellite network, a radio network, an optical network, a shortwave wireless network, a long-wave wireless network, combinations thereof,and / or the like. It is conceivable that in the near future, embodiments of the present disclosure may use more advanced networking topologies.

[0061] While the disclosure has been described in detail and refers to specific embodiments, it will be apparent to one skilled in the art that various changes and modifications can be made without departing from the spirit and scope of the embodiments. Thus, it is intended that the present disclosure covers the modifications and variations of this disclosure, provided they come within the scope of the appended claims and their equivalents.

[0062] It is to be understood that the invention may assume alternative variations and step sequences, unless specified to the contrary. It also is to be understood that the specific devices and processes illustrated in the attached drawings and described in this specification are simply exemplary embodiments of the invention. Hence, specific dimensions and other physical characteristics related to the embodiments disclosed are not to be limiting.

[0063] It should be understood that the invention may assume alternative variations and step sequences unless specified to the contrary. The specific devices and processes illustrated in the attached drawings and described in this specification should also be understood as exemplary embodiments of the invention. Hence, specific dimensions and other physical characteristics related to the embodiments disclosed are not to be limiting.

[0064] As further background, reference is made to Figure 1, which depicts a traditional machine learning pipeline 100. A machine learning pipeline 100 is a series of interconnected data processing and modeling steps designed to automate, standardize, and streamline the process of building, training, evaluating, and deploying machine learning models.

[0065] Traditional machine learning pipelines 100, as shown in Figure 1, have multiple stages. The first is a feature engineering stage 120 that converts raw data 110 to what are known as “feature(s)130.” As used herein, the term “feature(s) 130” refers broadly to any derived data element, variable, attribute, signal, or representation obtained or constructed from raw input data 110. The purpose of transforming raw data 110 to features 130 is to make them more amenable to being ingested by machine learning models 100. This is typically very labor intensive, requiring iterations through the entire pipeline, as well as early exploratory data analyses, in order to find suitable features. The next stage is model training 140 that generates an ML model 150. The inference stage 160 then generates model outputs 170, which if deemed unsuitable 175, necessitates the data scientist to revisit the feature engineering stage 120. The last analysis stage 180 takes the model outputs 170 to generate data intelligence 190. The key caveats of this traditional ML pipeline 100 are that feature engineering 120 is mostly manual and frequently requires trial and error iteration and rework. This is time consuming and an inefficient use of high cost data scientist expertise. It is also ad hoc, which results in less performant model outputs 170.

[0066] Deep neural ML pipelines 200, as shown in Figure 2, gained popularity in the mid- 2010s. These 200 were able to obviate the feature engineering stage 120 and directly fit models on raw data 110. But, feature engineering 120 was replaced with model architecture engineering 220 to create model architecture 230 and model training engineering 240 to create a deep learning ML model 250. Again, if the inference stage 160 generates model outputs 170 deemed unsuitable 275, this necessitates the machine learning engineer to revisit the model architecture engineering 220, the model architecture 230, and the training engineering 240 stages. Model engineering is iterative and often requires manual effort on the part of machine learning engineers, which is both time and cost expensive. Further, the deep neural models 200 have known reliability and interpretability issues, which prevent their direct usage in high-stakes settings.

[0067] Foundation model ML pipelines 300, as shown in Figure 3, are even more recent, and gained traction from 2018 onwards. These 300 not only eliminated the feature engineering stage 120 but also the model engineering stage 220. Foundation models 300 are monolithic large- scale models that are trained on a broad set of foundational data 305. The models 300 undergo pretraining 315. Then given a specific domain or task, these monolithic foundation models 325 are “fine-tuned” 335 i.e. slightly modified given data from the specific domain 110.

[0068] This foundation plus fine-tuned modeling framework 300 has been remarkably successful in domains such as images and text where the broad set of data has information that can be leveraged towards the specific domain. However, they 300 are ill-suited in settings where the data is inherently heterogeneous, and the assumption that there is a broad set of data that are all related to any specific domain breaks down. Prototypical instances of such heterogeneous data are numeric tabular data, and numeric time-series data. Numeric times series of stock prices have little to do with time series comprising the vital signs of a patient. In such cases, training a single monolithic foundation model over a large collection of such datasets is not meaningful. There is yet another caveat to foundation models 300: it is difficult to guarantee quality, safety, and legality of the provenance of foundation data 305. This is particularly relevant given possible targeted data poisoning attacks. The third caveat with foundation models 300 is that due to their scale, they make particular architectural choices that are more scalable for training on large datasets. In particular, in the key architecture underlying many modem foundation models 300, the transformer is inherently designed for discrete data such as text - which poses challenges when working with numeric data.

[0069] To address these shortcomings, various embodiments of systems and methods of the present invention comprises a new paradigm or architecture 500, which in some instances is referred to interchangeably herein as a “Wood Wide Process 500” and, when incorporated into anoverall ML pipeline, as a “Wood Wide models 1000” “Wood Wide paradigms 1000,” or “Wood Wide ML pipelines 1000,” (one example is illustrated in Figure 4). The specifications for the four rungs of one embodiment of the Wood Wide process 500 (which can be incorporated into a Wood Wide Model ML pipeline 1000) enable functionalities not possible with the earlier traditional ML 100, deep ML 200, and foundation ML 300 pipelines. Concrete implementations of each of the rungs are detailed, satisfying all their requirements, and using existing machine learning technologies.

[0070] As mentioned previously, various embodiments of the Wood Wide ML pipelines 1000 of the present invention are distinct from the recent foundation model pipelines 300, which are not well suited to discrete data, and data with a smaller number of samples, and that do not have a broader base of relevant data. In contrast, the approach of the present invention is applicable to heterogeneous data domains where there is a lot of numeric data, including the following nonlimiting examples:• tabular data, ranging over an entire set of (implicitly linked) relational tables within an enterprise relational database;• numeric time series; and• multimodal data domains with tables, time series and text, such as electronic health records, sales reports, stock analyst reports, or stock earnings transcripts.

[0071] In one embodiment, Wood Wide ML pipelines 1000 of the present invention can be used for numeric robust data decision support. For decision support in most modem enterprise settings, models are needed that can summarize a combination of numeric and text data e.g. numeric vital signs and doctors’ notes in the context of electronic health records, or sales or analystreports where text is also accompanied by tables and contextual numbers. These are ill-suited to modern foundation models 300 not only because of the heterogeneous data in modem enterprises, but also because particular classes of foundation models 300 such as large language models (“LLMs”) are not suited to reasoning with numbers. Further, when suggesting decisions to take, purely associational models are not robust enough to provide useful guidance: once the decisions are taken, this often causes a shift in the distribution of the data away from what the model was trained on. The classes of causal factors and models of the present invention provide guidance that is structurally robust to such interventions. Further, the classes can provide insights into decision outcomes in counterfactual settings, that is, being able to answer “what if’ questions.

[0072] As discussed previously, the three main ML pipelines (traditional ML 100, deep neural ML 200, and foundation ML 300) each have very critical disadvantages, so that the choice of which pipeline to choose is often determined by which disadvantages are less onerous for the specific application. For instance, for numeric tabular data, with a smaller number of samples, the disadvantages of foundation ML 300 and deep neural ML 200 are considered to be far greater than that of traditional ML 100, which is thus the most popular choice for such data, even though traditional ML 100 in turn has its disadvantages as outlined earlier.

[0073] There is a need for a pipeline that does not incur the disadvantages of these existing pipelines outlined in the earlier section. Mor specifically, the following needs (desiderata) exist in the field suggesting the creation of an ML pipeline that:• avoids iterative feature engineering (unlike traditional ML pipelines 100);• avoids model architecture and training engineering (unlike deep ML pipelines 200);• suitable to both continuous and discrete data, data domains with a smaller number of samples, and data domains that do not have a broader base of relevant data (unlike foundation ML pipelines 300);• ensures reliability, and interpretability of model outputs (unlike deep 200 and foundation 300 ML pipelines);• has the ability to incorporate human-interpretable domain knowledge as inputs to the models (unlike any of the pipelines (100, 200, or 300)); and• has the ability to incorporate changes and interventions in human-interpretable aspects of the input (unlike any of the pipelines (100, 200, or 300)).

[0074] The overall innovation of the present invention are systems and methods that comprise an embodiment of a new Wood Wide ML pipeline 1000 that contrasts from earlier pipelines in that it addresses all six desiderata outlined above. The embodiments of this Wood Wide ML pipeline 1000 remove the iterative and manual aspects of both feature 120 and model architecture 220 engineering, simplifying and speeding up the workflow, while guaranteeing accuracy, reliability, and providing interpretability. Various embodiment of the present invention Wood Wide ML pipeline 1000 also are applicable to general data types (continuous numeric data in particular), as well as data domains. Lastly, unlike any of the other pipelines, various embodiments of the Wood Wide ML pipeline 1000 have the ability to ingest human-interpretable domain knowledge (which can be considered as an inversion of interpretability, where users wish to extract human-interpretable knowledge from the model). Moreover, unlike any of the other pipelines, various embodiments of the Wood Wide ML pipeline 1000 are amenable to human-interpretable interventions to the data. For instance, for a credit decision Al system, a regulator might wish to understand the consequences to model outputs of changing some loan profiles with sometargeted interventions. Or a doctor might wish to understand how a model might behave in response to a change in operating room protocols. Such changes to the data, while human-interpretable, cause shifts in the distribution of the data that make the model outputs from all these existing pipelines unreliable. These last two desiderata are thus key to use in high-stakes settings, such as healthcare, finance, and defense, among many others. As detailed in the following discussion, there are fundamental reasons why these existing pipelines cannot satisfy all these desiderata. An innovation of the present invention is in setting up a hierarchical architecture with multiple rungs, each with specific requirements that together do satisfy all these desiderata. Moreover, concrete implementations of each of these rungs is provided.

[0075] To explain the nomenclature used herein, the overall architecture of the various embodiments of the Wood Wide Process 500 and the Wood Wide ML pipeline 1000 is inspired by the “wood wide web,” which is a phrase often used to describe an underground network of fungal threads that connect together many trees and plants. These trees are all independent of each other and yet share nutrients via the wood wide web. This stands in contrast to a large, shared concrete foundation on top of which one might build specialized buildings. Analogous to the latter concrete foundation, modem Al paradigms hinge on large foundation models upon which one might build specialized models. However, they are ill-suited in settings where the data is inherently heterogeneous and the assumption that there is a broad set of data related to any specific domain breaks down. Prototypical instances of such heterogeneous data are numeric tabular data and numeric time-series data. Numeric times series of stock prices have little to do with time series comprising the vital signs of an intensive care unit (“ICU”) patient. In such cases, training a single monolithic foundation model over an extensive collection of such datasets is notmeaningful. For such heterogeneous settings, the various embodiments of the Wood Wide ML pipeline 1000 fill the existing void.

[0076] In contrast, but analogous to the wood wide web, various embodiments of the Wood Wide ML pipeline 1000 comprise an architecture with many smaller wood wide ML models that are all independent and yet can share conceptual information with each other. To be able to share nutrients from the wood wide web, trees need a special root-based architecture that can connect to these fungal threads. Accordingly, to operationalize various embodiments of the Wood Wide ML pipeline 1000 a specific hierarchical four-runged architecture 500 is employed, which has specific requirements for each rung so that together they can address the desiderata described previously.

[0077] Various embodiments of the Wood Wide ML pipeline 1000 of the present invention relate to a computer-implemented enhanced machine learning pipeline method 1000 and system 2000, comprising a processor 2100 with memory 2200, that allow users to enable functionalities not possible with the earlier traditional ML 100, deep ML 200, and foundation ML 300 pipelines. Various embodiments of the Wood Wide ML pipeline 1000 comprise: (1) removal of the iterative and manual aspects of both feature 120 and model architecture engineering 220, simplifying and speeding up the workflow, while guaranteeing accuracy, reliability, and providing interpretability (while being applicable to general data types, continuous numeric data in particular, as well as data domains); (2) the ability to ingest human-interpretable domain knowledge (which can be considered as an inversion of interpretability, where one attempts to extract human-interpretable knowledge from the model),' and (3) amenability to human-interpretable interventions to the data. This invention further encompasses an embodiment of a system 2000 for implementation of the various embodiments of the Wood Wide ML pipeline 1000.

[0078] Building upon the analogy to the wood wide web, various embodiments of the present invention comprise an architecture with many smaller wood wide ML models that are all independent and yet share subtle elements with each other. To be able to share nutrients from the wood wide web, trees need a special root-based architecture that can connect to these fungal threads. Accordingly, to operationalize wood wide ML models, various embodiments of the present invention comprise a novel hierarchical neuro-symbolic architecture 500 - also referred to as “neuro-causal 500” - that synthesizes deep neural models 200 and causal graphical models 400, two powerful paradigms that have developed over the past two decades. The first rung 510 of the architecture 500 extracts identifiable neural features 528 (referred to herein interchangeably as (“identifiable neural vectors 528”) from raw input features 110. The role of the neural features 528 is to extract a compressed representation of the input 110 that synthesizes geometric and contextual knowledge of the domain. The second rung 540 of the architecture 500 extracts conceptual features 550 from these neural features 528. The role of these conceptual features 550 is to serve as disentangled semantic representations of the input 110 that can be grounded to human interpretable factors in the environment. The third rung 560 of the architecture 500 is the implementation of a structural causal model 562 over these conceptual factors 550 that synthesizes domain knowledge over these conceptual factors 550 and guides the lower-level conceptual and neural features 528. Conceptual factors 550 are the output of the second rung 540, which are features that are conceptual. Therefore, “conceptual factors 550,” “conceptual interpretable factors 550,” “conceptual interpretable features 550” and “conceptual features 550” are used interchangeably herein to describe this element and concept. At this structural causal layer 580, a web of wood wide models 1000 will be able to share conceptual ontologies and causal mechanisms among conceptual features 550.

[0079] The four rungs of one embodiment of a neuro-symbolic architecture 500 of one embodiment of a wood wide ML model 1000 or system 2000 are shown in Figure 4. For purposes of the various figures and discussion herein, the circles shown in Figures 4 through 8 indicate methods i.e. "vector extraction" refers to the methods that extract vectors), and the squares indicate the outputs i.e. "vectors"). The words “feature” or “features” simply refer to generic transformations of the data. “Vector” or “vectors” refer to features that satisfy particular requirements. “Factor” or “factors” refer to transformations of features that satisfy particular requirements. This means that factor(s) are also a type of feature (i.e. transformation of data). With that explanation, the four rungs of one embodiment of the architecture 500 shown in Figure 4 can be summarized as follows:First rung 510: this first rung 510 is where identifiable neural features 528 are extracted from given raw numeric data 110 and the output are identifiable vectors 528;Second rung 540: this second rung 540 is where disentangled conceptual interpretable features 550 are extracted from neural features 528;Third rung 560: this third rung 560 where casual relational models 562 (or structural causal models 562) are developed, over the conceptual features 550, which relational models synthesize and store high-level domain knowledge which are referred to as causal connectors 561 or casual connectors 561; andFourth Rung 580: this fourth rung 580 is where a shared prior 581 is developed over the connectors 561 of wood wide ML models 1000 (acting as mycelium 600.)

[0080] As illustrated in Figure 4, the novel hierarchical neuro-symbolic architecture 500 (or neuro-casual architecture 500) method / process can be incorporated into a wood wide machinelearning pipeline 1000 or computer device or system 2000 by using the architecture 500 as the training step 140 to generate an ML model 150. The inference stage 160 then generates model outputs 170, which are fed into the analysis stage 180 which uses the model outputs 170 to generate data intelligence 190.

[0081] As mentioned previously, due to the nature of the foundation model paradigm 300, foundation models 300 are suited to monolithic data domains where an aggregation of disparate datasets is meaningful. Moreover, the key architecture underlying many modem foundation models 300 - the transformer - is inherently designed for discrete data, which poses further challenges when working with numeric data. In contrast, the approach of the present invention methods 1000 and systems 2000 applies to heterogeneous data domains where there is a lot of numeric data.

[0082] Specifically, various embodiments of the present invention comprise methods 1000 and systems 2000 to train and perform inference (using the machine-learning models to answer questions over the data) via this novel machine-learning framework 500 of wood wide models for broad classes of heterogeneous data domains 110. One non- limiting example of how the present invention can be implemented is use for numeric robust data decision support. For decision support in most modem enterprise settings, models are needed that can summarize a combination of numeric and text data, e.g. , numeric vital signs and doctors’ notes in the context of electronic health records or sales or analyst reports where tables and contextual numbers also accompany the text. These are ill-suited to modem foundation models 300 because of the heterogeneous data 110 in modern enterprises and because particular classes of foundation models 300, such as large language models (“LLMs”), are not suited to reasoning with numbers. Further, when suggesting decisions to take, purely associational models are not robust enough to provide useful guidance.Once the decisions are taken, this often causes a shift in the distribution of the data away from what the model was trained on. The lack of identifiability means that two associational models that look identical on the training data may perform very differently when the decisions are taken. The Woodwide Al pipeline 1000 and system 2000 address all of these caveats. The vectors 528 of the first rung 510 are identifiable so that two models that behave similarly on training data 110 will also behave similarly on test data 110. The factors 550 of the second rung 540, and the causal connectors 561 over them in the third rung 560, provide structurally robust guidance to answering questions over the models in response to interventions that cause a shift in the distribution of the data away from what the model was trained on. Further, it can provide insights into decision outcomes in counterfactual settings, that is, being able to answer “what if’ questions. The mycelium prior 600 in the fourth rung 580 means that the model can learn these vectors 528, factors 550, connectors 561 with minimal data.

[0083] One embodiment is a computer-implemented machine learning model training method 500 for a machine learning pipeline 1000, as applied to a computing device or computer system 2000 having a memory 2200 coupled to one or more processors 2100 and storing a plurality of programs to be executed by the processors 2100. The computer system 2000 also having an input device and an output device 2400. The computer system 2000 programed to execute the method 500 comprising iterative training, by the programmed computer system 2000. For this embodiment, the iterative training comprises the following: (i) obtaining domain data; (ii) extracting identifiable neural vectors from the domain data; (iii) extracting conceptual interpretable factors from the identifiable vectors; (iv) applying a structural causal model for extraction of causal associations among the conceptual features to output causal connectors; and (v) employing a shared prior probability distribution over the connectors.

[0084] Another embodiment is a non-transitory computer-readable storage medium 2200 storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors 2100 of an electronic device having a display 2400, cause the electronic device to implement the iterative training of a machine learning pipeline 1000 by executing the following steps: (i) obtaining domain data 110; (ii) extracting identifiable neural vectors 510 from the domain data; (iii) extracting conceptual interpretable factors 540 from the identifiable vectors; (iv) applying a structural causal model 550 for extraction of causal associations among the conceptual features to output causal connectors; and (v) employing a shared prior probability distribution 580 over the connectors.

[0085] Overall Architecture. The four rungs of one embodiment of a hierarchical neuro- symbolic architecture 500, which can be incorporated into a wood wide model pipeline 1000 or system 2000 of the present invention, are discussed more fully below and illustrated in Figures 4 through 9.

[0086] First Runs 510 (Extractins Identifiable Vectors 528) (shown in Figures 4 and 5).

[0087] In one embodiment, the first rung 510 extracts identifiable neural features 528 (as outputs) given raw numeric data 110 (raw input features, e.g., table rows, time series, images) from heterogeneous data domains. Heterogeneous data domains are represented by the "input domain" box or “domain data 110.” Input domain is the space of all possible input data 110 in the setting under consideration (e.g. all possible patient records). Raw numeric data refers to the finite set of input data 110 that is provided or given (e.g. the set of patient records being analyzed.) Each input datum 110 consists of multiple data attributes. Heterogeneous data domains refers to the fact that these attributes can be very different (e.g. a patient record will have various kinds of blood tests,measurements such as blood pressure, etc.) This is in contrast to homogenous data where the data attributes are all similar (e.g. an image will have pixels as its attributes.) These features 528 can be viewed as vector-valued functions of the inputs 514 that map any given input 110 to an array of numbers, also referred to as vectors 528. In Figure 5, the function-space view of inputs 514 includes the input domain 512 (or data). Most modem Al models, such as deep neural networks 200, have an implicit feature learning component that maps raw inputs 110 to neural features 528 and a typically linear or some otherwise simple model on top of these features. There is, however, no explicit delineation between the two in most contexts, and indeed, stating the feature learning component as above is ill-posed, there are many possible feature functions for which a simple model on top might be performant. One embodiment of the present invention provides a rigorous framework for learning neural features 528, where it is formalized what is meant by good features and develop scalable algorithms to extract these features. These algorithms can vary and are based on known concepts for machine learning. In one embodiment of the present invention, there are two requirements for these neural features as illustrated in Figure 5.

[0088] Requirement 1 530: For one embodiment, the first requirement 530 is that the vectors 528 be identifiable (to avoid non-identifiability 532). This first requirement 530 means that if two feature functions are the same, the underlying model parameters are the same (see Figure 5). If the model class is not identifiable, then there will be many models all with same empirical performance: minor changes in the training setup (e.g. random seeds) could result in a different choice from this large set. If there are no guarantees translating empirical performance to deployment performance, then there would be variability in deployment performance among the equivalence class: some could be good, some could be very bad, and which could result in facets of random seeds. The prior art provides many examples of such deployment variabilitygiven non-identifiable models. The other caveat with non-identifiability is that if there is a large equivalence class that you cannot quantify, there is no hope of providing any guarantee of deployment performance: since one cannot even pinpoint what it is they are estimating. With modern large-scale neural models, the lack of identifiability is due in large part to their large scale: there are many possible models all with the same empirical performance on the training data. The Al community seems split between those who claimed identifiability was not necessary because of “good empirical performance” which misses the points made above, while others refer to “implicit inductive bias” that is perhaps learning identifiable models, but we don’t know how or why. In contrast, the wood wide Al pipeline provides an explicit requirement of identifiability of the neural features.

[0089] Requirement 2 534: For one embodiment of the present invention, the second requirement 534 is that there be an automatic construction of features similar to deep learning ML, but without the full model 536, and hence without the model architecture engineering (see Figure 5.) Modem deep neural networks have an implicit feature learning component that maps raw inputs 110 to neural features 528, and a typically simple model on top of these features. There is however no explicit delineation between the two in most contexts, and indeed, stating the feature learning component as above is ill-posed: there are many possible feature functions for which a simple model on top might be performant. In contrast, one embodiment of the present invention only requires that the feature construction be automatic 534, similar to deep ML, and also that it be identifiable 530, unlike deep ML.

[0090] There are many ways to satisfy both these requirements of automatically extracting identifiable neural features 528. One such is the class of self-supervised learning approaches 516, involving the specification of predict! ve / consistency tasks 518 to lead to identifiable conditions 520.The key insight is that neural features are finite summaries of typically infinite-dimensional function spaces over the input domain 514. For instance, if the input domain 110 comprises rows in a table 110, then a function space 514 would comprise some specification of numeric computations that can be performed given any particular row. Then, the task of feature learning reduces to learning or specifying these function spaces, and then extracting a summary, typically orthonormal dictionary 524 functions as finite summaries 522 of these spaces. Given a function space 514, there are standard “dictionary learning” tools 522 to extract finite-dimensional summaries 524 of these function spaces. The key question then is what the function spaces should be. The key idea in self-supervised learning is to specify functions that map parts or views of an input to the input itself. One embodiment of the present invention has been shown to yield identifiable neural features 528 (as outputs 526) and provide scalable reliable algorithms to do so. In one embodiment of the present invention, there are two implementations for automatically extracting identifiable features 528 without manual engineering by first specifying a principled function space 514. In one embodiment, this function space 514 is a spectrally transformed RKHS (Reproducing Kernel Hilbert Space) defined by kernels and unlabeled data, and according to the present invention it is augmentation-based self-supervised learning through an RKHS induced by invariances to augmentations. For additional background information the following publications are hereby incorporated by reference: (1) Runtian Zhai, Bingbin Liu, Andrej Risteski, Zico Kolter, and Pradeep Raviku-mar. Understanding augmentationbased self-supervised representation learning via rkhs approximation and regression. International Conference on Learning Representations (1CLR), 2024a; and (2) Runtian Zhai, Rattana Pukdee, Roger Jin, Maria-Florina Balcan, and Pradeep Ravikumar. Spectrally transformed kernel regression. arXiv preprint arXiv:2402.00645, 2024b. BothZ / zm 2024a and 2024b summarize theseinfinite-dimensional spaces via their leading eigenfunctions, yielding finite, identifiable features derived directly from the structure of the data.

[0091] Other embodiments of the present invention employ other spaces of functions 514 that matter in context, for example, those capturing similarities or invariances in the data 110, and then identifiably distill this space into a small set of core features, extracted automatically rather than hand-engineered, and any that satisfy the two requirements 530, 534 outlined above would qualify for the first rung 510 of the architecture 500 to output identifiable vectors 528.

[0092] Conceptual Interpretable Factors or Features 550 (as illustrated in Figures 4 and 6.)

[0093] In one embodiment of the present invention, the second rung 540 is focused on extracting conceptual features 550 from neural features 528. A key caveat of neural features 528 (including the present invention’s vectors 528) is that they are distributed in nature. This means that the salience of any relevant concept is distributed among all the neural features 528. On the one hand this allows for a compression storage of a lot of relevant concepts within a small set of neural features. But on the other hand, this has a key negative consequence of a lack of robustness: small changes to the compressed neural features are very meaningful, so that model outputs can be sensitive to small changes in model parameters. Correspondingly, small changes in the inputs can result in large changes in model outputs. Both of these result in a lack of reliability and robustness. Moreover, this sensitivity of distributed features to small changes in the input also typically entails that these features are not human-interpretable.

[0094] There is an importance of using identifiable neural features 528, specifically, if the neural features 528 are not identifiable, it would be difficult to learn to map them to conceptualfeatures 550 that are, in general, identifiable, often with a human-interpretable concept. Two further requirements are outlined below and illustrated in Figure 6.

[0095] Requirement 1 542: One embodiment of the present invention requires functions of the neural vectors 528 from the first rung that are stable 542, that is less sensitive to small changes in the input. Such stable features 542 are sometimes referred to as symbolic features, or herein as “conceptual factors 550” or “conceptual features 550,” which can be interpretable. Thus, in this second rung 540, the first requirement 542 entails the extraction of conceptual features 550 from neural features 528. This has long been a goal in Al to do so in a fully automated way. This requires some notion of what “conceptual” means. The prior art 544 is full of many different approaches which have aimed to “disentangle” the neural features into a set of conceptual features that are as far apart from each other as possible. However, this has been shown to not be identifiable, and moreover not always lead to useful features.

[0096] Requirement 2 546: One embodiment of the present invention further requires that these conceptual features 550 be those where causal interventions are meaningful 546. This requirement is imposed in order to satisfy the last two of the six desiderata of the pipeline. However, a crucial advantage of additionally imposing this second requirement 546 is that this also provides a very useful notion of what conceptual means: features are those that can be causally intervened upon (intervenable factors that are human interpretable 548.) This has the additional advantage of human interpretability 548, since one often only intervenes on factors that one conceptually understands. Thus, in this second rung 540, the extraction of causally relevant factors 550 from the first rung of neural feature vectors 528 is required.

[0097] For the present invention, the outputs of any approach that satisfy these requirements are referred to as “factors 550.” There are many recently proposed approaches 552that satisfy these requirements. One such is the class of causal representation learning approaches 554 that automatically extract such latent conceptual features given data that arise from multiple environments each of which corresponds to interventions on some unknown conceptual features. The key idea is that given multiple datasets each of which correspond to some interventions in the unobserved conceptual features, there is a signature of such interventions in the observed data and hence the neural features. The present invention makes this further scalable 556 by extracting the conceptual features 550 from neural features 528 rather than the raw inputs 110 themselves.

[0098] For various embodiments of the present invention, any approach that extracts causally relevant factors 550 from the first rung vectors 528 would qualify for the second - factors - rung 540 of the architecture 500. ft is instructive at this juncture to note the reason for imposing the requirement of identifiability requirements in the first rung 510 of the architecture 500: if the neural vectors 528 were not identifiable, it would be difficult to learn to map them to often human- interpretable conceptual features 550, which are by their very intervenability nature, also identifiable.

[0099] Third Rung 560 (Connectors 561)(illustrated in Figures 4 and 7).

[0100] For the last desideratum of the pipeline, one embodiment of the present invention requires the ability to incorporate changes and interventions in factors 550. This leads to the following requirements.

[0101] Requirement 1 566: One embodiment of the present invention requires the extraction of causal and statistical associations among the conceptual features 550 via a structural causal model 562 (“SCM 562”).

[0102] Requirement 2 579: A key additional requirement is that this is scalable to settings with high-dimensional raw inputs 110, and high-dimensional conceptual features 550.

[0103] The outputs of any approach that satisfies these requirements are referred to simply as “connectors 561.” There are many recently proposed approaches that satisfy these requirements. One embodiment is provided by Learning Linear Causal Representations from Interventions under General Nonlinear Mixing (Buchholz et al., 2024). This work does two things. It addresses the precursor task of extracting conceptual features 550 from high-dimensional raw inputs 110 via identifiable latent representations. Then, it leverages these representations within a structural causal model 562 to scalably recover causal and statistical associations 564 among the conceptual features 550. One embodiment of the present invention is enabled by a recent breakthrough in causal structure learning, where the combinatorial task of DAG discovery was reduced to a smooth continuous optimization problem 570 (see for background, Zheng, Xun, Chen Dan, Bryon Aragam, Pradeep Ravikumar, and Eric Xing. Learning sparse nonparametric dags. In International conference on artificial intelligence and statistics, pp. 3414-3425. PMLR, 2020; and Zheng, Xun, Bryon Aragam, Pradeep K. Ravikumar, and Eric P. Xing. Dags with no tears: Continuous optimization for structure learning. Advances in neural information processing systems 31 (2018).

[0104] By encoding the acyclicity constraint algebraically 574, this approach enables directly estimating structural causal models 562 (“SCMs”) that capture causal and statistical associations among conceptual features (Requirement 1 564) using scalable gradient-based solvers 578, by optimizing structure and parameters 576. Because the optimization 570 is continuous and amenable to standard numerical methods, the method further scales to high-dimensional raw inputs and conceptual feature spaces (Requirement 2 579), providing a practical path to causal discovery in complex domains.

[0105] To expand on the advantages of satisfying the last desideratum of the wood wide ML pipeline, such connectors answer interventional questions of what the model outputs will be in response to particular interventions in conceptual features. For instance, a shoe store chain might be interested in the consequences of intervening and changing the store layout of all their stores. Since this changes the data distribution, using a purely statistical association based model can lead to very misleading results. The connectors on the other hand will be able to provide relevant predictions due to their causal structure. Secondly, this also enables answering counterfactual questions of what the model outputs will be in response to a particular change in a specific input. For instance, a specific shoe store might be interested in how their sales would differ in the counterfactual setting where they changed the store layout. This is a subtlety different question and again requires causal structure to be able to answer accurately, which the connectors do.Fourth Rung 580 (Mycelium 600)(illustrated in Figures 4 and 8).

[0106] Recall the earlier discussion that disparate numeric data such as numeric time-series and numeric tables cannot share lower level features among themselves, so that it does not make sense to fit a single foundation model over all of them and therefore requires separate models for these. But one caveat with this is that this does not allow any sharing among even related wood wide models.

[0107] Requirement 1 598: One embodiment of the present invention requires that wood wide models 1000 be able to share information with each other.

[0108] Requirement 2 594: One embodiment of the present invention requires the ability to instantiate a wood wide ML model 1000 with very limited data 110.

[0109] These requirements together are important in settings with limited data, there might not be enough data to learn a wood wide model 1000 by itself, but it might be able to do so byreceiving shared information from other wood wide models 1000. Any system that satisfies these requirements is referred to as “mycelium 600” taking inspiration from the wood wide web. The present invention comprises systems and methods that satisfy both these requirements thanks to an important emergent ability that arises when wood wide models 1000 with the four rungs 510, 540, 560, 580 are fitted as outlined in earlier sections. Similar to the wood wide web that allows individual trees to pass nutrients among each other, the connector layer 580 allows separate wood wide models 1000 to pass associations among conceptual features 550 with each other. Thus, one system that will satisfy both these requirements is simply any scalable prior distribution over SCMs, that is available to and can draw from any wood wide model 1000 that wishes to participate in the wood wide web.

[0110] In one embodiment, a structural causal model 562 (“SCM 562”) is represented as a plurality of local mechanisms 584, each mechanism 584 comprising a conditional distribution 586 of a conceptual feature 550 given a subset of related features. The embodiment further comprises a prior distribution defined over said conditional distributions 588. To enable Requirement 2 (instantiation of models with limited data), the training of SCMs 562 in a connector layer 590 of a Woodwide model is regularized via a Kullback-Leibler (“KL”) divergence 592 between the connector model or layer 590 and the Woodwide prior 600. To enable Requirement 1 598 (information sharing across models), the prior 600 is updated using Bayesian inference conditioned on SCMs 596 obtained from other Woodwide models 1000. The key advantage of the fourth rung 580 of this embodiment wood wide ML pipeline 1000 is to provide the advantages of foundation ML pipelines 300, while extending them 300 to domains like numeric time series where a single monolithic model trained over a broad set of data is not meaningful. As outlined earlier, instead the sharing is at the higher level of connectors and factors.

[0111] While the disclosure has been described in detail and with reference to specific embodiments thereof, it will be apparent to one skilled in the art that various changes and modifications can be made therein without departing from the spirit and scope of the embodiments. Thus, it is intended that the present disclosure covers the modifications and variations of this disclosure, as well as other applications of the invention, provided they come within the scope of the appended claims and their equivalents.

Claims

CLAIMSWhat is claimed is:

1. A computer-implemented machine learning model training method for a machine learning pipeline, as applied to a computer system having a memory coupled to one or more processors and storing a plurality of programs to be executed by the processors and the computer system also having an input device and an output device, the computer system programed to execute the method comprising: iterative training, by the programmed computer system, wherein the iterative training comprises; obtaining domain data; extracting identifiable neural vectors from the domain data; extracting conceptual interpretable factors from the identifiable vectors; applying a structural causal model for extraction of causal associations among the conceptual features to output causal connectors; and employing a shared prior probability distribution over the connectors.

2. The method of Claim 1, wherein the step of extracting identifiable neural vectors from the domain data requires that the identifiable neural vectors having the same underlying model parameters are the same for vectors that are the same and that vector construction is automatic.

3. The method of Claim 1, wherein the conceptual interpretable factors have stable functions and support causal interventions.

4. The method of Claim 1, wherein applying a structural causal model further comprises extracting causal and statistical associations among the conceptual features in a manner that is scalable to settings comprising high-dimensional raw inputs and conceptual feature spaces.

5. The method of Claim 4, wherein the structural causal model is represented as a plurality of local mechanisms, wherein each of the mechanisms comprises a conditional distribution of the conceptual feature given a subset of related features.

6. The method of Claim 5, further comprising a prior distribution defined over the conditional distributions.

7. The method of Claim 6, further comprising: training the structural causal model and regularizing via a Kullback-Leibler divergence between the structural causal model and a model prior; and updating the model prior using Bayesian inference conditioned on a structural casual model obtained from a different computer-implemented machine learning model.

8. A computing device, comprising: one or more processors; memory coupled to the one or more processors; and a plurality of programs stored in the memory that, when executed by the one or more processors, cause the computing device to perform a plurality of iterative training operations comprising: obtaining domain data; extracting identifiable neural vectors from the domain data; extracting conceptual interpretable factors from the identifiable vectors;applying a structural causal model for extraction of causal associations among the conceptual features to output causal connectors; and employing a shared prior probability distribution over the connectors.

9. The device of Claim 8, wherein the step of extracting identifiable neural vectors from the domain data requires that the identifiable neural vectors having the same underlying model parameters are the same for vectors that are the same and that vector construction is automatic.

10. The device of Claim 8, wherein the conceptual interpretable factors have stable functions and support causal interventions.

11. The device of Claim 8, wherein applying a structural causal model further comprises extracting causal and statistical associations among the conceptual features in a manner that is scalable to settings comprising high-dimensional raw inputs and conceptual feature spaces..

12. The device of Claim 11 , wherein the structural causal model is represented as a plurality of local mechanisms, wherein each of the mechanisms comprises a conditional distribution of the conceptual feature given a subset of related features.

13. The device of Claim 12, further comprising a prior distribution defined over the conditional distributions.

14. The device of Claim 13, further comprising: training the structural causal model and regularizing via a Kullback-Leibler divergence between the structural causal model and a model prior; and updating the model prior using Bayesian inference conditioned on a structural casual model obtained from a different computer-implemented machine learning model.

15. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of an electronic device having a display, cause the electronic device to implement the iterative training of a machine learning pipeline by: obtaining domain data; extracting identifiable neural vectors from the domain data; extracting conceptual interpretable factors from the identifiable vectors; applying a structural causal model for extraction of causal associations among the conceptual features to output causal connectors; and employing a shared prior probability distribution over the connectors.

16. The storage medium of Claim 15, wherein the step of extracting identifiable neural vectors from the domain data requires that the identifiable neural vectors having the same underlying model parameters are the same for vectors that are the same and that vector construction is automatic.

17. The storage medium of Claim 15, wherein the conceptual interpretable factors have stable functions and support causal interventions.

18. The storage medium of Claim 15, wherein applying a structural causal model further comprises extracting causal and statistical associations among the conceptual features in a manner that is scalable to settings comprising high-dimensional raw inputs and conceptual feature spaces.

19. The storage medium of Claim 18, wherein the structural causal model is represented as a plurality of local mechanisms, wherein each of the mechanisms comprises a conditional distribution of the conceptual feature given a subset of related features.

20. The storage medium of Claim 19, further comprising a prior distribution defined over the conditional distributions.

Citation Information

Patent Citations

  • Techniques to train a neural network using transformations

    US20200293828A1

  • Identifying and quantifying confounding bias based on expert knowledge

    US20220101187A1

  • Efficient neural causal discovery

    US20240176994A1