Systems and methods for optimizers with enhanced neural estimation

A neural network-based optimization model addresses limitations in existing algorithms by performing efficient gradient calculations and direct search, improving prediction accuracy in portfolio management and other complex data scenarios.

JP2026504164APending Publication Date: 2026-02-03GOLDMAN SACHS & CO LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025543168
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-01-26
Filing Date
2024-01-23
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

Existing optimization algorithms impose significant limitations on the types of objective functions and constraints that can be employed, particularly in non-convex/concave scenarios, leading to inaccurate predictions due to the need for reformulating problems into infeasible forms or being computationally limited, often resulting in overly simplistic assumptions.

Method used

An optimization model utilizing a neural network architecture with multiple layers to perform training, differencing, loss recording, and metric calculation, enabling efficient gradient calculation and direct search to update weights for improved predictions.

Benefits of technology

The model provides accurate and efficient optimization by handling complex parameters and constraints, enhancing prediction accuracy in applications like portfolio management by minimizing transaction costs and capturing non-linear trading impacts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026504164000001_ABST
    Figure 2026504164000001_ABST
Patent Text Reader

Abstract

The method (600) includes receiving (614) a plurality of inputs (301, 403, 503) including domain parameters (404, 504) and initial weights (408, 508). The method also includes providing (620) the plurality of inputs to an optimization model (302, 402, 502). The method also includes using a first layer (304) of the optimization model to perform (622) a training and optimization process based on the plurality of inputs and based on a training objective. The method also includes using a second layer (306) of the optimization model to perform (624) a differencing operation on the output of the first layer. The method also includes using a third layer (308) of the optimization model to record (626) a loss based on a training objective used by the optimization model. The method also includes using a fourth layer (310) of the optimization model to calculate and store (628) metrics related to the training and optimization process. The method also includes using the optimization model to output (630) updated weights (418, 518).
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates generally to machine learning systems, and more particularly to a system and method for an optimizer with enhanced neural estimation. [Background technology]

[0002] Existing optimization algorithms impose significant limitations on the types of objective functions and constraints that users can employ, as well as the size of the data, given the high computational resource requirements of existing techniques. When dealing with non-convex / concave objectives, practitioners often face the difficult task of reformulating the problem into an infeasible quadratic or cone form, or are severely limited by computational considerations (e.g., a limited number of assets) associated with nonlinear optimization techniques. As a result, overly simplistic assumptions are often used in an attempt to obtain "optimal" results. However, such overly simplistic assumptions may ignore important information and provide inaccurate predictions. Summary of the Invention

[0003] The present disclosure relates to a system and method for an optimizer with enhanced neural estimation.

[0004] In a first embodiment, a method includes receiving a plurality of inputs including domain parameters and initial weights. The method also includes providing the plurality of inputs to an optimization model. The method also includes using a first layer of the optimization model to perform a training and optimization process based on the plurality of inputs and based on a training objective. The method also includes using a second layer of the optimization model to perform a differencing operation on an output of the first layer. The method also includes using a third layer of the optimization model to record a loss based on a training objective used by the optimization model. The method also includes using a fourth layer of the optimization model to calculate and store metrics related to the training and optimization process. The method also includes using the optimization model to output updated weights.

[0005] In a second embodiment, an apparatus includes at least one processor supporting optimization. The at least one processor is configured to receive a plurality of inputs, including domain parameters and initial weights. The at least one processor is also configured to provide the plurality of inputs to an optimization model. The at least one processor is also configured to use a first layer of the optimization model to perform a training and optimization process based on the plurality of inputs and based on a training objective. The at least one processor is also configured to use a second layer of the optimization model to perform a differencing operation on an output of the first layer. The at least one processor is also configured to use a third layer of the optimization model to record a loss based on a training objective used by the optimization model. The at least one processor is also configured to use a fourth layer of the optimization model to calculate and store metrics related to the training and optimization process. The at least one processor is also configured to use the optimization model to output updated weights.

[0006] In a third embodiment, a non-transitory computer-readable medium includes instructions supporting optimization that, when executed, cause at least one processor to receive multiple inputs including domain parameters and initial weights. The non-transitory computer-readable medium also includes instructions that, when executed, cause at least one processor to provide the multiple inputs to an optimization model. The non-transitory computer-readable medium also includes instructions that, when executed, cause at least one processor to perform a training and optimization process using a first layer of the optimization model based on the multiple inputs and based on a training objective. The non-transitory computer-readable medium also includes instructions that, when executed, cause at least one processor to perform a differencing operation on outputs of the first layer using a second layer of the optimization model. The non-transitory computer-readable medium also includes instructions that, when executed, cause at least one processor to use a third layer of the optimization model to record a loss based on a training objective used by the optimization model. The non-transitory computer-readable medium also includes instructions that, when executed, cause at least one processor to use a fourth layer of the optimization model to calculate and store metrics related to the training and optimization process. The non-transitory computer-readable medium also includes instructions that, when executed, cause the at least one processor to use the optimization model to output updated weights.

[0007] Other technical features may be readily apparent to those skilled in the art from the following figures, descriptions, and claims.

[0008] Prior to the following Detailed Description, it may be advantageous to set forth definitions of certain words and phrases used throughout this patent document. The terms "transmit," "receive," and "communicate," and their derivatives, encompass both direct and indirect communication. The terms "include" and "comprise," and their derivatives, mean inclusion without limitation. The term "or" is inclusive, meaning "and / or." The phrase "associated with," and its derivatives, means include, be included within, interconnect with, contain, be contained within, connect to or with, couple to or with, be communicable with, cooperate with, interleave, juxtapose, proximate to, be bound to or with, have, have a property of, have a relationship to or with, and the like.

[0009] Furthermore, various functions described below may be implemented or supported by one or more computer programs, each of which is formed from computer-readable program code and embodied in a computer-readable medium. The terms “application” and “program” refer to one or more computer programs, software components, sets of instructions, procedures, functions, objects, classes, instances, associated data, or portions thereof adapted for implementation in suitable computer-readable program code. The phrase “computer-readable program code” includes any type of computer code, including source code, object code, and executable code. The phrase “computer-readable medium” includes any type of medium accessible by a computer, such as read-only memory (ROM), random-access memory (RAM), hard disk drive, compact disc (CD), digital video disc (DVD), or any other type of memory. “Non-transitory” computer-readable medium excludes wired, wireless, optical, or other communication links carrying transient electrical or other signals. Non-transitory computer-readable media include media that can store data permanently and media that can store data and later be overwritten, such as rewritable optical disks or erasable memory devices.

[0010] As used herein, terms and phrases such as "have," "may have," "include," or "may include" a feature (such as a number, function, operation, or component) indicate the presence of the feature and do not exclude the presence of other features. Also, as used herein, the phrases "A or B," "at least one of A and / or B," or "one or more of A and / or B" may include all possible combinations of A and B. For example, "A or B," "at least one of A and B," and "at least one of A or B" may refer to (1) including at least one A, (2) including at least one B, or (3) including at least one A and at least one B. Furthermore, as used herein, the terms "first" and "second" modify various components, regardless of importance, and do not limit these components. These terms are used only to distinguish one component from another. For example, a first user device and a second user device may refer to different user devices, regardless of the order or importance of the devices. A first component may be referred to as a second component, and vice versa, without departing from the scope of this disclosure.

[0011] When an element (such as a first element) is referred to as being "coupled" or "connected" (operably or communicatively) to another element (such as a second element), it will be understood that it may be coupled or connected to the other element directly or through a third element. Conversely, when an element (e.g., a first element) is referred to as being "directly coupled" or "directly connected" to another element (e.g., a second element), it should be understood that there are no other elements (e.g., a third element) between the one element and the other element.

[0012] As used herein, the phrase “configured (or set) to” may be used synonymously with the phrases “suitable for,” “having the capacity to,” “designed to,” “adapted to,” “made to,” or “capable of,” depending on the context. The phrase “configured (or set) to” does not inherently mean “specially designed in hardware to.” Rather, the phrase “configured to” may mean that a device is capable of performing an operation in conjunction with another device or component. For example, the phrase “a processor configured (or set) to perform A, B, and C” may refer to a general-purpose processor (such as a CPU or application processor) that may perform operations by executing one or more software programs stored in a memory device, or a dedicated processor for performing operations (such as an embedded processor).

[0013] The terms and phrases used herein are provided only to describe some embodiments of the present disclosure and are not provided to limit the scope of other embodiments of the present disclosure. The singular forms "a," "an," and "the" should be understood to include plural references unless the context clearly dictates otherwise. All terms and phrases, including technical and scientific terms and phrases, used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the embodiments of the present disclosure belong. It will be further understood that terms and phrases as defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and should not be interpreted in an idealized or overly formal sense unless expressly defined as such herein. In some cases, terms and phrases defined herein may be interpreted to exclude embodiments of the present disclosure.

[0014] Definitions of other specific words and phrases may be provided throughout this patent document, and those skilled in the art should understand that in many, if not most, cases, such definitions apply to previous and future uses of such defined words and phrases.

[0015] Nothing in this application should be read as implying that any particular element, step, or function is an essential element required for inclusion in the scope of a claim. The scope of patented subject matter is defined solely by the claims. Furthermore, none of the claims are intended to invoke 35 U.S.C. §112(f) unless the precise words "means for" are followed by a participle. The use of any other term in the claims, including, but not limited to, "mechanism," "module," "device," "unit," "component," "element," "member," "apparatus," "machine," "system," "processor," or "controller," is understood by applicant to refer to structures known to those skilled in the relevant art and is not intended to invoke 35 U.S.C. §112(f). [Brief explanation of the drawings]

[0016] For a more complete understanding of the present disclosure and its advantages, reference should be made to the following description taken in conjunction with the accompanying drawings, in which like reference numerals indicate like elements and in which: [Figure 1] 1 illustrates an exemplary system that supports optimization using enhanced neural estimation, according to the present disclosure. [Figure 2] 1 illustrates an exemplary device that supports optimization using enhanced neural estimation, according to the present disclosure. [Figure 3] 1 illustrates an exemplary functional architecture for an optimization model according to the present disclosure. [Figure 4] 1 illustrates an exemplary optimization model process according to the present disclosure. [Figure 5] 1 illustrates an exemplary portfolio optimization model process according to the present disclosure. [Figure 6A] 1 illustrates an exemplary method for performing a training and optimization process using an optimization model according to the present disclosure. [Figure 6B] 1 illustrates an exemplary method for performing a training and optimization process using an optimization model according to the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0017] 1 to 6B and various embodiments of the present disclosure will be described with reference to the accompanying drawings. However, the present invention is not limited to the embodiments, and all modifications, equivalents, or alternatives should be understood to fall within the scope of the present disclosure. The same or similar reference numerals may be used throughout the specification and drawings to refer to the same or similar elements.

[0018] As mentioned above, existing optimization algorithms impose significant limitations on the types of objective functions and constraints that users can employ, as well as the size of data, given the high computational resource requirements of existing techniques. When dealing with non-convex / concave objectives, practitioners often face the difficult task of reformulating the problem into an infeasible quadratic or cone form, or are severely limited by computational considerations (e.g., a limited number of assets) associated with nonlinear optimization techniques. As a result, overly simplistic assumptions are often used in an attempt to obtain an "optimal" result. However, such overly simplistic assumptions may ignore important information and provide inaccurate predictions.

[0019] Various embodiments of the present disclosure provide an optimizer that provides enhanced neural estimation. The optimizer is an optimization model built using a neural network, which is used, for example, to calculate gradients and perform direct search in a highly efficient manner that takes into account various complex parameters, inputs, and constraints. In various embodiments, the optimization model can perform multi-period optimization on complex data provided by one or more feeder models to optimize the weights used by those feeder models to provide more accurate predictions and data correlations. The optimization model can receive as inputs from the feeder models domain parameters, i.e., the data used by and predicted using the feeder models, initial weights from the feeder models, and optional final weights from the feeder models, which can be modified and prepared for optimization by applying predetermined objectives and constraints to the inputs. The optimization model then performs various optimization processes as described in embodiments of the present disclosure to output updated model weights that are used to better predict results closely related to the original feeder models.

[0020] As one non-limiting example, in some embodiments, the optimization model can be used as a portfolio optimization network for stock trading in quantitative investment strategies. The portfolio optimization network can be used to generate trading recommendations for a portfolio manager that maximizes returns for a given risk level while minimizing transaction costs. Modern portfolio theory aims to maximize portfolio returns at a given risk level, and portfolio optimization remains a cornerstone of asset management. The original optimization problems addressed in modern portfolio theory can be solved using modern quadratic programming approaches. However, such approaches impose significant limitations on the types of objective functions and constraints a portfolio manager can employ in portfolio construction, as well as on the size of the portfolio given the high computational resource demands of such techniques. Therefore, when dealing with non-convex / concave objectives, practitioners often face the difficult task of reformulating the problem into an infeasible quadratic or conical form, or are severely limited by computational considerations associated with nonlinear optimization techniques, such as a limited number of assets. As a result, overly simplistic assumptions can often be used in an attempt to obtain an "optimal" portfolio. One such commonly used assumption is that transaction costs for constructing or rebalancing a portfolio are negligible.

[0021] In practice, trading costs are particularly significant for large asset-under-management (AUM) portfolios and can vary nonlinearly with trading positions, asset volatility, average daily trading volume, and so on. To mitigate the impact of high trading costs, portfolio managers often rebalance large positions over multiple days to reduce market-impact trading costs. To properly capture such costs, multi-day rebalancing should also take into account the longer-term impact of trades, as well as changes in expected returns (alpha decay) and risk over time. This further leads to both an increase in problem dimensionality and complex, non-convex objective functions. Therefore, realistic models should use multi-period optimization engines that are efficient at high-dimensional optimization problems under complex, nonlinear objective functions and constraints and can generate sequences of trades to be executed over multiple periods.

[0022] In practice, multi-period portfolio optimization is difficult to implement for several reasons. First, multi-period models are computationally intensive, especially when the population of assets considered is large. Second, the most common existing multi-period models do not handle real-world constraints. Finally, predicting returns, transaction costs, and risk over multiple days can be difficult in itself. Largely for these reasons, attempts to build robust multi-period optimizers have met with little success in the past. Moreover, even today, the majority of portfolio managers across the industry continue to rely on single-period optimizers.

[0023] However, the optimization model of the present disclosure can use fast and accurate calculation of gradients in neural networks, combined with increased computational power, to provide a direct search neural optimizer that addresses the challenges highlighted above. In various embodiments, the optimization model is built using a neural network, but does not provide direct inference. Rather, the optimization model uses a neural network-based approach to calculate gradients, perform direct search, and update neural network parameters in a highly efficient manner.

[0024] In various embodiments, the optimization model, when used for portfolio management, can use five categories of inputs: forecasts of future returns, equity volume and risk in the form of variance-covariance matrices, current positions / holdings, portfolio / account constraints, benchmarks, and volume estimates. All of these inputs can be defined for each time period (e.g., daily). In most cases, the inputs to the optimization model are generated by a quantitative investment strategy model (feeder model) already approved and used by the portfolio manager. The output of the optimization model can be the portfolio weights for each time period (e.g., each day).

[0025] Figure 1 illustrates an exemplary system 100 that supports optimization using enhanced neural estimation in accordance with the present disclosure. As shown in Figure 1, the system 100 includes multiple electronic devices 102a-102d, such as electronic computing devices, at least one network 104, at least one application server 106, and at least one database server 108 associated with at least one database 110. However, it should be noted that other combinations and arrangements of components may also be used herein.

[0026] In this example, each electronic device 102a-102d is coupled to or communicates through network(s) 104. Communication between each electronic device 102a-102d and the at least one network 104 may occur in any suitable manner, for example, via a wired or wireless connection. Each electronic device 102a-102d represents any suitable device or system used by at least one user to provide information to or receive information from an application server 106 or a database server 108. Any suitable number(s) and type(s) of electronic devices 102a-102d may be used in system 100. In this particular example, electronic device 102a represents a desktop computer, electronic device 102b represents a laptop computer, electronic device 102c represents a smartphone, and electronic device 102d represents a tablet computer. However, any other or additional types of electronic devices may be used in system 100. Each electronic device 102a-102d includes any suitable structure configured to transmit and / or receive information.

[0027] At least one network 104 facilitates communication between various components of system 100. For example, network(s) 104 may communicate Internet Protocol (IP) packets, Frame Relay frames, Asynchronous Transfer Mode (ATM) cells, or other suitable information between network addresses. Network(s) 104 may include one or more local area networks (LANs), metropolitan area networks (MANs), wide area networks (WANs), all or part of a global network such as the Internet, or one or more any other communication systems in one or more locations. Network(s) 104 may also operate according to any suitable communication protocol or protocols.

[0028] The application server 106 is coupled to at least one network 104 and is coupled to or otherwise in communication with the database server 108. The application server 106 can support various functions related to optimization using enhanced neural estimation embodied by at least the application server 106 and the database server 108. For example, the application server 106 can execute one or more applications 112, which can include an optimizer of various embodiments of the present disclosure and can be used to receive requests to calculate gradients, optimize weights, and perform direct searches in a highly efficient manner using the unique optimizer of various embodiments of the present disclosure. Data used by the unique optimizer can be received by the one or more applications 112 from a remote device, such as one or more of the electronic devices 102a-102d. In some embodiments, the data used by the unique optimizer can be stored in and retrieved from one or more databases 110 of the database server 108. The application server 106 may interact with the database server 108 as needed or desired to store and retrieve information from the database 110. Additional details regarding example functionality of the application server 106 are provided below. The one or more applications 112 may also present one or more graphical user interfaces to the user of the electronic devices 102a-102d, such as one or more graphical user interfaces that allow the user to retrieve and view data created and / or predicted using machine learning models optimized using the unique optimizer of various embodiments of the present disclosure from initial data, and display related results.

[0029] The database server 108 operates to store and facilitate retrieval of various information used, generated, or collected by the application server 106 and the electronic devices 102a-102d in the database 110. For example, the database server 108 may store various types of data used by the unique optimizer and other components of the present disclosure, such as information used in analyzing market data, statistics, signal processing, pattern recognition, econometrics, mathematical finance, weather forecasting, earthquake prediction, electroencephalography, control engineering, astronomy, communications engineering, and any area of ​​applied science and engineering that primarily involves temporal measurements, such as information including Internet of Things (IoT) device data and / or status, such as annual sales data, monthly subscriber numbers for various services, stock prices, temperature, rainfall, heartbeats per minute, and other measured metrics, stored in the database 110. Note, however, that in other embodiments, the database server 108 may be used to store information within the application server 106, in which case the application server 106 may store the information itself.

[0030] Some embodiments of the system 100 allow information to be collected or otherwise obtained from one or more external data sources 114 and pulled into the system 100, for example, for storage in the database 110 and use by the application server 106. In some embodiments, the one or more external data sources 114 may include a feeder model as described in this disclosure. Each external data source 114 represents any suitable source of information useful for performing one or more analyses or other functions of the system 100. At least a portion of this information may be stored in the database 110 and used by the application server 106 to perform one or more analyses or other functions using data stored in the database 110. Depending on the circumstances, the one or more external data sources 114 may be directly coupled to the network(s) 104 or indirectly coupled to the network(s) 104 via one or more other networks.

[0031] In some embodiments, the functionality of application server 106, database server 108, and database 110 may be provided in a cloud computing environment, for example, by using a proprietary cloud platform or by using a hosted environment such as the AMAZON WEB SERVICES (AWS) platform, the GOOGLE CLOUD platform, or MICROSOFT AZURE. In these types of embodiments, the described functionality of application server 106, database server 108, and database 110 may be implemented using a native cloud architecture, such as one that supports a web-based interface or other suitable interface. Among other things, this type of approach promotes scalability and cost efficiency while ensuring increased or maximum uptime. This type of approach allows electronic devices 102a-102d of one or more organizations (e.g., one or more companies) to access and use the functionality described in this patent document. However, different organizations may access different data or other different resources or functions within system 100.

[0032] In some cases, the architecture uses an architectural stack that supports the use of internal tools or datasets (meaning an organization's tools or datasets that access and use the described functionality) and third-party tools or datasets (meaning tools or datasets provided by one or more parties not using the described functionality). The datasets used in system 100 can have well-defined models and controls to enable effective import and use of the datasets, and the architecture can collect structured and unstructured data from one or more internal or third-party systems, thereby standardizing and combining the data source(s) with cloud-native data stores. The use of modern cloud-based and industry-standard technology stacks can enable smooth deployment and improved scalability of the described infrastructure, making it more resilient and achieving performance improvements, accelerating research and development efforts while reducing the time between new feature releases.

[0033] Among other possible use cases, native cloud-based architectures, or other architectures designed to use the optimizer and related methods disclosed herein, can be used to leverage data, such as market data, with advanced data analytics to make the investment process more reliable and reduce uncertainty. In these types of architectures, the described functionality can be used to achieve various technical benefits or advantages, depending on the implementation. For example, these approaches can be used to drive intelligence in investment or other processes by providing users and teams with information that could only be accessed through the application of data science and advanced analytics. Based on the described functionality, the approaches disclosed herein can meaningfully improve the sophistication of functions such as market selection, trade analysis, and risk and return management.

[0034] The value or benefits of data science and advanced analytics driven by the described approach can be highly useful or desirable. For example, deal sourcing can be driven by a deep understanding of market performance drivers to identify high-quality assets early in their lifecycle to increase or maximize investment returns. This can also position institutional or corporate investors to initiate outbound sourcing efforts to drive proactive partnerships with operating partners. Furthermore, with regard to deal analysis during the transaction due and execution stages, this can help optimize deal strategies by providing precision and clarity to the underlying market fundamentals.

[0035] While FIG. 1 illustrates an example system 100 supporting optimization using enhanced neural estimation, various modifications may be made to FIG. 1. For example, system 100 may include any number of electronic devices 102a-102d, a network 104, an application server 106, a database server 108, a database 110, an application 112, and an external data source 114. These components may be located in any suitable location or may be distributed over a wide area. Additionally, while FIG. 1 illustrates one example operating environment in which optimization using enhanced neural estimation may be used, this functionality may be used in any other suitable system.

[0036] 2 illustrates an example device 200 that supports optimization using enhanced neural estimation in accordance with the present disclosure. One or more instances of device 200 may be used to, for example, at least partially implement the functionality of application server 106 of FIG. 1. However, the functionality of application server 106 may be implemented in any other suitable manner. In some embodiments, device 200 illustrated in FIG. 2 may form at least a portion of electronic devices 102a-102d, application server 106, or database server 108 of FIG. 1. However, each of these components may be implemented in any other suitable manner.

[0037] 2, device 200 illustrates a computing device or system that includes at least one processing device 202, at least one storage device 204, at least one communication unit 206, and at least one input / output (I / O) unit 208. Processing device 202 may execute instructions loadable into memory 210. Processing device 202 may include any suitable number(s) and type(s) of processors or other processing devices in any suitable arrangement. Exemplary types of processing device 202 include one or more microprocessors, microcontrollers, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or discrete circuits.

[0038] Memory 210 and persistent storage 212 are examples of storage device 204, representing any structure(s) capable of storing and facilitating retrieval of information (such as data, program code, and / or other suitable information, either temporary or permanent). Memory 210 may represent random access memory or any other suitable volatile or non-volatile storage device(s). Persistent storage 212 may include one or more components or devices that support long-term storage of data, such as read-only memory, a hard drive, flash memory, or an optical disk. Device 200 may also access data stored in external memory storage locations with which device 200 is in communication, such as one or more online storage servers.

[0039] The communications unit 206 supports communications with other systems or devices. For example, the communications unit 206 may include a network interface card or a wireless transceiver that facilitates communications over a wired or wireless network. The communications unit 206 may support communications over any suitable physical or wireless communications link(s). As a particular example, the communications unit 206 may support communications over the network(s) 104 of FIG. 1.

[0040] I / O unit 208 allows for the input and output of data. For example, I / O unit 208 may provide a connection for user input via a keyboard, mouse, keypad, touchscreen, or other suitable input device. I / O unit 208 may also send output to a display, printer, or other suitable output device. Note, however, that if device 200 does not require local I / O, such as when device 200 represents a remotely accessible server or other device, I / O unit 208 may be omitted.

[0041] In some embodiments, the instructions executed by processing device 202 include instructions that implement functionality of application server 106. Thus, for example, the instructions executed by processing device 202 may cause device 200 to perform various functions related to optimization using enhanced neural estimation. As a particular example, the instructions may cause device 200 to receive multiple inputs including domain parameters and initial weights, provide the multiple inputs to an optimization model, use a first layer of the optimization model to perform a training and optimization process based on the multiple inputs and based on a training objective, use a second layer of the optimization model to perform a differencing operation on an output of the first layer, use a third layer of the optimization model to record a loss based on the training objective used by the optimization model, use a fourth layer of the optimization model to calculate and store metrics related to the training and optimization process, and use the optimization model to output updated weights. The instructions may also cause device 200 to present to a user of device 200 or a user of electronic device 102a-102d one or more graphical user interfaces, such as one or more graphical user interfaces that allow a user to retrieve and view results of the optimization process, such as presenting one or more results of a portfolio forecast.

[0042] Although Figure 2 illustrates one example of a device 200 that supports optimization using enhanced neural estimation, various modifications may be made to Figure 2. For example, computing and communication devices and systems come in a wide variety of configurations, and Figure 2 does not limit the disclosure to any particular computing or communication device or system.

[0043] FIG. 3 illustrates an example functional architecture 300 for an optimization model 302 according to the present disclosure. For ease of explanation, the functional architecture 300 of FIG. 3 may be implemented using or provided by one or more applications 112 executed by the application server 106 and / or the database server 108 of FIG. 1, and the application server 106 and the database server 108 may be implemented using one or more devices 200 of FIG. 2. However, the functional architecture 300 may be implemented using or provided by any other suitable device(s), such as electronic devices 102a-102d, or implemented in any other suitable system(s).

[0044] The architecture 300 includes an optimization model 302 that receives inputs 301, such as domain parameters, initial weights, and optional final weights, performs optimization on the inputs, and provides outputs 303, such as updated parameters or weights, as described in various embodiments of this disclosure. The optimization model 302 is a neural network-based optimizer that derives unknown arg min from the inputs. x We model an unknown function φ(·) that maps to f(x), where f is the desired objective function. Since the input x is a fixed constant x0, without loss of generality we can take x0 = (1, 2, ..., T), where T is the number of periods. Thus, at each optimization, neural network parameters are estimated such that φ is close to f. In this sense, there may not be a one-time neural network parameter estimation, but there may be hyperparameters related to the definition of φ.

[0045] First, φ is a neural network contained in the optimization model 302, which is composed of four layers. A neural network is a computational model inspired by the structure of the human brain, in that it similarly consists of a network of interconnected neurons that propagate information upon receiving a set of stimuli from neighboring neurons that approximate a mapping between input and output. Neural networks can be stacked using layers of neurons. There are various types of layers that can be used in a network, each suited to the type of data and features they process. For example, densely connected layers are often found in regression problems, and recurrently connected layers are often used in time series analysis.

[0046] The first layer of the optimization model 302 is the policy layer 304. In various embodiments, the policy layer 304 is a direct search policy layer. Depending on the problem to be optimized using the optimization model 302, the policy layer 304 can take different forms. For example, the policy layer 304 can be in a dense form 305, e.g., a dense neural network, or in a recursive form 307, e.g., a recursive neural network (RNN). In the dense form 305, the policy layer 304 is an n-th layer with m units. layerThe policy layer 304 may consist of fully connected layers. For example, in the dense configuration 305, the policy layer 304 may be a densely connected feedforward neural network in which neurons in the layer are connected to all neurons in the previous layer. Neurons function as general approximators in that they can be trained to approximate any given nonlinear input-output mapping. Mathematically, even a multilayer perceptron (MLP) with one hidden layer can approximate an arbitrarily close mapping in the limit for any continuous function. Neurons demonstrate their interpolation ability by generalizing even in sparse data space regions. In a densely connected layer, a neuron consists of four parts: input values, weights and biases, a weighted sum, and an activation function. A neuron operates by taking the outputs from all neurons in the previous layer. It then multiplies these inputs by their respective weights (known as a weighted sum).

[0047] These products are then added together with a bias. The result of this calculation is then passed to an activation function to generate the perceptron's output. Most activation functions are nonlinear, which plays a crucial role in capturing nonlinearities and improving the effectiveness of the network. Without a nonlinear activation function, the output would be a sum of simple linear functions, which would still output a linear function. Thus, activation functions allow the network to create complex functions by using a combination of multiple neurons. Such activation functions can include, but are not limited to, a sigmoid function, a hyperbolic tangent (tanh) function, and a rectified linear unit (ReLU) function.

[0048] In feedforward neural networks, information travels in only one direction: from the input layer through the hidden layer to the output layer. Information travels straight through the network and never touches the same node twice. Feedforward neural networks do not remember the inputs they receive and therefore have no concept of temporal order, but they allow for parallelization. Forward passes are used in the neural network inference step, while backpropagation is used in the model fitting step as described in this disclosure. Backpropagation involves moving backward through the neural network to find the partial derivative of the error with respect to the weights, which is then used by gradient descent, an algorithm that can iteratively minimize a given function. This allows the neural network to learn during the training process. A complete forward pass and backpropagation through the entire input data set is called an epoch. The number of epochs is a hyperparameter that defines the number of times the algorithm works through the entire input data set. In neural network training, there is typically a trade-off between the maximum number of epochs and model accuracy.

[0049] In the recursive form 307, the dense architecture is modified to include as input intermediate weights up to time t, for t=1,...,T, as shown in FIG. 3. The output of the final layer is the desired T x k tensor of weights, such as portfolio weights in some embodiments. RNNs are a class of neural networks useful for modeling sequence data. Derived from feedforward networks, RNNs behave similarly to how the human brain functions, with the creation of memory in the output sequential data. Unlike feedforward neural networks, in RNNs, information circulates in a loop. When an RNN makes a decision, it considers the current input and what it has learned from previously received inputs. In other words, an RNN has two inputs, the current and the most recent, and combines them to make a prediction. This can be important for implementations such as improving time series forecasting, as the data sequence contains important information about what will happen next.

[0050] Furthermore, RNNs can fine-tune their weights using both gradient descent and backpropagation through time (BPTT). BPTT involves performing backpropagation on an unrolled RNN. Unrolling is a visualization and conceptual tool illustrated by the recursive form 307 in Figure 3. Figure 3 shows that an RNN can be viewed as a sequence of neural networks trained one after the other using backpropagation. On the left side of the visualization of the recursive form 307 shown in Figure 3, the RNN is unrolled after the equals sign. Note that there is no cycle after the equals sign because different time steps are visualized and information is passed from one time step to the next. Within BPTT, the error is backpropagated from the last time step to the first time step, unrolling all time steps. This allows the error to be calculated for each time step, allowing the weights to be updated.

[0051] The second layer of the optimization model 302 is the processing layer 306. The processing layer 306 performs a parameter-free transformation on the output of the first layer 304 and therefore does not require calibration or estimation in various embodiments. The third layer is the loss layer 308. The loss layer 308 takes in an arbitrary function (the objective function) and records the loss. In various embodiments, the loss layer 308 does not have parameters that require estimation. The fourth layer is the metric layer 310. The metric layer 310 calculates and stores down metrics for the optimization. These typically include various components of the objective function and, in some embodiments, constraint values.

[0052] As will be appreciated from this description of the optimization model 302, in some embodiments, the calibration and estimation performed are related to the specific form of the first layer 304, i.e., hyperparameters governing the number of layers and number of units, as well as parameters related to the actual optimization process. The parameters may be selected based on the problem to be solved and other needs. In various embodiments, at the architectural level, these selectable parameters may include whether the first layer 304 is a dense or recurrent network, the number of hidden layers, the number of units per layer, whether to use a kernel initializer, a bias initializer, a bias, and / or the intermediate activation function to be used (sigmoid, tanh, ReLU, etc.). At the optimization level, selectable parameters of the optimization model 302 include the initial learning rate of the model fitting algorithm used by the optimization model, the exponential decay rate of the first moment estimates, the exponential decay rate of the second moment estimates, a numerical stability constant (e.g., a very small number to prevent division by zero in the implementation), a minimum delta value for the early termination convergence criterion, and / or a maximum number of epochs.

[0053] To perform these calibrations on the optimization model 302, the search space of hyperparameters can be defined, for example, according to Table 1 shown below. [Table 1]

[0054] As a preliminary step to addressing the main optimization problem, trial optimization problems can be run to determine hyperparameters that achieve a good balance between convergence speed and accuracy. These parameters can then be fixed as the parameters used for, for example, a multi-period portfolio. In some embodiments, the default optimization model 302 can be set as a default model using, for example, four dense hidden layers and 16 units per layer. As shown in Table 1 above, in some embodiments, the optimization model 302 is fit for up to 10,000 epochs. In some embodiments, both the weights and biases can be initialized using the He uniform variance scaling initialization method.

[0055] In various embodiments, the optimization model 302 performs a training / optimization process using inputs 301 (which may include data from feeder models and model parameters / weights) to find, in a highly efficient manner, network-trainable weights that minimize the loss function defined by the optimization process. The selection of such an optimization process to achieve the training / optimized weights while maintaining efficiency is an important consideration. For example, multi-period portfolio optimization with constraints on the weights of securities in the portfolio is a high-dimensional optimization problem that cannot be solved by traditional convex optimization techniques, such as quadratic programming. Furthermore, when using nonlinear optimization algorithms to solve the problem, computational time can increase exponentially with increasing dimensionality. However, because the optimization model 302 can leverage stochastic optimization, it is computationally more efficient and can scale in dimensionality (number of assets) with time. For example, as the number of assets in the portfolio increases, the computational time to fit using the optimization model 302 remains at a similar level.

[0056] As an example, the optimization model 302 may perform a model fitting process such as the Adam optimization algorithm. The model fitting process uses backpropagation or BPTT to fit the neural network. In some embodiments, the model fitting process is an extension of stochastic gradient descent. Stochastic gradient descent involves using a single learning rate for weight updates, where the learning rate for each network parameter (each weight) does not change during training. However, the model fitting process used by the optimization model 302 may adapt the parameter learning rate in real time based on the average of the first and second moments by calculating an exponential moving average of the gradients and the squared gradient. For example, the aggregate gradient m at time t t can be expressed as the following equation (1).

number

[0057] In addition to the cumulative sum of gradients, the model fitting process also calculates a running weighted average of the squared gradients v t This can be expressed as the following equation (2).

number

[0058] The parameters β1 and β2 can control the decay rate of both moving averages. In some embodiments, default values ​​such as β1=0.9 and β1=0.999 can be used. m t and v t may both be initialized as 0, so that as both β1 and β2 approach 1, they may be biased towards 0. The optimization model 302 generates a bias-corrected m t and v tThis problem is addressed by calculating: This bias correction is also performed to control the weights while reaching the global minimum to prevent high oscillations when close to the global minimum, which can be expressed as shown in equation (3) below.

number

[0059] The optimization model 302 then calculates the bias-corrected weight parameters as part of the model fitting process.

number

number

[0060] The optimization model 302 performs backpropagation combined with the optimization process described above to find neural network trainable weights that minimize the loss function. The loss function of the optimization model 304 is calculated by the objective function

number

number

number

[0061] With these inputs and outputs in mind, the custom policy layer 304, where the neural network trainable parameters reside, takes x0 as an input and W * The second layer, the processing layer 306, returns W * , which is a parameter-free custom layer used to apply a static differencing operation to obtain information such as the trades for each period. The third and fourth layers, loss layer 308 and metric layer 310, are custom parameter-free layers used to record loss function values ​​and any associated metrics, respectively. Incidentally, these custom layers are included in optimization model 302 to allow flexibility in defining arbitrary tensor-based loss functions and metrics that operate on weights (e.g., portfolio weights) or weight differences corresponding to optimization objectives and constraints without changing the architecture of the neural network.

[0062] The policy layer 304, in various embodiments, is a layer with trainable parameters that vary during adaptation. The input x0 passes through a sequence of several layers (e.g., densely connected layers or recurrently connected layers) and ultimately generates a policy layer W that minimizes the objective function.* This architecture definition can also be dynamically adjusted by the model to accommodate user-specified hard constraints. The soft constraints are

number

[0063] Furthermore, given a neural network, the optimization process reduces to constructing a neural network loss function with user-specified inputs that define the objective function (e.g., market inputs including return, risk, volume, etc.), and optionally combining this with any user-specified soft constraints. The model then propagates and backpropagates through the network using batches of 1 x0 for a number of numbered epochs, corresponding to one iteration of the optimization model 302. The maximum number of epochs and early termination criteria based on the loss function default to reasonable values, but can be adjusted to trade off optimization time and accuracy. The model has converged when the loss function no longer decreases significantly with each epoch, and the optimal weights W are calculated from x0. * We discovered trainable parameters for a neural network that allow for a mapping to

[0064] Two other important points about optimization model 302 are worth noting here. First, optimization model 302 is unique in that, despite using a neural network, the training / learning objective of optimization model 302 is not to achieve a resulting neural network that performs well out-of-sample. Rather, the focus of optimization model 302 is the fitting process, which provides weights that optimize the objectives / constraints. Second, the fitting process is performed not only for any changes to the objective function and constraints (e.g., adding constraints, changing risk sensitivity, etc.), but also for the underlying data (e.g., market data including returns, risk, etc.), because these together define the neural network's loss function and therefore change the problem and, in turn, any changes to the optimization problem.

[0065] Although FIG. 3 illustrates one example functional architecture 300 for optimization model 302, various modifications may be made to FIG. 3 . For example, various components and functions of FIG. 3 may be combined, further subdivided, duplicated, or reconfigured according to particular needs. Also, one or more additional components and functions may be included as needed or desired. Computing and machine learning architectures may be provided in a wide variety of configurations, and FIG. 3 does not limit this disclosure to any particular computing or machine learning architecture. For example, the components of architecture 300 illustrated in FIG. 3 may be depicted as being executed on a single electronic device, such as one of electronic devices 102a-102d, by one or more cloud servers, such as application server 106 and / or one or more of database servers 108, or by a combination of electronic devices and servers in a distributed environment. For example, input 301 provided to optimization model 302 may be provided by an electronic device, such as one of electronic devices 102a-102d, and the input is transmitted over a network to one or more remote servers, such as application server 106 and / or database server 108, which execute optimization model 302 to provide output 303 that is subsequently sent back to the electronic device that provided input 301.

[0066] 4 illustrates an example optimization model process 400 according to the present disclosure. For ease of explanation, process 400 is described as involving the use of one or more applications 112 executed by application server 106 and / or database server 108 of FIG. 1, and application server 106 and database server 108 may be implemented using one or more devices 200 of FIG. 2. However, process 400 may be performed using any other suitable device(s), such as electronic devices like electronic devices 102a-102d, and in any other suitable system(s).

[0067] As shown in FIG. 4 , in a first step, a feeder model 401 provides model inputs 403 to be used by an optimization network 402, e.g., the optimization model 302 described above. The model inputs 403 can include various inputs, such as domain parameters 404, e.g., parameters related to a particular field, industry, and / or problem to be solved. The model inputs 403 also include initial weights 408, which are weights provided by the feeder model 401 and used by the feeder model 401 in performing predictions. Optionally, the model inputs 403 can include optional final weights 410, which can be provided to apply to the rebalancing problem found by the optimization network 402, although a final result is given, and the path leading to it is provided; this can provide various benefits, such as reducing the transaction costs of the neural network based on the new weights provided by the optimization network 402.

[0068] In the next step, as shown in FIG. 4 , the objective 412 and constraints 414 are defined for the optimization performed by the optimization network 402. As described in this disclosure, the optimization network 402 is a neural network-based optimizer that can perform optimization for problems of large dimensionality, handle complex, highly nonlinear, real-world constraints, and solve optimization problems where the objective function 412 or constraints 414 can include stochastic parameters. For large numbers of assets, the optimization network 402 has been found to outperform traditional nonlinear optimization packages by orders of magnitude (up to 1,000 times), producing results in minutes that would previously have taken weeks. Overall, the model is robust and works with a wide range of objective functions 412 and constraints 414. For certain problems where a convex optimization solution exists, the optimization network 402 agrees with those results with a high degree of accuracy. The model output 416 provided by the optimization network 402 has also been found to be numerically stable. The model output 416, in various embodiments, includes updated weights 418 created using at least the initial weights 408 by performing the backpropagation and optimization process described in this disclosure. In some embodiments, the updated weights 418 can then be used by the feeder model 401 to provide an optimized, more accurate prediction. In some embodiments, the prediction optimized by the optimization network 402 may be output as part of the model output 416.

[0069] Although FIG. 4 illustrates one example of an optimization model process 400, various modifications may be made to FIG. 4 . For example, various components and functions of FIG. 4 may be combined, further subdivided, duplicated, or rearranged according to particular needs. Also, one or more additional components and functions may be included as needed or desired. Computing systems, machine learning systems, and related processes may be provided in a wide variety of configurations, and FIG. 4 does not limit the disclosure to any particular computing system, machine learning system, or related process. For example, the steps of process 400 illustrated in FIG. 4 may be depicted as being performed on a single electronic device, such as one of electronic devices 102a-102d, by one or more cloud servers, such as application server 106 and / or one or more of database servers 108, or by a combination of electronic devices and servers in a distributed environment. For example, the model input 403 provided to the optimization network 402 may be provided by an electronic device that executed the feeder model 401, such as one of the electronic devices 102a-102d, and the input 403 is sent over a network to one or more remote servers, such as the application server 106 and / or the database server 108, which execute the optimization network 402 and provide an output 416 that is later sent back to the electronic device that provided the input 403.

[0070] 5 illustrates an example portfolio optimization model process 500 according to the present disclosure. For ease of explanation, process 500 is described as involving the use of one or more applications 112 executed by application server 106 and / or database server 108 of FIG. 1, where application server 106 and database server 108 may be implemented using one or more devices 200 of FIG. 2. However, process 500 may be performed using any other suitable device(s), such as electronic devices like electronic devices 102a-102d, and in any other suitable system(s).

[0071] As shown in FIG. 5 , in a first step, a feeder model 501 provides model inputs 503 for use by a portfolio optimization network 502, e.g., the optimization model 302 or optimization network 402 described above. The model inputs 503 can include various inputs, such as domain parameters 504, e.g., parameters related to a particular field, industry, and / or problem to be solved. For example, in this example, the domain parameters are related to market data and include inputs such as return 505, risk 506, and volume 507. The model inputs 503 also include initial weights 508, which are weights provided by the feeder model 501 and used by the feeder model 501 in performing inference and prediction. Optionally, the model inputs 503 can include optional final weights 510, which can be provided to apply to the rebalancing problem found by the portfolio optimization network 502, although a final result is given, and the path leading to it is determined by the portfolio optimization network 502; this can provide various benefits, such as reducing the trading costs of the neural network based on the new weights provided by the portfolio optimization network 502.

[0072] In this example, a portfolio optimization network 502 is used to optimize a portfolio forecast. As shown in FIG. 5 , the portfolio optimization network 502 can use five categories of inputs and optimization criteria: forecasts of future returns 506; stock volume in the form of a variance-covariance matrix 507 and risk 506; current positions / holdings 509; objectives 512, such as benchmarks and volume estimates; and portfolio / account constraints 514. All of these inputs can be defined on a time-period basis (e.g., for each day). In various embodiments, the inputs to the portfolio optimization network 502 are generated by one or more feeder models 501 already approved and used by risk management operations. Model inputs 503 are provided in a first step, and in a next step, the objectives 512 and constraints 514 can be defined and / or input for the optimization performed by the portfolio optimization network 502, as shown in FIG. 5 . In a next step, the optimization network 502 performs the optimization using the inputs.

[0073] In a next step, output 516 is provided by portfolio optimization network 502. In this example, model output 516 includes updated weights 518 optimized using domain parameters 504, position parameters 509 (initial weights 508 and optional final weights 510). For example, updated weights 518 may be multi-period weights for each day's portfolio. Updated weights 518 are created using at least initial weights 508 by performing the backpropagation and optimization process described in this disclosure. In some embodiments, updated weights 518 are then used by feeder model 501 to provide optimized, more accurate forecasts. In some embodiments, the forecasts optimized by portfolio optimization network 502 may be output as part of model output 516.

[0074] The model inputs 503 provided by one or more feeder models 501 and used by the neural optimizer may depend on the objective function and constraints of the particular optimization problem. For example, in the case of multi-period portfolio optimization, the inputs 503 typically include a list of assets and their corresponding forecasts of return 505, risk 506, and volume 507. Specifically, in such embodiments, the inputs include expected future returns 505 (u t ), the expected future variance-covariance matrix for the asset 506 (Σ t ), and the expected future trading volume for individual assets 507(v t ).

[0075] In addition to multi-period portfolio optimization, the optimization model of various embodiments of the present disclosure can handle multiple other optimization processes, such as alpha harvesting, portfolio rebalancing, portfolio replication, etc. Depending on the type of optimization problem to be solved, the portfolio optimization model 502 may use a priori known deterministic parameters (u t ,Σ t ,v t ), or the optimization model 502 can receive stochastic parameters (u) that are affected by a wide range of exogenous random factors (e.g., the overall volatility or liquidity of the market over the period) or endogenous levels. t ,Σ t ,v t ) can be used. For example, v t can be a random vector with a given distribution.

[0076] Depending on the type of portfolio optimization problem / objective 512 (e.g., alpha harvesting, portfolio rebalancing, portfolio replication), the initial weights 508 and / or final weights 510 of the portfolio, or the weights of a benchmark portfolio, can be provided as inputs 503. Similarly, additional parameters or metadata specifying portfolio constraints 514 can be provided. For example, risk concentration by industry can use an input parameter specifying a threshold for concentration, allowing additional data to be used to map each asset to a corresponding industry. For deterministic inputs, future returns 505, covariance matrices 506, and volumes 507 can be provided for each time period. For example, a one-month optimization with daily portfolio changes can use return, risk, and volume forecasts for each day during the month. Thus, standard inputs can be asset vectors (for returns and volumes) or matrices (for risk) with additional time dependencies. These inputs 503 can be generated by alpha and risk feeder models 501, which can be pre-validated for use.

[0077] In some embodiments, for input data, the optimization network 502 can apply checks to ensure the consistency of the input data. For example, the optimizer may expect predicted returns and volumes to be within certain ranges (e.g., predicted volumes are not negative) and the variance-covariance matrix to be positive semi-definite. In some embodiments, the model can perform data quality checks to check for consistency in the number of securities (the number and order of securities appearing in all inputs match), validity checks of volumes (all volumes are non-negative), positive definite checks of variance-covariance matrices (input variance-covariance matrices are positive definite), feasibility checks of constraints (constraints are feasible if a weight solution exists), and infinity checks of inputs 503 (market information inputs are not infinite and position weight inputs sum to 1).

[0078] In some embodiments, a key assumption of the optimization network 502 is that the local optima found by the model are close to the true global solution of the problem, even when the dimensionality of the problem is large. The output 516 of the optimization model 502 for high-dimensional single-step mean-variance optimization problems has been found to yield model outputs that are extremely close to the global solution. For multi-step optimization, results show that the optimization network 502 significantly outperforms benchmarks such as Monte Carlo benchmarks. In some embodiments, another key assumption is that the output of the optimization network 502, which relies on numerical techniques and algorithms, is numerically stable, and the numerical stability of the optimization network 502 has been confirmed.

[0079] Furthermore, in the example of FIG. 5 , where the optimization network 502 is used to optimize market and portfolio forecasts, several assumptions may be made related to the market impact of trading activity. For example, it may be assumed that the market impact is a nonlinear function of trade size and inversely proportional to the average daily trading volume. Another assumption may be that the market impact continues over several trading periods and slowly asymptotically converges to a certain percentage of the first-day impact (i.e., to a long-term trading impact threshold). Another assumption may be that the model inherits all assumptions from its feeder model 501. That is, if assumptions related to the input data of the optimizer 502 (such as assumptions related to alpha forecasts, daily volume, risk, etc.) are violated, it may adversely affect the optimization results.

[0080] In some embodiments, the target users of the portfolio optimization network 502, or the beneficiaries of the results provided by the portfolio optimization network 502, are stock valuation, asset management, and / or asset trading personnel. Such personnel may manage client stock portfolios using various quantitative strategies. The quantitative strategies employed by such personnel may aim to generate excess returns by customizing portfolio exposure to selected factors. Personnel responsible for optimal execution of trades may manage stock trades with large total trading values ​​per day. Given the size and daily trading volume of some portfolios, the transaction costs of trading due to market impact may be significant and suitable for optimization using the optimization network 502. The portfolio optimization network 502 may provide multi-period optimization (e.g., multi-period weights 518) and reduce transaction costs by taking these multi-period outputs provided by the portfolio optimization network 502 into account in decisions made by such personnel.

[0081] The optimization performed by the optimization model 502 is a process of searching inputs, such as input 503, in a given search space to maximize or minimize a given function or objective, such as objective 512. Constrained optimization, such as using constraints 514, further restricts the search space by requiring additional equalities and / or inequalities to hold for the optimal solution. Optimization methods can be described in terms of global optimization and local optimization. The goal of global optimization is to find a minimum or maximum value, if one exists, throughout the entire search space. A global search is performed in each dimension of the objective function to obtain the best solution. Existing global optimization methods include uniform grid search, exhaustive (enumerative) search strategies, and successive approximation methods. While these existing methods converge under moderate assumptions, their computational cost increases exponentially as the dimension increases, making them impractical for higher-dimensional problems. Unlike global optimization, which finds the best solution across a given set, local optimization attempts to find the optimal solution within a neighboring set of candidate solutions to reduce computational cost. Existing local search algorithms belonging to deterministic optimization, such as the Nelder-Mead algorithm and the Broyden-Fletcher-Goldfarb-Shanno (BFGS) algorithm, often find locally optimal solutions of varying quality, i.e., they are easily trapped in local minima depending on the starting point of the search.

[0082] However, the optimization model of various embodiments of the present disclosure, including the optimization network 502, involves stochastic optimization that avoids local optimum traps and introduces randomness into the process. This allows less optimal local decisions to be made within the search procedure, increasing the probability of finding a global optimum for the objective function. This is achieved by allowing the optimization model 502 to escape local optima by making locally suboptimal steps or moves in the search space. The use of randomness in the stochastic optimization performed by the optimization network 502 does not mean that the algorithm is random. Rather, it means that some decisions made during the search procedure contain a portion of randomness. For example, the move from the current point to the next point in the search space made by the portfolio optimization network 502 may be made according to a probability distribution for optimal moves.

[0083] In various embodiments, the portfolio optimization network 502 is a stochastic local optimizer implemented using a neural network framework. The portfolio optimization network 502 starts from an arbitrary point and searches for an optimal solution. The inherent stochasticity built into the neural network formulation (e.g., stochastic gradient descent, stochastic weight initialization) increases the likelihood that the algorithm will avoid local extreme points. While there may be no guarantee of finding a global optimum, the optimization results of the portfolio optimization network 502 have been found to tend to be more stable and less dependent on initial values ​​than traditional local optimizers. Thus, the portfolio optimization network 502 computes gradients and performs a direct search in a highly efficient manner.

[0084] In particular, when the portfolio optimization network 502 is used to output multi-period weights 518, the portfolio optimization network 502 aims to find a set of security weights for each period to achieve the maximum net return for a given risk tolerance of the portfolio in the presence of multi-period market impacts. The objective function of objective 512 can include three parts: total portfolio return (aiming to maximize the weighted sum of expected asset returns over multiple periods fed from one or more feeder models 501), total portfolio risk (aiming to minimize portfolio risk, which is the weighted sum of expected asset-asset variance-covariance matrices over multiple periods for a given risk tolerance fed from one or more feeder models 501), and total market impact (aiming to minimize loss of portfolio return due to short-term and long-term impacts of securities trading), as described below.

[0085] In various embodiments, using the optimization network 502 involves determining the portfolio weights at time t for k assets in the portfolio.

number

number

number

number

number

[0086] Similarly, the total risk is defined as the sum over all periods, which can be expressed as shown in equation (6) below:

number

[0087] In addition to risk and return, the portfolio optimization network 502 can consider transaction costs resulting from the market impact of trades for multi-period portfolio optimization; such transaction costs are generally ignored in single-period portfolio optimization approaches. Transaction costs can include two terms: the immediate effect of a trade and its long-term impact. Portfolio transaction costs can be assumed to be additive across assets and time, with the time dependency being captured by the long-term impact. The total transaction cost is calculated as τ tot (W). In some embodiments, the optimization problem to be solved using the portfolio optimization network 502 can be written as the following optimization problem:

number

[0088] The first constraint above (initial weights 508) indicates that a "day 0" portfolio is given, which is used to evaluate the trading impact on day 1. The second constraint above (final weights 510) is optional; the final portfolio is given, but the path to it can be applied to the rebalancing problem found by the portfolio optimization network 502. In some embodiments, if optional final target portfolio weights are not provided, the multi-period portfolio optimization performed using the portfolio optimization network 502 becomes an alpha harvesting problem that seeks excess returns for a given forecast. In some embodiments, if optional final target portfolio weights are provided, the optimization problem can become a rebalancing problem performed over multiple periods to reduce trading costs. The third constraint above indicates that the security weights in the portfolio always sum to 1 throughout the optimization period. In some embodiments, the last constraint above restricts the portfolio to not consider either leverage or shorting. If one wants to enable shorting and leverage, the last two constraints can be easily relaxed within the framework.

[0089] In practice, the portfolio manager may use additional constraints resulting from either investment views or internal guidelines. These soft constraints are therefore expressed as a soft constraint penalty function f pen Also, in portfolio optimization network 502, these constraints can be any function acting on the weights w, but in some embodiments may be restricted to linear min-max constraints, which can be expressed as shown in equations (7) and (8) below.

number

number

[0090] Regarding trading costs, several types of costs may be involved when implementing a long / short equity strategy, such as leverage costs (funding spreads paid for long positions and short fees paid for short positions), dividend deduction taxes (which negatively impact long positions), and transaction costs (commissions paid to brokers and market impact). Among these costs, market impact costs, which include both immediate and long-term market impact, are important to quantitative asset managers because they increase more than proportionally with the capital allocated to the trading strategy. Other costs are proportional to the allocated capital and shift net returns down by the same percentage. Properly controlling these costs using optimization techniques can significantly improve portfolio performance. For sufficiently large order sizes, also known as "metaorders," a major component of trading costs appears in the form of market impact, which can be measured as the difference between the average execution price of the order and the price prevailing before the order was initiated. Once a metaorder is executed, market impact decays slowly (over several days), possibly incompletely, but may reverse. This has implications for estimating trading costs.

[0091] For example, splitting a large amount of capital for a target portfolio into several trades will result in autocorrelation in the execution prices of those trades. Ignoring the slow decay of market impact for past trades can lead to underestimation of transaction costs and poor trade scheduling decisions. To illustrate how market impact affects orders, consider a large institutional metaorder to purchase stock A where liquidity constraints prevent the entire order from being completed in one day. Instead, the order must be split evenly over two days, with an estimated market impact of 10 bps for each trade. The single-period transaction cost model assumes that the market impact fully reverts after the order is executed. That is, if stock A is purchased over two days, the total market impact cost is 50% × 10 bps + 50% × 10 bps = 10 bps. In reality, this underestimates the cost on day two, which should be greater than 10 bps.

[0092] Because the order on Day 1 increases the price of Stock A, this impact will only be reversed slightly by Day 2. For simplicity, assume that the price reversed 2 bps from the previous day before the market opened on Day 2. The true market impact cost of purchasing Stock A on Day 2 is 10 - 2 + 10 = 18 bps. Therefore, the total market impact cost of executing the entire trade is 50% × 10 + 50% × 18 bps = 14 bps, which is higher than 10 bps. The underlying cause of underestimating transaction costs is ignoring the slow decay of market impact, which leads to overoptimism in cost estimates. While it is true that splitting a metaorder over multiple periods can reduce transaction costs, assuming the market fully reverses before the next period, the benefits of order splitting are likely overestimated. When the slow decay of market impact is taken into account, as shown in the example above, order splitting may not reduce transaction costs by as much as intended.

[0093] Regarding instantaneous market impact, in some cases, the impact of large trades on stock market prices has been found to follow a concave power function of trade size relative to average daily trading volume, as shown in various literature. In some cases, a concave market impact has been observed that roughly matches the square root formula. Furthermore, market impact can be studied from two different perspectives. The first perspective addresses the impact on the price formation process when a metaorder is executed. This impact, commonly referred to as temporary market impact, is an important explanatory variable for price discovery. Temporary market impact is a major driver of trading costs, and models based on empirical measurements can be used in optimal trading schemes or by investment firms to understand their trading costs. In various embodiments, the portfolio optimization model 502 ensures that market impact increases approximately in proportion to the square root of the trade size. That is, the instantaneous market impact of a metaorder per unit for a given portfolio of market capitalization C, in dollars, can be expressed using the square root law, which can be expressed as shown in equation (9) below:

number

[0094] In some cases, empirical estimates indicate that k=10. Therefore, in some embodiments, the portfolio optimization model 502 may include k as a constant for all securities. If the constant k is sensitive to the portfolio market capitalization C, then k is set to 1.0 and the remainder is the market impact sensitivity γ 2It is understood that in some embodiments, the choice of k is not important in the portfolio optimization network 502 because a portfolio manager-defined trading cost sensitivity (γ2) calibrated to the universe of traded securities can be used. What is important in such cases is the nonlinear functional dependence on trading capital as a percentage of daily trading volume.

[0095] Market impact can also be related to the persistence of price shifts after a metaorder is fully executed, referred to as long-term market impact, which reflects the price retracement after metaorders are executed. Long-term market impact has been found to be a square-root function of the trading duration. Furthermore, market impact decay can occur quickly after a metaorder is completed. By the end of the same day, decay averages two-thirds of the peak impact, but decay continues into the next day following a power-law function on short timescales, converging to half the impact at the end of the first day over a longer period of approximately 50 days. Similar behavior has been observed, i.e., market impact slowly converges to a fraction of the first-day impact (the "permanent" impact), as shown in various literature studies. This slow decay of market impact can significantly increase purchase costs for subsequent trades and, therefore, can be important for modeling when optimizing portfolios.

[0096] Taking the above into account regarding long-term market impact and decay, in various embodiments, the portfolio optimization network 502 may be calibrated to express the decay of the market impact of a metaorder executed at time t and the long-term market impact measured at time t′ as shown in equation (10) below:

number

[0097] In some embodiments of the present invention, the permanent impact can be considered 0 for simplicity (the return occurs slower than the number of periods considered), and η = 0.05 is chosen so that the impact is 50% of its original value after approximately 10 days, which is consistent with empirical observations. In some cases, it has been observed that stock prices accumulate rapidly as trading continues and do not return to their permanent price by the second day. The longer the trading period, the higher the price of the asset, which in turn increases the cost of trading. Furthermore, the longer the trading period, the longer it takes for the price to return to its permanent price, which represents a slow decay of the extended market impact. Therefore, the total market impact can be expressed as shown in Equation (11) below:

number

[0098] With respect to daily trading volume, some forecast of average daily volume can be used to calculate trading costs. Unlike trading costs, which are part of the objective function portfolio optimization network 502, forecasts of volume can be used as feeder models 501. In some embodiments, volume can be forecast using approaches including moving averages or complex standalone volume models. The portfolio optimization network 502 can work with any type of volume model as the feeder model 501, and the output is either deterministic volume data or a forecast of a particular distribution of volume.

[0099] As described in this disclosure, for example, with respect to FIG. 3 and with respect to equations (1)-(4), the portfolio optimization model 502 uses an optimization / fitting process to optimize inputs 503 provided to the portfolio optimization model 502 and provide updated and enhanced parameters, such as multi-period weights 518, that are used for more accurate portfolio forecasting. As described above, the process 500 involves constructing a neural network loss function with user-specified market inputs (return, risk, volume) that define the portfolio objective function, and optionally combining this with any user-specified soft constraints. The model then propagates and backpropagates through the network multiple times (epochs), corresponding to one iteration of the portfolio optimization model 502, using batches of 1 x0. In some embodiments, the maximum number of epochs and early termination criteria based on the loss function default to reasonable values ​​but can be adjusted to trade off optimization time and accuracy. The model has converged when the loss function no longer significantly decreases with each epoch, fitting the optimal weights W from x0. * We discovered trainable parameters for a neural network that allow for a mapping to

[0100] Two other important points about the optimization model 502 are worth noting here. First, despite the portfolio optimization network 502 using a neural network, the optimization model is unique in that the training / learning objective of the portfolio optimization network 502 is not to achieve a resulting neural network that performs well out-of-sample. Rather, the focus of the optimization model 502 is the fitting process, as it provides weights that optimize the objectives / constraints. Second, the fitting process is performed not only for any changes to the objective function and constraints (e.g., adding constraints, changing risk sensitivity, etc.), but also for the underlying data (e.g., market data including returns, risk, etc.), as these together define the neural network's loss function and therefore change the problem and, in turn, any changes to the optimization problem.

[0101] Although FIG. 5 illustrates an example of a portfolio optimization model process 500, various modifications may be made to FIG. 5 . For example, various components and functions of FIG. 5 may be combined, further subdivided, duplicated, or rearranged according to particular needs. Also, one or more additional components and functions may be included as needed or desired. Computing systems, machine learning systems, and related processes may be provided in a wide variety of configurations, and FIG. 5 does not limit the disclosure to any particular computing system, machine learning system, or related process. For example, the steps of process 500 illustrated in FIG. 5 may be depicted as being performed on a single electronic device, such as one of electronic devices 102a-102d, by one or more cloud servers, such as application server 106 and / or one or more of database servers 108, or by a combination of electronic devices and servers in a distributed environment. For example, the model input 503 provided to the portfolio optimization network 502 may be provided by an electronic device that executed the feeder model 501, such as one of the electronic devices 102a-102d, and the input 503 is sent over a network to one or more remote servers, such as the application server 106 and / or the database server 108, which execute the portfolio optimization network 502 and provide an output 516 that is later sent back to the electronic device that provided the input 503.

[0102] 6A and 6B illustrate an example method 600 for performing a training and optimization process using an optimization model according to the present disclosure. For ease of explanation, the method 600 illustrated in FIGS. 6A and 6B is described as being performed using a processor of an electronic device, such as using one or more applications 112 executed by the application server 106 and / or the database server 108 of FIG. 1, where the application server 106 and the database server 108 may be implemented using one or more devices 200 of FIG. 2. However, the method 600 may be performed in any other suitable system(s) using any other suitable device(s), such as electronic devices like electronic devices 102a-102d.

[0103] At block 602, a processor of the electronic device defines an optimization problem for optimization to be performed using an optimization model, such as optimization model 302, optimization network 402, and / or optimization network 502 of the present disclosure. For example, the optimization problem may be a variety of optimization problems, such as a multi-period optimization problem, an alpha harvesting problem, a rebalancing problem (such as a portfolio rebalancing problem), a replication problem (such as a portfolio replication problem), etc.

[0104] At decision block 604, the processor determines whether final weight parameters have been or will be received as inputs to the optimization model. In various embodiments of the present disclosure, the final weight parameters may be provided depending on the type of optimization problem to be solved by the optimization model. For example, if final weight parameters are not provided, an optimization problem such as multi-period portfolio optimization may become an alpha harvesting problem, which seeks excess returns for a given forecast. If final weight parameters (e.g., target portfolio weights) are provided, the optimization problem may become a rebalancing problem, where a final result (the final weight parameters) is given, but the path to get there is found by the optimization model, which is implemented over multiple periods to reduce trading costs.

[0105] If, at decision block 604, the processor determines that final weight parameters have been or will be received as inputs to the optimization model, then method 600 moves to block 606. At block 606, the processor updates the optimization problem to be solved by the optimization model and / or the constraints used, and the process moves to decision block 608. If, at decision block 604, the processor determines that final weight parameters have not been or will not be received as inputs to the optimization model, then method 600 moves to decision block 608. At decision block 608, the processor determines whether to set hyperparameters for the optimization model.

[0106] For example, such hyperparameters may include at least one of whether the first layer of the optimization model is a densely connected layer or a recurrently connected layer, the number of hidden layers in the first layer of the optimization model, the number of units per hidden layer in the first layer of the optimization model, and the maximum number of epochs for training purposes. These hyperparameters may also include setting a kernel initializer, a bias initializer, whether to use a bias, an intermediate activation function, an initial learning rate(s) for the optimization model, an exponential decay rate for the first moment estimation performed by the optimization model, an exponential decay rate for the second moment estimation performed by the optimization model, a numerical stability constant (e.g., a very small number to prevent division by zero in the implementation), and / or a minimum delta value for an early termination convergence criterion.

[0107] If, at decision block 608, the processor determines that hyperparameters should not be set, method 600 moves to block 610, where the processor sets default hyperparameters. For example, some default hyperparameters may include, by way of example only, setting the optimization model to use four hidden layers and densely connected layers with 16 units per layer. In various embodiments, these defaults can be changed, such as if it is determined over time that the optimization problems most frequently applied using the optimization model utilize particular parameters, which are set to the optimization model's default parameters. In some embodiments, default parameters may be stored for different optimization problems, different feeder models, etc. Method 600 moves from block 610 to block 614. If, at decision block 608, the processor determines that hyperparameters should be set, method 600 moves to block 612, where the processor sets the received selected hyperparameters, for example, based on user input, transmission, or other interaction with the electronic device. Method 600 then moves to block 614.

[0108] At block 614, the processor receives a plurality of inputs including domain parameters and initial weights. For example, when the optimization model is performing a training and optimization process that includes performing multi-period portfolio optimization, the plurality of inputs may include return data, risk data, volume data, and initial weights from a feeder model. At decision block 616, the processor determines whether to perform one or more data consistency checks on the input data received at block 614. If at decision block 616 the processor determines that one or more consistency checks should be performed on the input data, method 600 moves to block 618. At block 618, the processor performs one or more data consistency checks on the input data. For example, in some embodiments, the processor may perform at least one data consistency check including one or more of: checking that the predicted returns and volumes are within certain ranges, checking that a variance-covariance matrix associated with the risk data is positive semi-definite, checking that a variance-covariance matrix associated with the risk data is positive definite, checking for consistency in the number of securities, checking for validity of volume data, checking for feasibility of one or more constraints, and / or performing an infinity check on a plurality of inputs. Method 600 then moves to block 620.

[0109] If, at decision block 616, the processor determines that one or more consistency checks should not be performed on the input data, method 600 moves to block 620. At block 620, the processor provides multiple inputs to the optimization model. At block 622, the processor performs a training and optimization process based on the multiple inputs and based on a training objective using a first layer of the optimization model. In some embodiments, this includes performing model fitting based on the multiple inputs to adapt the neural network by backpropagation. For example, in some embodiments, performing model fitting includes adapting a parameter learning rate in real time based on an average of the first and second moments by calculating an exponential moving average of the gradient and a moving average of the squared gradient, controlling the decay rate of the exponential moving average of the gradient and the moving average of the squared gradient, bias-correcting one or more weight parameters, and updating the initial weights using the bias-corrected one or more weight parameters. As merely one example of the present disclosure, performing the training and optimization process may include performing a multi-period portfolio optimization, where the multiple inputs include return data, risk data, volume data, and initial weights from one or more feeder models, and the goal of the optimization model is to output multi-period weights over a defined period that maximizes the return for a given forecast based on the market impact forecast.

[0110] In block 624, the processor uses a second layer of the optimization model to perform a differencing operation on the output of the first layer. In block 626, the processor uses a third layer of the optimization model to record losses based on the training objective used by the optimization model. In block 628, the processor uses a fourth layer of the optimization model to calculate and store metrics related to the training and optimization process. In block 630, the processor uses the optimization model to output updated weights. For example, the optimization model can output multi-period weights over a defined period of time, or can output other outputs depending on the optimization problem.

[0111] For example, as described in this disclosure with respect to FIG. 3 and with respect to equations (1)-(4), the optimization model uses an optimization / fitting process to optimize inputs provided to the portfolio optimization model and provide updated extended parameters, such as multi-period weights, that are used for more accurate predictions, such as portfolio forecasts. As described in this disclosure, the optimization process can include constructing a neural network loss function using user-specified market inputs (return, risk, volume) that define the objective function, and optionally combining this with any specified soft constraints. The model then propagates and backpropagates through the network using batches of 1 x0 for a number of epochs corresponding to one iteration of the portfolio optimization model. In some embodiments, the maximum number of epochs and the early termination criteria based on the loss function default to reasonable values, but can be adjusted as one of the hyperparameters in blocks 608, 612, for example, to trade off optimization time and accuracy. The model has converged when the loss function no longer significantly decreases with each epoch, and the neural network weights W are fitted to the optimal weights W from x0. * We discovered trainable parameters for a neural network that allow for a mapping to

[0112] While the portfolio optimization model uses a neural network, the optimization model is unique in that the training / learning objective of the optimization model is not necessarily to achieve a resulting neural network that performs well out-of-sample. Rather, the focus of the optimization model is the fitting process, as it provides weights that optimize the objectives / constraints. The fitting process is also performed not only for any changes to the objective function and constraints (e.g., adding constraints, changing risk sensitivity, etc.), but also for the underlying data (e.g., market data including returns, risk, etc.), as these together define the neural network's loss function and therefore modify the problem, and thus any changes to the optimization problem. Method 600 ends at block 632.

[0113] 6A and 6B illustrate an example of a method 600 for performing a training and optimization process using an optimization model, although various modifications may be made to FIGS. 6A and 6B. For example, while shown as a series of steps, the various steps in FIGS. 6A and 6B may overlap, occur in parallel, occur in a different order, or occur any number of times. For example, a decision block 604 may occur before block 602, thereby determining whether final weight parameters should be used before defining the optimization problem in block 602. Similarly, in some embodiments, a block 614 for receiving multiple inputs may occur before block 602, thereby considering the received inputs when defining the optimization problem to be solved using the optimization model. Furthermore, defining the optimization problem in block 602 may also include setting one or more constraints, as described in various embodiments of the present disclosure.

[0114] In one exemplary embodiment, a method includes receiving a plurality of inputs including domain parameters and initial weights; providing the plurality of inputs to an optimization model; using a first layer of the optimization model to perform a training and optimization process based on the plurality of inputs and based on a training objective; using a second layer of the optimization model to perform a differencing operation on an output of the first layer; using a third layer of the optimization model to record a loss based on a training objective used by the optimization model; using a fourth layer of the optimization model to calculate and store metrics related to the training and optimization process; and using the optimization model to output updated weights.

[0115] In one or more of the above examples, the method further includes setting one or more hyperparameters of the optimization model, the one or more hyperparameters including at least one of whether a first layer of the optimization model is a densely connected layer or a recurrently connected layer, the number of hidden layers in the first layer of the optimization model, the number of units per hidden layer in the first layer of the optimization model, and a maximum number of epochs for training purposes.

[0116] In one or more of the above examples, performing the training and optimization process using the first layer of the optimization model includes performing model fitting based on a plurality of inputs to adapt the neural network by backpropagation.

[0117] In one or more of the above examples, performing the model fitting includes adapting a parameter learning rate in real time based on an average of the first moment and the second moment by calculating an exponential moving average of the gradient and a moving average of the squared gradient; controlling a decay rate of the exponential moving average of the gradient and the moving average of the squared gradient; bias-correcting one or more weight parameters; and updating the initial weights using the bias-corrected one or more weight parameters.

[0118] In one or more of the above examples, the plurality of inputs further includes final weight parameters, and performing the training and optimization process includes performing a rebalancing process in which a path for achieving the final weight parameters is determined by the optimization model.

[0119] In one or more of the above examples, performing the training and optimization process includes performing a multi-period portfolio optimization, where the multiple inputs include return data, risk data, volume data, and initial weights from a feeder model, and the optimization model outputs multi-period weights over a defined period of time.

[0120] In one or more of the above examples, the multi-period portfolio optimization maximizes returns for a given forecast based on market impact forecasts.

[0121] In one or more of the above examples, the method further includes performing at least one data consistency check on at least one input of the plurality of inputs, the at least one data consistency check including one or more of: checking that predicted returns and volume are within certain ranges; checking that a variance-covariance matrix associated with the risk data is positive semi-definite; checking that a variance-covariance matrix associated with the risk data is positive definite; checking consistency of the number of securities; checking validity of volume data; checking feasibility of one or more constraints; and performing an infinity check on the plurality of inputs.

[0122] In another exemplary embodiment, an apparatus comprises at least one processor supporting optimization, the at least one processor being configured to receive a plurality of inputs including domain parameters and initial weights; provide the plurality of inputs to an optimization model; use a first layer of the optimization model to perform a training and optimization process based on the plurality of inputs and based on a training objective; use a second layer of the optimization model to perform a differencing operation on an output of the first layer; use a third layer of the optimization model to record a loss based on the training objective used by the optimization model; use a fourth layer of the optimization model to calculate and store metrics related to the training and optimization process; and use the optimization model to output updated weights.

[0123] In one or more of the above examples, the at least one processor is further configured to set one or more hyperparameters of the optimization model, the one or more hyperparameters including at least one of whether a first layer of the optimization model is a densely connected layer or a recurrently connected layer, the number of hidden layers in the first layer of the optimization model, the number of units per hidden layer in the first layer of the optimization model, and a maximum number of epochs for training purposes.

[0124] In one or more of the above examples, to perform the training and optimization process using the first layer of the optimization model, the at least one processor is further configured to perform model fitting based on the plurality of inputs to adapt the neural network by backpropagation.

[0125] In one or more of the above examples, to perform the model fitting, the at least one processor is further configured to: adapt a parameter learning rate in real time based on an average of the first moment and the second moment by calculating an exponential moving average of the gradient and a moving average of the squared gradient; control a decay rate of the exponential moving average of the gradient and the moving average of the squared gradient; bias-correct one or more weight parameters; and update the initial weights using the bias-corrected weight parameters.

[0126] In one or more of the above examples, the plurality of inputs further includes final weight parameters, and to perform the training and optimization process, the at least one processor is further configured to perform a rebalancing process in which a path for achieving the final weight parameters is determined by the optimization model.

[0127] In one or more of the above examples, to perform the training and optimization process, the at least one processor is further configured to perform multi-period portfolio optimization, wherein the plurality of inputs include return data, risk data, volume data, and initial weights from a feeder model, and the optimization model outputs multi-period weights over a defined period of time.

[0128] In one or more of the above examples, the multi-period portfolio optimization maximizes returns for a given forecast based on market impact forecasts.

[0129] In one or more of the above examples, the at least one processor is further configured to perform at least one data consistency check on at least one input of the plurality of inputs, the at least one data consistency check including one or more of: checking that predicted returns and volume are within certain ranges; checking that a variance-covariance matrix associated with the risk data is positive semi-definite; checking that a variance-covariance matrix associated with the risk data is positive definite; checking for consistency of the number of securities; checking for validity of volume data; checking for feasibility of one or more constraints; and an infinity check on the plurality of inputs.

[0130] In another exemplary embodiment, a non-transitory computer-readable medium includes instructions supporting optimization that, when executed, cause at least one processor to receive a plurality of inputs including domain parameters and initial weights; provide the plurality of inputs to an optimization model; use a first layer of the optimization model to perform a training and optimization process based on the plurality of inputs and based on a training objective; use a second layer of the optimization model to perform a differencing operation on an output of the first layer; use a third layer of the optimization model to record a loss based on a training objective used by the optimization model; use a fourth layer of the optimization model to calculate and store metrics related to the training and optimization process; and use the optimization model to output updated weights.

[0131] In one or more of the above examples, the non-transitory computer-readable medium further includes instructions that, when executed, cause the at least one processor to set one or more hyperparameters of the optimization model, the one or more hyperparameters including at least one of whether a first layer of the optimization model is a densely connected layer or a recurrently connected layer, the number of hidden layers in the first layer of the optimization model, the number of units per hidden layer in the first layer of the optimization model, and a maximum number of epochs for training purposes.

[0132] In one or more of the above examples, to perform the training and optimization process using the first layer of the optimization model, the non-transitory computer-readable medium further includes instructions that, when executed, cause the at least one processor to perform model fitting based on the plurality of inputs to adapt the neural network by backpropagation; and to perform the model fitting, the instructions, when executed, further cause the at least one processor to adapt a parameter learning rate in real time based on an average of the first and second moments by calculating an exponential moving average of the gradient and a moving average of the squared gradient; control a decay rate of the exponential moving average of the gradient and the moving average of the squared gradient; bias-compensate one or more weight parameters; and update the initial weights using the bias-compensated weight parameter or parameters.

[0133] In one or more of the above examples, to perform the training and optimization process, the non-transitory computer-readable medium further includes instructions that, when executed, cause the at least one processor to perform multi-period portfolio optimization, wherein the plurality of inputs include return data, risk data, volume data, and initial weights from a feeder model, and the optimization model outputs multi-period weights over a defined period, and wherein the multi-period portfolio optimization maximizes returns for a given forecast based on a market impact forecast.

[0134] Although the present disclosure has been described with exemplary embodiments, various changes and modifications may be suggested to those skilled in the art. The present disclosure is intended to cover such changes and modifications as fall within the scope of the appended claims.

Claims

1. 1. A method comprising: receiving a plurality of inputs including domain parameters and initial weights; providing the plurality of inputs to an optimization model; performing a training and optimization process based on the plurality of inputs and based on a training objective using a first layer of the optimization model; performing a differencing operation on the output of the first layer using a second layer of the optimization model; using a third layer of the optimization model to record losses based on the training objectives used by the optimization model; using a fourth layer of the optimization model to calculate and store metrics related to the training and optimization process; outputting updated weights using the optimization model; A method comprising:

2. setting one or more hyperparameters of the optimization model; and wherein the one or more hyperparameters are: whether the first layer of the optimization model is a densely connected layer or a recurrently connected layer; the number of hidden layers in the first layer of the optimization model; the number of units per hidden layer of the first layer of the optimization model; and Maximum number of epochs for the training purpose The method of claim 1 , comprising at least one of:

3. 2. The method of claim 1 , wherein performing the training and optimization process using the first layer of the optimization model comprises performing model fitting based on the plurality of inputs to adapt a neural network by backpropagation.

4. Performing the model fitting includes: Adapting the parameter learning rate in real time based on the average of the first and second moments by calculating an exponential moving average of the gradients and a moving average of the squared gradients; controlling the decay rates of the exponential moving average of the gradient and the moving average of the squared gradient; bias-correcting one or more weight parameters; updating the initial weights using the bias-corrected one or more weight parameters; The method of claim 3, comprising:

5. 2. The method of claim 1 , wherein the plurality of inputs further includes final weight parameters, and wherein performing the training and optimization process includes performing a rebalancing process in which a path for achieving the final weight parameters is determined by the optimization model.

6. 2. The method of claim 1, wherein performing the training and optimization process includes performing a multi-period portfolio optimization, wherein the plurality of inputs include return data, risk data, volume data, and the initial weights from a feeder model, and wherein the optimization model outputs multi-period weights over a defined period of time.

7. The method of claim 6 , wherein the multi-period portfolio optimization maximizes returns for a given forecast based on market impact forecasts.

8. and performing at least one data consistency check on at least one input of the plurality of inputs, the at least one data consistency check comprising: Checking that projected returns and volumes are within specified ranges; checking that a variance-covariance matrix associated with said risk data is positive semi-definite; checking that the variance-covariance matrix associated with the risk data is positive definite; Checking the consistency of the number of securities; checking the validity of said volume data; Checking the feasibility of one or more constraints; and performing an infinity check on said plurality of inputs; The method of claim 6, comprising one or more of:

9. 1. An apparatus comprising: at least one processor supporting optimization, said at least one processor comprising: receiving a plurality of inputs including domain parameters and initial weights; providing the plurality of inputs to an optimization model; performing a training and optimization process based on the plurality of inputs and based on a training objective using a first layer of the optimization model; performing a differencing operation on the output of the first layer using a second layer of the optimization model; using a third layer of the optimization model to record losses based on the training objectives used by the optimization model; using a fourth layer of the optimization model to calculate and store metrics related to the training and optimization process; outputting updated weights using the optimization model; An apparatus configured to:

10. The at least one processor Setting one or more hyperparameters of the optimization model and wherein the one or more hyperparameters are whether the first layer of the optimization model is a densely connected layer or a recurrently connected layer; the number of hidden layers in the first layer of the optimization model; the number of units per hidden layer of the first layer of the optimization model; and Maximum number of epochs for the training purpose 10. The apparatus of claim 9, comprising at least one of:

11. 10. The apparatus of claim 9, wherein to perform the training and optimization process using the first layer of the optimization model, the at least one processor is further configured to perform model fitting based on the plurality of inputs to adapt a neural network by backpropagation.

12. To perform the model fitting, the at least one processor: Adapting the parameter learning rate in real time based on the average of the first and second moments by calculating an exponential moving average of the gradients and a moving average of the squared gradients; controlling the decay rates of the exponential moving average of the gradient and the moving average of the squared gradient; bias-correcting one or more weight parameters; updating the initial weights using the bias-corrected one or more weight parameters; The apparatus of claim 11 , further configured to:

13. 10. The apparatus of claim 9, wherein the plurality of inputs further includes final weight parameters, and to perform the training and optimization process, the at least one processor is further configured to perform a rebalancing process in which a path for achieving the final weight parameters is determined by the optimization model.

14. 10. The apparatus of claim 9, wherein to perform the training and optimization process, the at least one processor is further configured to perform multi-period portfolio optimization, wherein the plurality of inputs include return data, risk data, volume data, and the initial weights from a feeder model, and the optimization model outputs multi-period weights over a defined period of time.

15. The apparatus of claim 14 , wherein the multi-period portfolio optimization maximizes returns for a given forecast based on market impact forecasts.

16. The at least one processor is further configured to perform at least one data integrity check on at least one input of the plurality of inputs, the at least one data integrity check comprising: Checking that projected returns and volumes are within specified ranges; checking that the variance-covariance matrix associated with said risk data is positive semi-definite; checking that the variance-covariance matrix associated with the risk data is positive definite; Checking the consistency of the number of securities, checking the validity of said volume data; Checking the feasibility of one or more constraints; and Infinity check for multiple inputs 15. The apparatus of claim 14, comprising one or more of:

17. 1. A non-transitory computer-readable medium comprising instructions to support optimization, the instructions, when executed, causing at least one processor to: receiving a plurality of inputs including domain parameters and initial weights; providing the plurality of inputs to an optimization model; performing a training and optimization process based on the plurality of inputs and based on a training objective using a first layer of the optimization model; performing a differencing operation on the output of the first layer using a second layer of the optimization model; using a third layer of the optimization model to record losses based on the training objectives used by the optimization model; using a fourth layer of the optimization model to calculate and store metrics related to the training and optimization process; outputting updated weights using the optimization model; A non-transitory computer-readable medium for causing

18. When executed, the at least one processor: further comprising instructions for setting one or more hyperparameters of the optimization model, the one or more hyperparameters comprising: whether the first layer of the optimization model is a densely connected layer or a recurrently connected layer; the number of hidden layers in the first layer of the optimization model; the number of units per hidden layer of the first layer of the optimization model; and Maximum number of epochs for the training purpose 20. The non-transitory computer-readable medium of claim 17, comprising at least one of:

19. To perform the training and optimization process using the first layer of the optimization model, the non-transitory computer-readable medium further includes instructions that, when executed, cause the at least one processor to perform model fitting based on the plurality of inputs to adapt a neural network by backpropagation, and to perform the model fitting, the instructions, when executed, cause the at least one processor to: Adapting the parameter learning rate in real time based on the average of the first and second moments by calculating an exponential moving average of the gradients and a moving average of the squared gradients; controlling the decay rates of the exponential moving average of the gradient and the moving average of the squared gradient; bias-correcting one or more weight parameters; updating the initial weights using the bias-corrected one or more weight parameters; 20. The non-transitory computer-readable medium of claim 17, further comprising:

20. 18. The non-transitory computer-readable medium of claim 17, wherein the non-transitory computer-readable medium further includes instructions that, when executed, cause the at least one processor to perform a multi-period portfolio optimization, wherein the plurality of inputs include return data, risk data, volume data, and the initial weights from a feeder model, the optimization model outputting multi-period weights over a defined time period, and the multi-period portfolio optimization maximizing returns for a given forecast based on a market impact forecast.