A method and system for generating generalised additive models (GAMS)

The method automatically generates well-conditioned Generalised Additive Models (GAMs) that balance accuracy with interpretability, smoothness, and monotonicity, addressing the challenges faced by existing models in risk prediction tasks.

WO2025114810A1PCT designated stage expired Publication Date: 2025-06-05PORCUPINE UNION (PTY) LTD

Patent Information

Application Number
PCT/IB2024/061551
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-28
Filing Date
2024-11-19
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

Existing machine learning and artificial intelligence models struggle to balance accuracy with interpretability, smoothness, and monotonicity, which are crucial for regulatory compliance and trust in risk prediction tasks, particularly in finance and insurance.

Method used

A computer-implemented method and system for automatically generating Generalised Additive Models (GAMs) that are both accurate and well-conditioned, by using a loop-based process with steps for loss minimization, automatic conditioning of weights, and model convergence checking.

Benefits of technology

The method produces GAMs that are not only accurate but also interpretable, smooth, and monotonic, enhancing model transparency and robustness, which is essential for regulatory compliance and trust in predictive models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2024061551_05062025_PF_FP_ABST
    Figure IB2024061551_05062025_PF_FP_ABST
Patent Text Reader

Abstract

A method for automatically producing an accurate and well-conditioned Generalised Additive Model (GAM) comprises receiving a dataset which will be used to fit the GAM. The dataset comprises rows and columns containing data. Each row containing at least one or more different predictions using the GAM and one or more features which are values that the GAM uses to produce a prediction. The GAM is initialised by creating starting model baseline weights including a vector and an intercept. The GAM is then constructed by executing a loop with plural steps, in which a loss function is minimised with respect to epoch weight. A baseline commit step includes adding a fraction of the conditioned epoch weights to the model baseline weights to produce updated model baseline weights. Model convergence checking determines whether the changed baseline weights are sufficiently accurate and well conditioned to stop the loop, thereby yielding the final GAM.
Need to check novelty before this filing date? Find Prior Art

Description

[0001]A Method and System for Generating Generalised Additive Models (GAMs) FIELD OF INVENTION The invention relates to a computer-implemented method and system for generating Generalised Additive Models (GAMs). BACKGROUND OF INVENTION In industries such as banking and insurance, there is perhaps no greater challenge than to accurately predict the risk associated with doing business with prospects, for example: ^ in the short-term motor insurance industry, the insurance company must accurately predict the future frequency and severity of the perilous events that will occur to a prospect; and ^ in the banking industry, the bank that intends to loan money to a prospect needs to accurately predict the probability that the prospect will default on their loan obligation. It is known that computerised models can be constructed from a set of prior observations to predict an outcome. There has, in recent years, been a surge in methods to construct such models, spanning the fields of Machine Learning (ML), Artificial Intelligence (AI), and classical statistical methods. Before these models can be used, they must be evaluated and declared fit for service. The criterion by which many ML and AI models are evaluated is their accuracy. There are many applications, however, where accuracy is not the only desirable trait for a model to have. As an example, there may be three other traits against which a model can be evaluated: • Interpretability: It is often difficult or impossible to explain how a prediction made by an ML- or AI-based model is related to the prediction’s input features. In sectors like financial services, where models and their accompanying predictions undergo scrutinisation by auditors, this presents a significant disadvantage. Regulatory bodies and auditors require transparency to ensure fairness and compliance. There are a couple of interpretability techniques, such as: o Feature Importance or feature attribution techniques: Identifying which input features contribute most to a model’s predictions. Techniques like feature importance scores (e.g., SHAP values) help highlight salient features. o Partial Dependence Plots: Visualizing how a model’s output changes with variations in a single feature while keeping other features constant. o Simpler Models: Using interpretable models (e.g., decision trees, linear regression) as proxies for complex models. o Rule-Based Systems: Representing model decisions as a set of rules (e.g., CASE or IF-THEN statements). o Built-in methods that form part of the model architectures such as the attention mechanisms found in models utilising neural network-based transformers and accompanying attention masks. Interpretability enables trust, accountability, and regulatory compliance in computerised models. It empowers practitioners to make informed decisions, especially in sectors where transparency matters most. • Smoothness: A smooth model is a model that produces similar results for feature sets that are alike. Smoothness refers to the regularity or continuity of a function or model. It describes how consistent or predictable the function is across its domain Smoothness is desirable because it simplifies analysis, optimization, and interpretation of models. A hypothetical example considers two assets having the same feature sets except for their age. The model must predict the number of adverse events this asset is going to experience in the next timeframe. If their age differs by only one year, it is expected that their predictions will be similar, differing only by a small amount. In practice, however, the Applicant has experienced that ML and AI models can produce wildly different predictions for these two feature sets. This is unsatisfying and can lead to brand damage for the entity making decisions based on this model. It can also exacerbate model distrust. • Monotonicity: This refers to the property of a model producing monotonically increasing or decreasing predictions as a certain ordinal or numerical feature in the input space of the model increases in value. A popular example would be the effect of excess on insurance premium. If a quote has been obtained, and the insurance prospect decides to increase their excess payment on their policy, the recalculated quoted premium must, without exception, decrease. A model or ensemble / combination of models that does not adhere to this monotonicity requirement may be perceived as unfair and illogical, potentially resulting in brand damage. The three traits mentioned above can be categorised as traits pertaining to model conditioning. A well-conditioned model would typically perform well when evaluated against all three of the above-mentioned traits, as well as accuracy. Conditioning GAMs around interpretability, smoothness, and monotonicity promotes model transparency, robustness, and alignment with domain knowledge. GAMs have simple mathematical representations, and as such they can be constructed in a well- conditioned manner. Even though GAMs have simple mathematical representations, it is an immense technical challenge to determine the parameters of said representation. Traditionally, GAMs often require tedious manual intervention, feature selection and setting of hyperparameters. SUMMARY OF INVENTION The present invention provides a computer-implemented method for producing an accurate and well-conditioned Generalised Additive Model (GAM) automatically, the method implemented by a computer system comprising a processor, the method comprising: receiving, by a data parsing module provided by the processor, a dataset, which will be used to fit the GAM, wherein: the dataset comprises rows and columns containing likely heterogeneous tabular data, categorical, ordinal or a combination thereof in nature; each row contains at least two kinds of values, namely (1) one or more different observations to be predicted using the GAM and (2) one or more features which are values that the GAM uses to produce a prediction for an observation; and a single dataset row in the dataset therefore contains different features and observations associated with it; initialising the GAM, by a modelling engine provided by the processor, by creating a starting model baseline weights including a vector and an intercept; and constructing, by the modelling engine, the GAM by executing a loop with plural steps, each time the loop has been completed, an epoch has occurred, the plural steps comprising: a loss step, in which a loss function is minimised with respect to epoch weights, including vector and an intercept, of the GAM, resulting in a set of epoch weights; epoch weights automatic conditioning, in which the resulting epoch weights are automatically changed to improve the conditioning of the epoch weights; baseline commit step, in which a fraction of the conditioned epoch weights is added to the model baseline weights to produce updated model baseline weights; baseline weights automatic conditioning, in which the updated model baseline weights are automatically changed to improve the conditioning of the GAM; and model convergence checking, in which an automatic check is performed to determine whether the changed baseline weights are sufficiently accurate and well-conditioned to stop the loop, thereby yielding the final GAM. The modelling engine may be configured to repeat the loop until a final condition (or exit condition) is fulfilled. The method may be implemented by the computer system including a computer program configured to direct the operation of the computer processor. The invention extends to a computer system including at least one computer processor and a computer program configured to direct the operation of the computer, the computer system being configured to implement the method as defined above. The invention extends to a non-transitory computer-readable medium having stored thereon a computer program which, when executed by at least one computer processor, implements the method as defined above. BRIEF DESCRIPTION OF DRAWINGS The invention will now be further described, by way of example, with reference to the accompanying diagrammatic drawings. In the drawings: FIG.1 shows a schematic diagram of a computer system for producing an accurate and well-conditioned Generalised Additive Model (GAM) automatically, in accordance with the invention; FIG.2 shows a flow diagram of a method of producing an accurate and well- conditioned Generalised Additive Model (GAM) automatically, in accordance with the invention; FIG.3 shows a schematic view of a computer system within which a computer program, for causing the computer system to perform any one or more of the methodologies discussed herein, may be executed; FIG.4 shows a chart of epoch feature relative weights as a function of the feature's unique values, as part of the method of FIG.2; FIG.5 shows a chart of model weights from feature 1 before and after conditioning using a median filter, as part of the method of FIG.2; FIG.6 shows a chart of weight values vs the unique values for feature 1 before and after conditioning using a Savitsky-Golay filter, as part of the method of FIG. 2; FIG.7 shows a chart of the weight values vs the unique values for feature 1 before and after implementing a monotonically increasing conditioning function, as part of the method of FIG.2; FIG.8 shows a chart of the weight values vs the unique values for feature 1 before and after implementing a +-+ turning point conditioning function, as part of the method of FIG.2; FIG.9 shows a chart of the baseline weight values vs the unique values for feature 1 after committing the conditioned epoch weights after epoch 0, 5, 15 and 30, as part of the method of FIG.2; and FIG.10 shows a chart of the baseline weight values vs the unique values for feature 1 before and after conditioning the baseline weights using a turning point monotonicity conditioning function, as part of the method of FIG.2. DETAILED DESCRIPTION OF EXAMPLE EMBODIMENT(S) The following description of an example embodiment of the invention is provided as an enabling teaching of the invention. Those skilled in the relevant art will recognise that changes can be made to the example embodiment described, while still attaining the beneficial results of the present invention. It will also be apparent that some of the desired benefits of the present invention can be attained by selecting some of the features of the example embodiment without utilising other features. Accordingly, those skilled in the art will recognise that modifications and adaptations to the example embodiment are possible and can even be desirable. Thus, the following description of the example embodiment is provided as illustrative of the principles of the present invention and not a limitation thereof. FIG.1 illustrates a computer system 100 in accordance with the invention. Although the computer system 100 is illustrated as a single device or system, it may well be distributed among a number of devices, e.g., being cloud-based. The computer system 100 has a computer processor 110 coupled to a computer-readable medium 120. A computer program is stored on the computer-readable medium 120 and is configured to direct the operation of the processer 110 when executed thereon. More specifically, the computer program provides a modelling engine 122 to construct a Generalised Additive Model (GAM). The modelling engine 122 may be a conceptual module corresponding to a functional task performed by the processor 110. It is to be understood that the processor 110 may be one or more microprocessors, controllers, Digital Signal Processors (DSPs), Field Programmable Gate Array Devices (FPGAs) or any other suitable computing device, resource, hardware, software, or embedded logic. The computer-readable medium 120 may be main memory and / or a storage device. The computer-readable medium 120 further provides a data parsing module 124 which is configured to communicate data between the modelling engine and a source or destination, like a database 114. The computer system includes a communications module 112 (e.g., a network interface) and is also communicatively coupled to a database 114. The database 114 may have stored thereon one or more datasets 130 and / or one or more GAMs 132. The dataset 130 may be communicated to the database 114 via the communications module 112 and the constructed GAM 132 may be communicated from the database 114 again via the communication module 112. An option is to provide the computer system 100 as a service, in which a client may provide the dataset 130 and receive the constructed GAM 132 in return. FIG.2 illustrates a method 200 of constructing the GAM 132, in accordance with the invention. The method 200 will be described below with reference to detailed examples. Detailed description of the step of Receiving a dataset (block 202 of FIG.2) The purpose of constructing a GAM is to predict an observation. There are many kinds of observations that one may want to predict. Some illustrative examples are: • The probability that a person will be involved in a motor vehicle accident for which the person has taken out an insurance policy. • The probability that a person will pay or have a payment default for their insurance premium for said policy. • The monetary severity of an accident, given that an accident has occurred. • The probability that a claim lodged by a person against said policy is a fraudulent claim. • The probability that a claim is to be paid out after logging of said claim. This can be influenced by the voluntary excess amount, additional excess values that form part of the policy terms or other such conditions of cover. To build a GAM for predicting an observation, a dataset is used. A dataset consists of rows and columns. The columns that do not contain observations are referred to as feature columns. Each row groups together several features associated with one or multiple kinds of observations. These features represent data associated with each observation. Specifically, some examples of what these features represent are given below: 1. Biological Information: 1.1. Includes details about an individual’s health, medical history, and genetic factors. 1.2. Examples: Age, gender, pre-existing conditions. 2. Environmental Information: 2.1. Pertains to the surroundings and context in which an event occurs. 2.2. Examples: Location, climate, pollution levels. 3. Meteorological Information: 3.1. Specific to weather conditions and natural events. 3.2. Examples: Average temperature, average precipitation, storm occurrences. 4. Physical Event Information: 4.1. Describes specific incidents or accidents. 4.2. Examples: previous vehicle collisions, previous property damage. 5. Geographical Information: 5.1. Relates to the geographic location of insured assets. 5.2. Examples: ZIP code, proximity to flood zones. 6. Asset Specifications: 6.1. Details about the insured property or vehicle. 6.2. Examples: Engine specifications, safety features. 7. Abstract Risk Information: 7.1. Factors that impact overall risk but may not be directly observable. 7.2. Examples: Credit score, payment history. 8. Nonphysical Event Information: 8.1. Behavioural data related to financial decisions. 8.2. Examples: Switching between providers, claim history. 9. Financial Information: 9.1. Monetary aspects associated with the insured item. 9.2. Examples: Market value, replacement cost. 10. Personal Information: 10.1. Demographic details about the insured individual. 10.2. Examples: Age, gender, marital status. This dataset has already been manipulated and engineered such that it can be used to construct the GAM. It is known that data engineering and data manipulation is important and will influence any resulting model. The present invention, however, does not require any specific way in which the data has been manipulated or prepared. This dataset may also include features that are not considered primary features for use in the GAM, but rather to perform a supporting function to the GAM scaling function. These extra features are referred to belonging to the metadata dataset. A good example of such metadata features is the data required for modelling the sensitivity of an outcome to a feature. One would have a sensitivity factor as a feature and multiply this with the weights of the GAM. The sensitivity factor is therefore part of the metadata dataset. Variations on the detail description The features mentioned above may be represented digitally by a simple data type, such as an integer, floating point value, Null value, or a string value. Each feature may also be digitally represented by a composite type, such as a JSON object, where the JSON’s keys are represented by simple data types and their corresponding values represented digitally by either simple data types as mentioned above, or composite types themselves. As mentioned above, this dataset may be prepared in any way. Examples of ways in which the dataset may be prepared are: 1. One may, for example, group the values of an entire feature column into near equal count bins and replace the feature value with a value associated with its count bin, such as the bin’s left value, right value, starting percentile, ending percentile, or the ordered index of the bin’s left value among all the other bins’ left values. 2. One may, for example, change certain feature value types based on expert knowledge. An example would be changing a credit score of -1 from a continuous / ordinal type to a categorical type because a credit score of -1 has a meaning that is not related to its ordinal position among the other credit score values. 3. One may, for example, replace certain feature values by other values. An example of this would be to replace features with a Null value with the median of the numerical values associated with that feature column. 4. One may, for example, change certain values of a feature because of there not being enough occurrences of that value to model the GAM on. An example may be changing all “home loan count” values larger than 2 to 2. The GAM would then treat all rows having “home loan count” values larger than 2 as if it has “home loan count” values equal to 2. 5. One may, for example, alter a feature that acts as an observation to be normalised according to another exposure feature. Exposure in this context normally refers to the percentage of the time period associated with each row that an insurer was responsible for ensuring a policyholder should the event represented by the observation occur. 6. One may, for example, fuse two or more feature columns together, thereby resulting in an interaction feature. An example of this would be an interaction between “age” and “vehicle type”. A person aged 18 having a “Model A” vehicle might then be represented by the fused feature “18_ Model A”. 7. One may, for example, create new features based on the geospatial distance between entities described in the dataset. 8. One may, for example, create new features based on ML techniques such as Principal Component Analysis (PCA) or clustering. 9. One may, for example, create a new feature based on specified rules applied to a combination of other features to create a bonus-malus score (https: / / en.wikipedia.org / wiki / Bonus–malus). Physical manifestation The dataset is stored on a computer-readable device. Examples of such devices are: 1. A hard disk drive; 2. A solid-state drive; 3. A flash drive; 4. A compact disk; 5. A memory embedded in a processor; 6. A distant storage accessible in the cloud; or 7. A combination of the above The act of “receiving” the dataset includes loading the dataset into an ephemeral memory device, such as DRAM, SRAM, VRAM, or any other RAM used to store data which a processor uses. The processor will be involved in this loading process. The dataset can be stored in any data format, as long as the code generating the GAM can execute its intended function on the dataset.. Examples of such formats includes, CSV, Parquet, Excel, SQLServer, PostgreSQL, DuckDB, HDF Store, Pandas or Polars DataFrame, etc. Detailed description of the step of Initialising the GAM (block 204 of FIG.2) A GAM is described by its model weights and link function. The link function is often determined based on how the observation is statistically distributed. As an example, a GAM that predicts an observation that is modelled as having a Poisson or Gamma statistical distribution can have the form described in Equation 1: ^^ೕ ∙^^ In Equation 1 above, ^ ^^^^is the GAM’s prediction for the dataset row number ^^; ^ ^^ is the exponential function; ^ ^^ represents the dataset that contains the features that have feature weights associated to them; ^ ^^ is an index that identifies a specific feature column in ^^; ^ ^^ is the total number of feature columns in ^^; ^ ^^^represents the index of unique value number ^^ in feature column number ^^; ^ ^^^,^ೕrepresents the GAM feature weight associated with unique value ^^ in feature column ^^; ^^^(^^, ^^୫, ^^, ^^, ^^) is a weight scaling function that scales the feature weight, ^^^,^ೕ fordataset row ^^; ^ ^^୫represents the metadata dataset; the metadata dataset contains zero or more feature columns that are used in the scaling function, but which do not have weights associated with them. References made to the metadata dataset are explicit. All other references to a dataset refer to ^^; ^ and ^^ represents the model intercept, which is a constant term. In the simplest terms, the action of constructing a GAM (block 210 of FIG.2) comprises setting each weight value, ^^^,^ೕ, and the intercept, ^^. In Equation 1 above, the entire set of weights can be represented together in a weight vector, as shown in Equation 2 below: ^^ ⋯ ⋯ ⋯ ⋯ ^^ The weight vector therefore contains as many values as there are unique values in the dataset. The present invention distinguishes between three different kinds of weight vectors: ^ The baseline weight vector: This vector is represented mathematically by ^^ୠୟ^^୪୧୬^. ^ The epoch weight vector: This vector is represented mathematically by ^^^୮୭ୡ୦. ^ The composite weight vector: This vector is represented mathematically by ^^ୡ୭୫୮୭^୧^^. The composite weight vector is the sum of the baseline weight vector and the epoch weight vector, Equation 3 illustrates this: ^^ୡ୭୫୮୭^୧^^ = ^^ୠୟ^^୪୧୬^ + ^^^୮୭ୡ୦ Eq. (3) The present invention is also an iterative method. Each iteration of the method loop (not to be confused with each possible iteration of the loss step), is referred to as an epoch. The weight vectors are therefore expressed as being dependant on the epoch number, ^^: ^^ୡ୭୫୮୭^୧^^,^ = ^^ୠୟ^^୪୧୬^,^ + ^^^୮୭ୡ୦,^ Eq. (4) where ^^ is the epoch Before commencing with the first epoch, the baseline weight vector needs to be set. This is referred to as baseline initialisation. Normally, the baseline would be set as a vector of zeros: ^^ୠୟ^^୪୧୬^,^ = ^ ^⃗^ Eq. (5)In Equation 5 above, ^ ^⃗^ is a vector where the weight associated with each uniquefeature has been set to zero. The model intercept is handled in the same way as the weight vector. There are three model intercept values, ^^ୠୟ^^୪୧୬^,^, ^^^୮୭ୡ୦,^, and ^^ୡ୭୫୮୭^୧^^,^. The composite intercept is calculated as follows: ^^ୡ୭୫୮୭^୧^^,^ = ^^ୠୟ^^୪୧୬^,^ + ^^^୮୭ୡ୦,^ Eq. (6) The model intercept is usually initialised such that the GAM, if it were only to predict based on the intercept, produces predictions equal to the mean of the observation column being modelled. For the link function shown in Equation 1, the starting model intercept would be selected as follows: ^^ୠୟ^^୪୧୬^,^ = ln൫mean(^^)൯ Eq. (7)In Equation 7, ln(∙) is the natural logarithm function, mean(∙) is an aggregation function that calculates the mean of a vector, and ^^ is the set of all observations in the dataset.In one example of the present invention, the scaling function, ^^(^^, ^^୫, ^^, ^^, ^^), has theform as expressed in Equation 8 below: ^^ ^1, ^^^,^ = In Equation 8 above, ^^^,^is the feature value of the ^^th feature in dataset row ^^, and ^^^ೕis the ^^th unique value of feature column number ^^. Equation 8 ensures that the weights contributing to the ^^th row’s GAM prediction are only the ones corresponding to the feature values that actually occur in the ^^th row. The set of unique values within each feature column can be expressed as done in Equation 9 below. ^^^ = ^ ^^^ೕ , ^^^ೕ , ⋯ ^^^ೕ^ Eq. (9)There is a one-to-one correspondence between the unique feature values and the feature weight vector: ^^^ = ^ ⋯ Variations on the detailed description Equation 1 shows one possible example of a GAM’s mathematical form that is typically used when the observation variable is distributed according to the Poisson, Gamma or Tweedie statistical distributions. There are many possible statistical distributions that can be used for building GAMs, some of them include: ^ Gaussian distribution ^ Poisson distribution ^ Gamma distribution ^ Tweedie distribution ^ Negative binomial distribution ^ Binomial distribution ^ Inverse Gaussian distribution There are many mathematical forms GAMs can be expressed in. In general, a GAM can be expressed in the following form: ^^^^^^ ^ ^^ ^ In Equation 11, ^^(∙) is the link function for the specific GAM, the symbol ^^^represents a constant offset that is associated with dataset row ^^. In most cases, this offset will be zero. In some formulations of GAMs, the link function is rather defined as follows: ^ Where ^^ୟ୪^(∙) These two formulations are equivalent. There are cases where ^^^can be nonzero. Examples of such cases are: ^ If it is desired to construct a GAM that serves as an improvement over an existing model. In this case, ^^^will be the existing model’s prediction, transformed into the same domain as the argument of ^^(∙). ^ If it is desired to model an outcome’s sensitivity to a feature, you may choose to fit a first model without said feature and produce initial predictions from this first model. One would then fit a second GAM while setting ^^^equal to row ^^’s prediction from the first GAM. The baseline weights may be initialised as the zero vector and the intercept is initialised to such that the GAM produces a prediction corresponding to the observation column’s mean value. However, the baseline weights and intercept can be initialised in many ways. Some examples of ways in which the baseline weights can be initialised are discussed below: ^ The baseline weights and / or model intercept can be initialised by copying values from a different GAM. During the copying process, there are many possible reasons why changing the values being copied may be desired; possible reasons are: o If the unique feature values in the dataset does not match up with the unique feature values in the GAM you are copying from; o If the GAM you are copying from uses a different link function. o If the GAM you are copying from was trained on an observation column with different statistical properties than the observation column from the current dataset. ^ The baseline weights and / or model intercept can be initialised from expert knowledge or according to industry best practices. The present invention does not claim or require any specific form for the scaling function. Some examples of different scaling functions are provided below. One example of a different scaling function is given below: ^^ ^ , ^^^,^ = In Equation 12 above, the scaling function returns the feature value in the ^^th row and the ^^th feature column in the dataset. In this case, ^^^ೕneeds to be a numeric type or needs to be converted to a numeric type prior to starting the modelling for this scaling function. Another example of the scaling function uses a penalized spline function to transform the feature value: ^^(^^, ^^୫, ^^, ^^, ^^) = ^ ^^^(^^^ೕ), ^^^,^ = ^^^ೕ0, ^^^,^ ≠ ^^^ೕEq. (13) In Equation 13, one transformation function ^^^corresponds to each feature column. The function ^^^can be a penalised regression spline function. As another example, a scaling function can be used, and then multiplied with a drop- out layer, as shown in Equation 14 below. ^^ௗ(^^, ^^୫, ^^, ^^, ^^) = ^^ ∙ ^^(^^, ^^୫, ^^, ^^, ^^) (14) In Equation 14 above, a scaling function is multiplied by a dropout variable, ^^. The dropout variable randomly sets the scaling function value, and hence the weight contribution of certain weights, to zero. Dropout is used to build models that are less prone to overfitting. The dropout variable is usually distributed according to a Bernoulli distribution. A scaling function using dropout will only be used in the loss step, and not normally the prediction step. As another example, a scaling function that returns a value from the metadata dataset can be returned, as shown in Equation 15 below: ୫^^ ^ ^^^,^ = In Equation 15 above, ^^୫^^,^^represents the value in row ^^ and column ^^ from the metadata dataset. This example is typically used with price sensitivity modelling, where a price change factor is multiplied with the feature weight to capture the effect of price sensitivity. Physical manifestation The loss step is performed using a processing unit. Examples of such a processing unit are CPU, GPU or a GPU grid, FPGA device, ASIC chip, specialised chips such as SOC’s or NOC’s, AI specialised chips such as TPU’s, etc. The processing unit executes this step based on programming code. This programming code may be located on a computer readable medium, many of which have been mentioned before. The processing unit may need to retrieve or store data from or to an ephemeral memory unit. The processing unit may need to temporarily store and retrieve data or code on a computer readable medium. This step is often performed on a single computer. This step can also be performed on several computers simultaneously. An amalgamation step may then be included. This step can be performed on a local computer, or on a computing device in the cloud where one may not have exact knowledge of the computing device’s location, ownership or architecture. Detailed description of the Loss step (block 212 of FIG.2) The Loss step is the first step performed in every epoch. The loss step starts by initialising the epoch weights to zero. The term epoch weights include a vector of feature weights and an intercept. The epoch feature weights are initialised to the zero vector and the epoch intercept to zero: =^ ^⃗^ = 0 In Equation 16, ^^ is a counter variable that keeps track of the epoch number. After the epoch weights have been initialised, the loss step continues. The loss step is concerned with minimising the GAM’s loss function. The loss function must be minimised with respect to the epoch weights, but the loss values are calculated based on the composite weights. An example of a loss function for a GAM predicting a Poisson distributed observation is shown in Equation 16 below. = ூ− ^^^ ^^^ = ∑^ ^ୀ^ ^Σ^^ೕೕୀ^^^ୡ୭୫୮୭^୧^^,^^,^ೕ ∙ ^^(^^, ^^, ^^, ^^)^ + ^^ୡ୭୫୮୭^୧^^,^ Eq. (18)In Equation 17 above, ^^ is one less than the total number of rows in the dataset. In Equation 18 above, ^^ୡ୭୫୮୭^୧^^,^^,^ೕrefers to the composite weight in epoch ^^ that is associated with the ^^th unique value in feature column ^^. Finding a set of epoch weights that minimises Equation 17 can be expressed as done in Equation 19 below. ^^^ ∗ ^^^ ∗ = In Equation 19 above ^^^∗୮୭ୡ୦and ^^^∗୮୭ୡ୦are the epoch feature weights and intercept that minimises the loss function in Equation 17. The present invention does not claim or require any specific method by which Equation 19’s result is achieved. Any optimisation algorithm can be used. Examples of some algorithms that can be used are gradient descent, stochastic gradient descent, or proximal gradient descent algorithms. Non-gradient descent methods, such as using Gradient Boosted Trees (GBMs) with a tree depth of 1, is another example. The loss step ends when the outcome specified in Equation 19 has been achieved. Practically, however, it is often impossible to know which epoch weights and intercept lead to a globally optimum solution. Therefore, an appropriate stopping criterion is relied on to decide when the result in Equation 19 has been achieved. The present invention does not claim or require any specific stopping criterion. Some examples of stopping criteria are: ^ Running the loss step for a fixed period of time; ^ Running the loss step for a fixed period of iterations; ^ Running the loss step until the change in best solution is less than a defined tolerance value; this tolerance value may be compared to each weight in the weight vector, or it may be compared to the aggregate of the weight vector. The absolute values of the weight vector are used to compare against the tolerance. ^ Running the loss step until the loss score starts to deteriorate on a different, “holdout” set. This can be achieved by training only on a limited subset of rows and keeping the remainder of the observations for determining when convergence has been achieved. Variations on the detail description The present invention does not claim or require any specific loss function. A non- exhaustive list of loss functions that may be applicable can be found in the documentation of a popular ML Library (PyTorch, 2023): ^ L1 Loss ^ Mean Squared Error Loss ^ Cross Entropy Loss ^ Connectionist Temporal Classification ^ Negative log likelihood loss ^ Negative log likelihood loss with Poisson distribution of target. This is the loss function from Equation 16 when the keyword log_input is set to True ^ Gaussian negative log likelihood loss ^ Kullback-Leibler divergence loss ^ Binary Cross Entropy loss ^ A loss function combining a sigmoid layer and the Binary Cross Entropy loss ^ Margin Ranking Loss ^ Hinge Embedding Loss ^ Multi Label Margin Loss ^ Huber Loss ^ Smooth L1 Loss ^ Soft Margin Loss ^ Multi Label Soft Margin Loss ^ Cosine Embedding Loss ^ Multi Margin Loss ^ Triplet Margin Loss ^ Triplet Margin with Distance Loss ^ Gamma loss function (not in Pytorch documentation) ^ Tweedie Loss function (not in Pytorch documentation) More complex loss functions can also be used. Sometimes, more complex loss functions aim to preserve model conditioning during optimisation. A more complex loss function will typically use a simple loss function, like one listed above, with additional penalty functions to preserve or promote model conditioning. Examples of such loss functions can be found in (US 2021 / 0365822 A1). The loss step may be concerned with minimising the loss function with respect to the epoch weights and intercept. An equivalent action would be to maximise the logloss function with respect to the epoch weights and intercept. The loss function may be calculated on all the rows in the dataset. However, the loss function may also perform this loss step on a subset of the dataset. The subset may be chosen randomly, or it may be chosen to train on a less or more balanced dataset by subsampling or oversampling the observation rows that have certain observation values. This subset may be selected based on certain feature values. For instance, an existing GAM may not predict accurately for rows with certain features. The dataset may be filtered to include only observations containing these observation rows. The epoch weights and intercept may be determined to minimise the loss function. This does not always need to be the case. The following three scenarios are also possible: ^ It may be decided to only determine the model epoch feature weights and leave the intercept at its initial value. This may be applicable when modelling on a dataset which has subsampled or oversampled data. ^ It may be decided only to determine the model intercept and leave the epoch feature weights at its initial value. This may be the case when importing a GAM from a different modelling engine, or when the imported GAM was modelled on a dataset with a different target column mean. ^ It may be decided to only determine a subset of the epoch feature weights as opposed to the entire feature weight vector. This may be appropriate when the model weights assigned to some features are deemed acceptable, but the model feature weights associated with the remaining features are not deemed acceptable. Detailed description of the Epoch weights automatic conditioning (block 214 of FIG.2) Once the optimal epoch weights and intercept have been determined in the previous step, this step performs automatic model conditioning on the epoch weights. The automatic conditioning can be expressed as a function acting on the feature weights. ^^^^ = ℎ(^^^, ^^^) Eq. (20)In Equation (20, ℎ(∙,∙) is the model conditioning function, ^^^^are the model weights for feature column ^^ after being changed by ℎ(∙,∙) , and, as already mentioned, ^^^and ^^^are the weights and unique values associated with feature column ^^. It is generally considered true that altering a model’s weights – without considering the effects on the loss function – reduces the model’s accuracy. Normally, this alteration step is performed manually. This step of the present invention employs an automatic model conditioning step, whereby some or all the weights corresponding to each feature column is changed to improve model conditioning. The loss function is not taken into account during the model conditioning step in the present invention. This necessarily means that the model conditioning step causes a reduction in model accuracy. FIG.4 illustrates a chart 400 of a set of epoch weights from one feature column after the loss step has converged. The optimal weights for a single feature, called Feature 1, are shown as a function of the feature’s unique values. Unique feature value ^^^భis denoted as “u_0_1” in FIG.4. This means “unique value number 0 of Feature 1.” As is evident from FIG. 4, the feature values are not smooth. They exhibit large and seemingly erratic jumps between adjacent unique feature values. As one example of performing automatic model conditioning, the weights from FIG.4 are filtered using a median filter. A median filter is an algorithm that uses a rolling window to replace the centre of each window with the window’s median value. The window size can, for example, be set to 7. In other words, the rolling window considers 7 adjacent weights and “picks” the median value of the 7 weights as the new weight for the centre of the window. FIG. 5 illustrates a chart 500 which illustrates the model weights before and after applying the median filter is shown. The weights after applying the median filter, indicated using stars as markers, have fewer abrupt jumps between adjacent unique feature values. From this step alone, the model’s conditioning has improved. As another example, a Savitzky-Golay filter can be used to automatically change the feature weights. A Savitzky-Golay filter uses a rolling window to select discrete weight values for a subsequent calculation step. A polynomial function is then fit to the values inside the rolling window. The polynomial is then evaluated at the centre of the rolling window to get a new weight value for the centre of the rolling window. The Savitzky- Golay filter’s windows size can be set to any value, with the requirement that the window size must be larger than the polynomial degree being fit. In FIG.6, an example chart 600 is shown of the model weights for a feature column before and after having been filtered using a Savitzky-Golay filter. The filter has been applied to the weights after the previous conditioning step, i.e., filtering using a median filter. The weight vector values before and after applying the Savitsky-Golay filter are plotted against the unique feature values. The filtered weights, indicated by stars, are much smoother than the non-filtered weights. The weights for this feature now contain fewer abrupt discontinuities. As another example, it may be desirable for Feature 1 to exhibit monotonic behaviour. One may want the weights from Feature 1 to increase monotonically as the unique feature values increase. This can be achieved by implementing a function that executes one of the following methods: ^^^^ೕ , ^^^ೕ ≥ ^^^ି^ೕ or ^^^^ೕ , ^^^ೕ ≤ ^^^ା^ೕ In Equation 20 and Equation 21, ℎା,^୭୰^ୟ୰^^ ൫^^^ , ^^^൯ and ℎ^ା,ୠୟୡ୩^ୟ୰^are the conditioning functions that produce a monotonically increasing weight vector through a forward or backwards pass respectively. The unique feature values must be sorted in ascending order. The conditioning function must be implemented sequentially for each weight. In other words, if ℎା,^୭୰^ୟ୰^^ ൫^^^, ^^^൯ is used, the value of the second weight must first bechanged, then the third, etc. If ℎା,ୠୟୡ୩^ୟ୰^^ ൫^^^ , ^^^൯ is used, the value of the second lastweight must first be changed, then the third last weight, then the fourth last weight etc. The weight values at the edges where the algorithm starts are normally left as-is. A combination of Equation 20 and Equation 21 may also be used, where the weight vector is split at a specific unique feature value into a first portion and a second portion. The first portion of the weights is then conditioned using a backwards pass, and the second portion is conditioned using a forward pass. Furthermore, this can be repeated iteratively by splitting the weight vector into their first and second portions on all the unique feature values. The “splitting value” can be chosen that results in the conditioned weights being closest to the original weights. This will often be the best monotonic representation of the weights. An example of using a monotonic increasing conditioning function onto the Feature 1 weights after filtering using the Savitsky-Golay filter is shown in chart 700 of FIG.7. The weight values are now monotonically increasing from the first unique feature value to the last unique feature value. The final feature value on the right-hand side appears to decrease again. This final feature value is, however, not a number, but a “NULL” value. It can sometimes happen that certain feature values do not possess ordinality. This value is therefore not altered by the monotonicity conditioning function. As another example, the number of desired turning points inside a weight vector and the monotonicity between those turning points can be specified. A similar approach as shown in Equation 20 and Equation 21 can be used to achieve this. The difference is that, instead of using only Equation 20 and Equation 21 on the first or second portion of the weight vector, different monotonic constraint functions on the first and second portions can be implemented. This may require each portion to itself be split two portions. This method works by recursively calling the entire conditioning step on each portion. This allows implementation of flexible conditioning steps. For example, it may be specified that the feature vector has two turning points. The weights must have the following monotonic behaviour between turning points: Increasing – turn – decreasing – turn – increasing This conditioning function can be summarised as implementing a “+-+” turning point conditioning function. The resulting weights of implementing a “+-+” turning point conditioning function are shown in chart 800 of FIG.8. The model now have two turning points, thereby allowing them to coincide with the pre-monotonic weights for a large portion of the feature’s unique values. Different conditioning functions may be applied to different features. For instance, a median filter may first be applied, then a Savitzky-Golay filter, and then a turning point function for a feature with many unique values. For a feature with few unique values, a simple polynomial fit on the weights may be performed and the polynomial’s evaluated values assigned as the weights. For some features, it may be desired not to perform any conditioning. There are many different model conditioning functions that can be used. New model conditioning functions may be constructed based on expert knowledge. Some examples of model conditioning functions can be found in the documentation for a popular signal processing Python Library [Scipy.signal]: ^ Order filter ^ Median filter ^ Median filter 2D, when you want to filter two features simultaneously ^ Wiener filter ^ Symiirorder1 filter ^ Symiirorder2 filter ^ lfilter filter ^ filtfilt, this applies a digital filter to a signal using a forward and backwards pass ^ Savitzky-Golay filter ^ Deconvolve filter ^ SOS filter ^ Sosfiltfilt ^ Hilbert filter, this computes the analytic signal using the Hilbert transform. ^ Detrend filter Some of these filters may be applicable, others not. It depends on the domain of the problem. There are many more filters, including window functions, that may be used. New model conditioning functions may also be created. The present invention does not claim or require specific model conditioning functions to be used. Several consecutive model conditioning steps may be applied on a feature weight vector. Alternatively, zero model conditioning steps may be applied. A constant weight adjustment to all the weights of a feature vector may be performed. This may also be accompanied by a compensating adjustment in the model intercept, to render the resulting GAM predictions exactly the same as prior to the adjustment. Detailed description of the Baseline commit step (block 216 of FIG.2) Most prior art approaches to model fitting are, for all intents and purposes, finished after the loss step. The present invention, however, performs comprehensive model conditioning on the epoch weights after the loss step. It has already been said that automatic model conditioning will result in loss of model accuracy. The problem of accuracy loss is solved by only committing a fraction of the resulting epoch weights to the baseline weights. In other words, after epoch number ^^’s automatic conditioning step (block 214 of FIG.2), the new baseline weights can be calculated as shown in Equation 22 below. ^^ୠୟ^^୪୧୬^,^ା^ = ^^ୠୟ^^୪୧୬^,^ + ^^ ∙ ^^^୮୭ୡ୦,^ Eq. (23) In Equation 22 above, the next epoch’s baseline weight vector is the sum of the previous baseline weight vector and a fraction of the current epoch weight vector after conditioning (i.e. after executing block 214 of FIG.2). The fraction, ^^, is a number between 0 and 1, including the edge values. The intercept terms are handled in the same way. FIG.9 illustrates a chart 900 of an example of the baseline weight vector as the epochs progress. The baseline weights grow further and further away from their initial values as the epochs progress. The reasoning behind committing only a fraction of the epoch weights is now explained further. It is known that performing model conditioning without regard for model accuracy causes loss of model accuracy. The present invention performs what may be described as an extreme form of this, as discussed in the detail description of the epoch weights automatic conditioning step (block 214 of FIG.2). It would be intuitive, therefore, to regard the resulting epoch weights after conditioning as the direction in which our final weights should progress rather than the solution to the mathematical problem. Once the baseline weights have been moved a small amount in the direction of a well-conditioned solution, the loss step in the next epoch can be relied on to regain the accuracy which has been lost. The present invention works because the loss step, optimising the epoch weights on top of the nonzero baseline vector, will increasingly find accurate solutions closer and closer to the well- conditioned starting point after each epoch. Another way of describing the commit step is to say that the starting point of the loss step is increasingly changed such that the loss step increasingly finds accurate solutions that are also well conditioned. The commit fraction, ^^, may change from epoch to epoch. In one example of the present invention, the commit fraction will start off small, maybe 0.1 or 0.2, during the initial epochs, and slowly grow to 1 as the epochs progress. As the baseline weights grow, the epoch weights will become smaller and smaller, hence requiring a larger commit fraction to make an impact on the baseline weights. It has been mentioned that the commit fraction should be between 0 and 1. However, the commit fraction may be set to 0 for some or all of the feature weights. This would typically happen when it is desired to update certain feature weights but leave other feature weights completely unaltered. The commit fraction may also be set to 1 on the first epoch, which is the same as saying the result of the model conditioning step represents the entire solution. Detailed description of the step of Baseline weights automatic conditioning (block 218 of FIG.2) This step comprises the similar model conditioning methods as described in the epoch weights automatic conditioning step, a difference being that the model conditioning functions are applied directly on the new baseline weights (see Equation 22), and not on the epoch weights from the loss step. FIG.10 illustrates a chart 1000 showing a turning point conditioning function applied to the baseline weights (not applied to the epoch weights). This step may be required to enforce certain conditioning traits of the final weights that are not applicable to the epoch weights. It may be desired to condition the epoch weights to be smooth, but only apply monotonic constraints to the baseline weights. This reduces potential loss of accuracy, because even though the epoch weights might not exhibit the desired monotonic behaviour, there is a possibility that the baseline weights may exhibit the desired monotonic behaviour after the commit step. Unnecessary loss of accuracy can be reduced by waiting to see if the monotonic conditioning step is really necessary on the epoch weights. By way of variation, no conditioning may be performed during this step. This will typically be the case if there is no interest in models with monotonic constraints, only smoothness. Detailed description of the step of Model convergence checking (block 220 of FIG. 2) This step serves to determine when the end of the building process has been reached, and no more epochs are required. If further steps are required, at branch 220.1, the epoch counter is incremented and the method repeats, e.g., from the loss step at block 212. However, if an exit condition is met and no further epochs are needed, at branch 220.2, the GAM is considered fully defined by taking the most recent baseline weights from the baseline weights automatic conditioning step and the baseline intercept. There are many possible stopping criteria which can be used to exit the building loop. Some examples of exit criteria are: 1. Exit the loop after the total time exceeds a set time period; 2. Exit the loop after a set number of epochs have been run; 3. Exit the loop after the accuracy, which can be measured by a variety of metrics, starts to decrease, or increase below a certain minimum value. This can be measured against the dataset itself, or a “holdout” set which isn’t used in the loss step; 4. Exit the loop once the epoch weights, individually or aggregated, become smaller than a prescribed tolerance value. FIG.3 illustrates a diagrammatic representation of a computer system 300 within which a set of instructions, for causing the computer system 300 to perform any one or more of the methodologies described herein, may be executed. In a networked deployment, the computer system 300 may operate in the capacity of a server or a client machine in server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The computer system 300 may be a personal computer (PC), a tablet PC, a set-top box (STB), a Personal Digital Assistant (PDA), a cellular telephone, a tablet, a web appliance, a network router, switch or bridge, or any computer system 300 capable of executing a set of instructions (sequential or otherwise) that specify actions to be taken by that computer system 300. Further, while only a single computer system 300 is illustrated, the term “computer” shall also be taken to include any collection of computers that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methodologies discussed herein. The example computer system 300 includes a computer processor 302 (e.g., a central processing unit (CPU), a graphics processing unit (GPU) or both, a main memory 304 and a static memory 306, which communicate with each other via a bus 308. The computer system 300 may further include a video display unit 310 (e.g., a liquid crystal display (LCD)). The computer system 300 also includes an alphanumeric input device 312 (e.g., a keyboard or touchscreen), a user interface (UI) navigation device 314 (e.g., a mouse or touchscreen), a disk drive unit 316, a signal generation device 318 (e.g., a speaker) and a network interface device 320. The disk drive unit 316 includes a computer-readable medium 322 on which is stored one or more sets of instructions and data structures (e.g., computer software or a computer program 324) embodying or utilised by any one or more of the methodologies or functions described herein. The computer software 324 may also reside, completely or at least partially, within the main memory 304 and / or within the processor 302 during execution thereof by the computer system 300, the main memory 304 and the processor 302 also constituting computer-readable media. The computer software 324 may further be transmitted or received over a network 326 via the network interface device 320 utilising any one of a number of well-known transfer protocols (e.g., HTTP, FTP). While the computer-readable medium 322 is shown in an example embodiment to be a single medium, the term “computer-readable medium” should be taken to include a single medium or multiple media (e.g., a centralised or distributed database, and / or associated caches and servers) that store the one or more sets of instructions. The term “computer-readable medium” shall also be taken to include any medium that is capable of storing, encoding or carrying a set of instructions for execution by the computer system 300 and that cause the computer system 300 to perform any one or more of the methodologies of the present embodiments, or that is capable of storing, encoding or carrying data structures utilized by or associated with such a set of instructions. The term “computer-readable medium” shall accordingly be taken to include, but not be limited to, solid-state memories and optical and magnetic media. The computer-readable medium may be a non-transitory medium. The computer system 100 may include at least some of the components of the computer system 300. COMMENTS AND ADVANTAGES The accurate assessment of risk in the industries mentioned in this document, is fundamentally a technical problem as described by the following examples of aspects with a technical nature: ^ Data complexity: Modern risk assessment must process and analyse vast amounts of heterogeneous data from multiple sources with different formats, scales and reliability. The data includes historical claim data, physical attributes of vehicles and buildings, geographical information and socio-economic factors amongst others. ^ Multidimensional analysis: The interplay between these factors creates a high- dimensional problem space that cannot be adequately analysed using simple statistical methods or human intuition alone. ^ Non-linear relationships: Many risk factors have complex, non-linear relationships that require advanced mathematical modelling techniques to capture accurately. ^ Temporal dynamics: Risk factors change over time, requiring models that can account for temporal variations and trends. ^ Computational challenges: The scale of data and complexity of models create significant computational challenges, requiring optimization of algorithms and data structures. ^ Predictive modelling: Developing accurate predictive models that can generalize to new, unseen scenarios is a complex technical task involving statistical inference and / or Machine Learning (ML). ^ Real-time adaptation: Modern risk assessment often requires real-time updates based on changing conditions, presenting challenges in streaming data processing and online learning algorithms. ^ Specialised hardware: Because of the vast size of the datasets these models are constructed from, specialized hardware such as Graphics Processing Units (GPUs), Tensor Processing Units (TPUs) or even custom hardware designs implemented on Field Programmable Gate Arrays (FPGAs) are becoming more prevalent. Moreover, because of the cost of these processing units, companies often resort to “hiring” these processors from a cloud computing company. Any improvement in the time it takes to fit these models therefore directly translate to a cost saving for the company. The business decision-making aspects such as pricing and policy decisions, are outcomes of this technical process rather than the core of the innovation itself. Different companies solve these challenges in different ways with a whole industry of software providers, consultants and specialised career professionals that has developed around this problem. The difference between a profitable company and a company veering towards bankruptcy is intimately tied to how these technical challenges are addressed in order to enable sound decision-making. An advantage of the present invention is the fact that the loss function, used in the loss step, can be simple. Conventional approaches to building GAMs use complex loss functions containing penalty terms and functions that penalise poorly conditioned solutions. Loss steps with complex loss functions are often more computationally expensive than loss steps with simple loss functions. The present invention does not have to use a complex loss function because it uses two separate model conditioning steps to achieve model conditioning. Model conditioning steps are computationally inexpensive. As an example, conventional loss functions often involve calculating the gradient of the loss function with respect to: 1. every weight in the weight vector 2. and every row in the dataset 3. and every term of the loss function The weight gradients for each row of data are then summed, yielding a total gradient for each weight. If the loss function is complex, the computational intensity of the gradient calculations can be illustrated as done below: ^ Suppose the dataset has 50 feature columns and 1,000 rows of data; ^ There are, therefore, 50 weight gradients that need to be calculated for each row of data during the loss step; ^ Suppose the loss step is composed of 2 parts: a part penalising for inaccuracy, and a part penalising non-smooth solutions. The latter part is a nonlinear function of the weight vector. ^ This means each loss iteration consists of 1,000 + 1,000 x 50 gradient computations = 51,000 calculations. In contrast to this conventional approach, if a simple loss function is used, such as the one proposed in Equation 16 using the scaling function expressed in Equation 8, the gradient for every row of data is the same for each feature. This is because the solution is linear and additive and an incremental change in each weight would have the same effect on the loss function for the entire row of data. Also, there is no second term in the loss function, thereby saving these computations. The number of computations to perform are therefore 1,000x1x1 = 1,000. There are approximately 50 times less calculations to perform for this illustrative example. The loss step should, all other things being equal, be approximately 50 times faster to fit. Another advantage of the present invention is that it leads to models that has better conditioning than prior art methods. Because the model conditioning step is not constrained by the model’s loss function, one can prepare models that are extremely well conditioned in terms of traits such as model smoothness, monotonic constraints etc. Another advantage of the present method over methods that use loss functions with regularisation embedded into it is that the method can be less sensitive to hyperparameter selection. Hyperparameters are parameters that are usually set before the training process is started. They control the learning algorithm and structure of the architecture. The parameters can differ significantly between different model architectures, cannot be inferred beyond intuition and are non-differentiable with respect to the objective function. Hyperparameters are usually tuned iteratively, if not fixed intuitively. The present method’s conditioning functions, however, can mostly be set intuitively. It is intuitive to state that a Savitsky-Golay filter or a median filter should have a certain window length. It is also intuitive to specify monotonicity constraints. On the other hand, loss functions that include regularization in them often have dimensionless scaling factors that depend on the size of the dataset and is therefore nontrivial to determine. The present invention may be used sequentially or in parallel with itself. For example, one may start off by building a GAM using a fixed set of feature columns, say 50. The GAM building loop may then be exited, thereby resulting in a first GAM. One can then evaluate this GAM and remove certain features or add new features that were excluded from the first GAM. A second GAM would then be initialised by copying the relevant feature weights from the first GAM. The present method will then be used to fit a new GAM with the old GAM as starting point. Some examples given may have implied that all feature weights need to be a) changed in the loss step, b) conditioned, c) committed to the baseline, and then d) conditioned again in the baseline step. The present invention can, however, be used to alter only a subset of the GAM’s weights whilst keeping the others constant. This would typically happen when a model is to be fine-tuned for new feature columns.

Claims

CLAIMS What is claimed is:

1. A computer-implemented method for producing an accurate and well-conditioned Generalised Additive Model (GAM) automatically, the method implemented by a5computer system comprising a processor, the method comprising: receiving, by a data parsing module provided by the processor, a dataset, which will be used to fit the GAM, wherein: the dataset comprises rows and columns containing likely heterogeneous tabular data, categorical, ordinal or a combination thereof in nature; each row contains at least two kinds of values, namely (1) one or more different observations to be predicted using the GAM and (2) one or more features which are values that the GAM uses to produce a prediction for an observation; and a single dataset row in the dataset therefore has different features and observations associated with it; initialising the GAM, by a modelling engine provided by the processor, by creating a starting model baseline weights including a vector and an intercept; constructing, by the modelling engine, the model by executing a loop with plural steps, each time the loop has been completed, an epoch has occurred, the plural steps comprising: a loss step, in which a loss function is minimised with respect to epoch weights, including vector and an intercept, of the GAM, resulting in a set of epoch weights; epoch weights automatic conditioning, in which the resulting epoch weights are automatically changed to improve the conditioning of the epoch weights;baseline commit step, in which a fraction of the conditioned epoch weights is added to the model baseline weights to produce updated model baseline weights; baseline weights automatic conditioning, in which the updated model5baseline weights are automatically changed to improve the conditioning of the model; and model convergence checking, in which an automatic check is performed to determine whether the changed baseline weights are sufficiently accurate and well-conditioned to stop the loop, thereby yielding the final GAM.

2. The method as claimed in claim 1, in which the modelling engine is configured to model the GAM as having a Poisson, Gamma, Binomial, Negative Binomial, Gaussian, Inverse Gaussian or a Tweedie statistical distribution with a set of weights which can be represented together in a weight vector.

3. The method as claimed in claim 2, in which the modelling engine provides three weight vectors, namely a baseline weight vector, an epoch weight vector, and a composite weight vector, wherein the composite weight vector is the sum of the baseline weight vector and the epoch weight vector, and wherein these weight vectors are dependent on an epoch number.

4. The method as claimed in claim 3, in which the modelling engine is configured to provide three intercepts for the loss function, namely a composite intercept, a baseline intercept, and an epoch intercept, wherein the composite intercept is the sum of the baseline intercept and the epoch intercept.

5. The method as claimed in claim 2, in which the modelling engine is configured to provide a scaling function that scales a feature weight for a particular feature in a row having a unique value in a corresponding column of the dataset, to ensurethat the weights contributing to the prediction of that row for the GAM are only the weights corresponding to the feature values that actually occur in that row.

6. The method as claimed in claim 1, in which the modelling engine is configured to5provide the baseline weights are initialised as the zero vector and the intercept is initialised such that the model can produce a prediction corresponding to the observation column’s mean value.

7. The method as claimed in claim 1, in which the modelling engine is configured to initialise the baseline weights: by copying values from a different GAM; and / or from expert knowledge or according to industry best practices.

8. The method as claimed in claim 1, in which the modelling engine is configured to implement the loss step by initialising the epoch weights to zero, including the vector of feature weights and the intercept based on the epoch number.

9. The method as claimed in claim 1, in which the modelling engine is configured to repeat the loop until a final condition is fulfilled.

10. The method as claimed in claim 9, in which the final condition includes: a total time which the modelling engine has been constructing the model exceeds a predefined time threshold; and / or a total number of epochs completed by the modelling engine exceeds a predefined epoch threshold.

11. The method as claimed in claim 2, in which the method includes determining, by the modelling engine, an accuracy of the GAM and in which the final condition includes a decrease in the accuracy below a predefined accuracy threshold.

12. The method as claimed in claim 2, in which the final condition includes epoch weights, individually or aggregated, becoming smaller than a prescribed epoch weigh threshold. 5 13. A system configured to produce an accurate and well-conditioned Generalised Additive Model (GAM) automatically, the system comprising a computer processor and a computer-readable medium, wherein the system comprises: a data parsing module provided by the processor, the data parsing module configured to receive a dataset which will be used to fit the GAM, wherein: the dataset comprises rows and columns containing likely heterogeneous tabular data, categorical, ordinal or a combination thereof in nature; each row contains at least two kinds of values, namely (1) one or more different observations to be predicted using the GAM and (2) one or more features which are values that the GAM uses to produce a prediction for an observation; and a single dataset row in the dataset therefore has different features and observations associated with it; and a modelling engine provided by the processor, the modelling engine configured to: initialise the GAM by creating a starting model baseline weights including a vector and an intercept; and construct, by the modelling engine, the model by executing a loop with plural steps, each time the loop has been completed, an epoch has occurred, the plural steps comprising: a loss step, in which a loss function is minimised with respect to epoch weights, including vector and an intercept, of the GAM, resulting in a set of epoch weights;epoch weights automatic conditioning, in which the resulting epoch weights are automatically changed to improve the conditioning of the epoch weights; baseline commit step, in which a fraction of the conditioned5epoch weights is added to the model baseline weights to produce updated model baseline weights; baseline weights automatic conditioning, in which the updated model baseline weights are automatically changed to improve the conditioning of the model; and model convergence checking, in which an automatic check is performed to determine whether the changed baseline weights are sufficiently accurate and well-conditioned to stop the loop, thereby yielding the final GAM.

14. A non-transitory computer-readable medium which, when executed by a processor, causes the processor to perform the method as claimed in claim 1.

Citation Information

Patent Citations

  • GNSS Signal Processing with Synthesized Base Station Data

    US20120154215A1

  • Regression-tree compressed feature vector machine for time-expiring inventory utilization prediction

    US20200090116A1

  • Machine learning system using a stochastic process and method

    US20210042590A1

  • Adaptively adjusting influence in federated learning model updates

    US20210287114A1

Cited By

  • Data prediction and early warning method, device and equipment suitable for deep-sea mining long-distance transmission and medium

    CN122241575A

  • Data prediction and early warning method, device and equipment suitable for long-distance transportation of deep sea mining and medium

    CN122241575B