Privacy-preserving computation of ordinal regression

A cryptographic system using secure multi-party computation addresses the challenge of efficiently fitting and evaluating ordinal regression models by employing oblivious intercept selection and multi-party exponentiation, achieving privacy-preserving and computationally efficient ordinal regression computations.

WO2025120006A1PCT designated stage expired Publication Date: 2025-06-12ROSEMAN GRP BV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2024/084737
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-05
Filing Date
2024-12-04
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Existing technologies face challenges in efficiently fitting and evaluating ordinal regression models in a privacy-preserving manner, particularly when using secure multi-party computation (MPC) due to computational limitations and trade-offs imposed by cryptography.

Method used

The development of a cryptographic system that utilizes secure multi-party computation (MPC) to perform privacy-preserving computations, specifically by obliviously selecting intercepts and using multi-party exponentiation protocols to compute likelihood gradients for ordinal regression models.

Benefits of technology

This approach enables efficient fitting and evaluation of ordinal regression models while maintaining privacy, reducing computational costs, and overcoming the limitations of direct cryptographic MPC engines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024084737_12062025_PF_FP_ABST
    Figure EP2024084737_12062025_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a cryptographic system (010) for computing a likelihood gradient of an ordinal regression model as a cryptographic secure multi-party computation between multiple cryptographic devices. The model is parametrized by weights and intercepts. The devices obliviously select a lower intercept and / or a higher intercept, and apply a multi-party exponentiation protocol based on the lower intercept and the higher intercept, respectively, to determine lower and higher exponentials. Gradients of one or more weights and / or one or more intercepts are determined under the multi-party computation based on the lower and higher exponential.
Need to check novelty before this filing date? Find Prior Art

Description

[0001]PRIVACY-PRESERVING COMPUTATION OF ORDINAL REGRESSION FIELD OF THE INVENTION The invention relates to a cryptographic system for performing a privacy- preserving computation on secret data. The invention further relates to a cryptographic device for use in such a system; to a corresponding computer-implemented method; and to a computer-readable medium. BACKGROUND OF THE INVENTION There is a growing demand for privacy enhancing technologies (PETs), i.e., data processing techniques that intrinsically protect the privacy of the data they operate on. For example, with the cryptographic technique of secure multi-party computation (MPC), multiple parties can perform a computation on their joint input using a distributed cryptographic protocol, such that each party learns nothing beyond the output of computation and his own (private) input. One reason for the growing demand for PETs is that citizens are becoming increasingly dependent on the digital information stored about them by various companies and institutions. Because of this increasing dependence, the consequences of a breach of personal data are getting increasingly severe. And due to the worldwide surge of cybercrime and nation-state-sponsored cyber espionage, the risk of a data breach has increased sharply in recent years. Also, data-based collaborations between separate entities (like companies, hospitals, local governments) usually implies that personal data is copied between the entities, which poses the risk of uncontrolled spreading of data, in particular personal information. PETs can enable data collaboration between entities without the need for sharing the data in clear-text form. Another factor driving demand for PETs is the emergence of legal frameworks for data protection, such as the European GDPR and the Californian CCPA legislation, and their mandatory compliance. In the context of such frameworks, PETs are valuable as technical safeguards, and typically provide concrete instantiations of abstract legal notions. One important functionality for PETs is the fitting of regression models to a distributed dataset comprising contributions of multiple inputters. For example, the multiple inputters may each provide different information about a common set of records (horizontally partitioned data), may each provide the same information but about different records (vertically partitioned data), or any combination thereof. By fitting a regression model on such a distributed dataset, a common regression model representing the distributed dataset can be obtained, without the need to combine the data in a single place. The fitted model can also be applied to a record while keeping the record and / or the fitted model private. More generally, also other machine learnable models can be applied in this way. An example of a known type of regression model is an ordinal regressionmodel. Such a model can consider K features X1, ... , XK which can for example becontinuous, normalized continuous, or binary variables; as well as a response variable Y. The response variable can take on a categorical yet ordered value. For instance, the response might represent a movie rating, where the possible options are ’terrible’, ’bad’, ’okay’, ’good’, ’great’. An ordinal regression model can provide a way to investigate the relation between the features (predictor variables) and the response variable, and to predict the response variable for new records for which no response is observed. Unfortunately, to obtain an efficient implementation of a given functionality using MPC, it is typically not possible to simply take an existing algorithm that implements the functionality, and run that on top of an available cryptographic MPC engine. Instead, it is typically needed to specifically adapt existing methods to the specific computational limitations and trade-offs imposed by the underlying cryptography. For example, MPC has certain limitations, e.g., it is not directly possible to branch on a secret guard, because that would reveal the value of the guard to the parties carrying out the MPC. Moreover, the relative costs of basic operations when performing a computation as a distributed cryptographic protocol on secret data under MPC, are typically different from when performing the computation directly on the plain data. This makes it necessary to design customized cryptographic approaches that take into account the specifics of using MPC. In the case of regression, for some particular types of regression, techniques are known to implement this regression using MPC. For example, the paper "High Performance Logistic Regression for Privacy-Preserving Genome Analysis" by M. de Cock et al., Cryptology ePrint Archive 2020 / 171, proposes a cryptographic protocol that implements logistic regression training using multi-party computation. SUMMARY OF THE INVENTION It would be desirable to provide techniques for performing a privacy- preserving computation on secret data that enable more types of regression model to be fitted and evaluated efficiently. In particular, it would be desirable to provide support for efficiently fitting and evaluating an ordinal regression model. In accordance with a first aspect of the invention, a cryptographic system for performing a privacy-preserving computation is provided, as defined by claim 1. In accordance with further aspects of the invention, a cryptographic device for use in such a system; and a cryptographic method of performing such a privacy-preserving computation are provided, as defined by claims 13 and 14, respectively. In accordance with an aspect of the invention, a computer-readable medium is provided, as defined by claim 15. The techniques described herein make use of cryptographic secure multi- party computation (MPC) to perform a privacy-preserving computation on secret data. As is known per se, MPC is a cryptographic technique in which a computation is performed in a distributed way between multiple cryptographic devices in such a way that the inputs, intermediate values, and / or outputs of the computation remain hidden from the parties performing the computation. Such values that remain hidden from the parties may be referred to as the secret values of the MPC. In general, a secret value of the MPC may have the property that a limited number of parties, up to a given threshold, does not know the secret value. However, a number of parties that exceeds the threshold may be able to derive the secret value. A secret value can for example be a threshold encryption, of which the decryption key is distributed among the parties; or a secret sharing, also referred to herein simply as a sharing. A sharing may be defined as a distributed representation of a value into shares of the respective parties such that a limited number of the shares, up to the given threshold, does not allow to derive the represented value. Although the term "secret share" is most commonly used for MPC techniques using so-called arithmetic secret sharing, also other MPC techniques such as garbled circuits are considered herein to operate on secret shares. In multi-party computation, through the use of secret values and of various protocols that allow to perform operations on secret values, e.g., in secret-shared or threshold encrypted form, various computations can be performed, while keeping the underlying values hidden from the parties that perform them, thus providing privacy-preserving computation. The design of efficient cryptographic protocols for implementing specific MPC operations is the topic of a significant amount of research. Various aspects may relate to an ordinal logistic regression model. The ordinal logistic regression model may be configured to be applied to a record comprising one or more respective feature values for a set of one or more respective, e.g., numeric, features. The ordinal logistic regression model may be defined with respect to a set of multiple labels. The labels may be ordered, e.g., according to a linear ordering. The labels are typically discrete, in other words categorical, e.g., there is a finite number of different labels. Given a record, the ordinal logistic regression model may be configured to classify the record according to the set of labels by outputting data indicating a correspondence of the record to the labels, e.g., a most likely label, and / or respective likelihood values for respective labels. The ordinal logistic regression model may be parametrized by a set of parameters. One or more of the parameters may be obtained by fitting the logistic regression model to a labelled dataset. The parameters may include one or more respective weights corresponding to the respective possible labels. The parameters may further include one or more intercepts corresponding to the set of labels, in particular with a respective intercept corresponding to a respective boundary between two subsequent ordered labels. Various aspects may relate to the computation of a likelihood gradient of an ordinal regression model, in particular an ordinal logistic regression model under MPC. The likelihood gradient may be a gradient of a function that is based on the likelihood function of the ordinal regression model. For example, the likelihood gradient may be based on the log- likelihood, e.g., on a logarithm of the likelihood. The likelihood function for a given record with respect to a given set of parameters may be indicative of an extent to which the model matches the record, in particular of a probability of observing the feature values of the record in combination with the label of the record. The likelihood gradient may be computed for a record with respect to a set of parameters. The record may be a labelled record, comprising feature values and a label. For example, the label may be a ground truth label. Parameter values may be obtained by fitting the model to a dataset, e.g., containing the record, or may represent intermediate values of the fitting process. More generally, the likelihood gradient may be computed as part of various larger operations, in particular the fitting of the model to a dataset, the fitting of a larger model that comprises the model as fixed or trainable sub-model, a quality evaluation of the model, etc. Various aspects may relate to the determination of a lower intercept and a higher intercept using the multi-party computation. One or both intercepts may be obliviously selected from the intercepts in the model parameters, based on the label of the record. These lower and higher intercepts may then be used to compute a lower and a higher exponential, by applying a multi-party exponentiation protocol based on the lower and the higher intercept, respectively. Interestingly, by obliviously selecting a lower and / or a higher intercept and using them for a multi-party exponentiation, the computation of the likelihood gradient can be performed relatively efficiently. The inventors realized that, in the specific computational setting of multi-party computation, it is relatively costly, e.g., in terms of computational and / or communication cost, to evaluate an exponential function. At the same time, the formulas for the gradients for a record typically comprise several exponentials, and which exponentials are needed depends on the shape of the gradient function, and on the particular label that arecord has. However, using multi-party computation, the label may remain hidden to thecryptographic devices that compute the gradients, and accordingly, the devices may not know which particular shape of gradient function to use, and may not be able to leave the other gradients unevaluated. Although it may be possible for the cryptographic devices to compute the respective gradients for the respective possible labels, and then obliviously select the correct gradient based on the label that the record actually has, the inventors realized that this would be inefficient, in particular, since respective gradient formulas for the respective possible labels include evaluations of the exponential function on different inputs. Interestingly, however, the inventors envisaged to evaluate the gradients by obliviously determining the inputs of the exponentials, namely based on selecting a lower intercept and / or a higher intercept from the intercepts of the ordinal regression model; applying a multi-party exponentiation protocol based on these intercepts; and then using the resulting lower and higher exponentials to determine the gradient. In this way, instead of separately evaluating the different possible gradients for the different possible labels, it can be exploited that these respective gradients have in common that they comprise exponentials, albeit on different inputs. Thus, by computing the inputs of the exponentials under multi-party computation depending on which formula is needed; by evaluating the exponentials on these inputs regardless of the formula; and then using the resulting exponentials depending on the formula, still using multi-party computation, the evaluation of the exponential formula may effectively be shared between the different possible formulas for the gradient, while still leaving it hidden to the cryptographic devices which formula is effectively being evaluated. In other words, by obliviously selecting inputs of the multi-party exponential protocol, the number of applications of the protocol can be made to scale in the maximum number of exponentials that are needed in any one case, as opposed to the number of exponentials that occur in the different possible formulas for the gradients. Interestingly, an oblivious selection makes this possible without the cryptographic devices having plaintext access to the label, or without even knowing what shape of formula for the gradient (corresponding to the first label, last label, or neither) applies. In particular, the label may indicate a lower intercept and / or a higher intercept to select. The lower and higher intercept may represent the intercept terms corresponding to boundaries between the label and adjacent labels, when applicable. By using oblivious selection based on multi-party computation, the lower intercept and the higher intercept may be computed by the set of cryptographic devices without knowing the computed values for the intercepts, and / or without knowing which of the intercepts were selected. This way, the label of the record can be used in a privacy-preserving way, e.g., without the cryptographic devices learning the label. Having obliviously determined the lower intercept and the higher intercept, a lower and a higher exponential may then be computed by applying a multi-party exponentiation protocol based on the lower and the higher intercept. Optionally, also a reciprocal of the lower exponential may be computed. Based on these exponentials, and optionally also this reciprocal, gradients of one or more weights and / or one or more intercepts may be determined using the multi-party computation. In particular, the computed higher exponential may contribute to gradients of one or more weights and / or one or more intercepts at least in cases where the label is not the last label, and in particular may be used both in the cases where the label is the first label and in cases where the label is neither the first nor the last label, despite a different formula for the gradient applying. The value of the higher exponential may in this case affect the value of the determined gradients. Also in the case where the label is the last label, the computed higher exponential may be used to determine the gradient, but it may be used in such a way that it does not affect the value of the determined gradients. Because the cryptographic devices may use the higher exponential in either case, but in a multi-party computation without knowing whether it affects the determined gradients, the label may remain hidden from the cryptographic devices. Similarly, the lower exponential may affect one or more determined gradients at least in cases where the label is not the first or last label. The reciprocal of the lower exponential may affect one or more determined gradients at least in the case where the label is the first label. As above, the lower exponential and / or its reciprocal may be used in such a way that the cryptographic devices do not know whether or not these values affect the determined gradients, thus contributing to the confidentiality of the label. Optionally, using the multi-party computation, a reciprocal of the lower exponential may be computed by applying a multi-party exponentiation protocol based on the lower intercept. The reciprocal may be used to compute gradients of one or moreweights. The reciprocal can also be computed from the lower exponential using a multi-partyreciprocal protocol, but the inventors realized that it is advantageous in the specific computational setting of multi-party computation, to compute the reciprocal using a multi- party exponentiation protocol. Namely, the exponentiation protocol itself may be implemented more efficiently than a reciprocal protocol, as discussed in more detail elsewhere; and moreover, by not computing the reciprocal from the lower exponential, these two values can be computed in parallel, reducing latency, which can be of particular importance in the multi-party computation setting. Optionally, to determine the gradients of the weights and / or intercepts, a product of the lower exponential and the higher exponential may be computed using the multi-party computation. In this case, the inventors found that it is more advantageous to compute the product from the exponentials than directly using an exponentiation protocol, given the relatively low overhead and round complexity of computing the product. Optionally, the oblivious selection of the lower and / or higher intercept may be performed based on determine a unit vector representation of the class using a multi-party index-to-unit-vector protocol. Such a protocol may, on input a label, output a set of respective indicator values for possible labels, indicating whether the label matches the possible label. The oblivious selection may be implemented particularly efficiency based on such a unit vector representation by avoiding respective comparisons of the label against respective possible labels. For example, the computation of the lower and / or the higher intercept may be implemented as an inner product of respective indicator values with respective intercepts. Optionally, the unit vector representation may be used in multiple computations of the likelihood gradient for the record. By re-using this information multiple times, e.g., in multiple respective iterations of fitting the ordinal regression model, overall performance can be improved. For example, the unit vector representation for a record may be precomputed, e.g., before the operation to compute the likelihood gradient is initiated, thereby reducing the latency from the moment of initiating of the operation. Optionally, the lower exponential and the higher exponential are computed at least in part in parallel. In the specific computational setting of multi-party computation, this is particularly advantageous, since the computation of the exponential is a relatively significant part of the overall computation, especially in terms of rounds of communication between the cryptographic devices. Interestingly, by computing the inputs to the exponentiations under multi-party computation, such that the cryptographic devices do not learn the label, and by then performing the exponentiations in parallel, the round complexity due to exponentiations occurring in the likelihood gradient can be reduced regardless of where they occur in the mathematical expression for the gradient, and regardless of whether they in the end contribute to the computed gradients. As discussed above, when also computing the reciprocal of the lower exponentiation by exponentiation, this exponentiation is preferably performed in parallel with the other two as well. Optionally, a further oblivious selection may be performed to determine a reciprocal of a gradient denominator. As also discussed elsewhere, the shape of the formula of the likelihood gradient for ordinal regression can depend on the label at hand, and in particular on whether the label is the first label, the last label, or neither. In particular, in the respective cases, the formula for a gradient may contain respective fractions that have denominators of different shapes. By obliviously determining a gradient denominator based on whether the label is the first label, the last label, or neither, and then applying a multi- party reciprocal protocol to compute a reciprocal of the gradient denominator, a reciprocal may be determined that can be used to computing the gradients, regardless of whichparticular shape of formula is needed for the record at hand, and without needing tocompute respective reciprocals for respective formulas that may apply to the record. Interestingly, this can be done without the cryptographic devices knowing which shape of formula to apply. In particular, the gradient denominator may be computed by computing multiple denominator candidates of respective shapes and obliviously selecting one of the denominator candidates of a particular shape based on whether the label of the record is the first label, the last label, or neither label of the set of ordered labels, e.g., using indicator values for the respective cases that can for example be derived from a unit vector representation of the class of the record. E.g., three denominator candidates may be computed regardless of the number of labels. Thus, effectively, the computation of the reciprocal may be shared between the different shapes of formula, thereby reducing the number of applications of a multi-party computation reciprocal protocol that are needed. In the computational setting of multi-party computation, the computation of a reciprocal is relatively expensive in terms of communication and / or computation, such that this sharing of the reciprocal between formulas can significantly improve the performance of the overall cryptographic protocol. For example, the reciprocal of the gradient denominator may be used to compute a gradient of a weight, and / or a gradient of the intercept for the first label, and / or a gradient of the intercept for the last label. Optionally, using the multi-party computation, a class probability and / or a classification of the ordinal regression model may be computed for a further record, different from the record(s) for which the likelihood gradient is computed. For example, the gradient may be computed as part of a fitting of the ordinal regression model, and the fitted model may be applied to the further record. When computing the class probability and / orclassification, the parameters of the ordinal regression model, the further record, or both,may be kept secret under the multi-party computation and may accordingly remain unknown to the cryptographic devices that apply the ordinal regression model. For example, the class probability and / or the classification may be computed by computing a cumulative class probability for a class based on the set of parameters of the ordinal logistic regression model. For example, a class probability may be derived by subtracting adjacent cumulative class probabilities, and / or a classification may be derived by determining, under the multi- party computation, which class has the highest class probability. Optionally, the likelihood gradient computation may be performed as part of fitting the ordinal logistic regression model to multiple records. The fitting may comprise computing many likelihood gradients, e.g., by performing multiple iterations, wherein an iteration comprises computing multiple likelihood gradients, e.g., for all or for a selection of records, and updating the set of parameters of the model. By fitting, a model may be obtained that corresponds to the provided records. Optionally, the fitting may be performed as a numeric optimization. The numeric optimization may be performed by updating, using the multi-party computation, the set of parameters based on an optimization direction and a step size. The step size may be determined from the optimization direction and from a curvature value. As explained further elsewhere in this specification, by taking the performance aspects of the MPC setting into account, such an optimization can greatly reduce the number of numeric operations that are needed, thus leading to a particularly efficient fitting of the model. It will be appreciated by those skilled in the art that two or more of the above- mentioned embodiments, implementations, and / or optional aspects of the invention may be combined in any way deemed useful. Modifications and variations of any system and / or any computer readable medium, which correspond to the described modifications and variations of a corresponding computer-implemented method, can be carried out by a person skilled in the art on the basis of the present description, and the other way round as well. BRIEF DESCRIPTION OF THE DRAWINGS These and other aspects of the invention will be apparent from and elucidated further with reference to the embodiments described by way of example in the following description and with reference to the accompanying drawings, in which: Fig.1 shows a cryptographic device; Fig.2 shows a cryptographic device; Fig.3 shows a cryptographic system; Fig.4 shows a detailed example of fitting and / or applying an ordinal regression model; Fig.5 shows a detailed example of computing a likelihood gradient; Fig. 6 shows a computer-implemented method;Fig.7 shows a computer-readable medium comprising data. It should be noted that the figures are purely diagrammatic and not drawn to scale. In the figures, elements which correspond to elements already described may have the same reference numerals. DETAILED DESCRIPTION OF EMBODIMENTS Fig. 1 shows a cryptographic device 100 for use in a cryptographic system asdescribed herein, e.g., in Fig.3. The cryptographic system may be for performing a privacy- preserving computation on secret data. The computation may be performed as a cryptographic secure multi-party computation between multiple cryptographic devices, including device 100. The privacy-preserving computation may comprise computing a likelihood gradient of an ordinal regression model for a record with respect to a set of parameters. The record may comprise one or more feature values and a label. The label may be comprised in a set of multiple ordered labels. The set of parameters may comprise one or more weights corresponding to the features and one or more intercepts corresponding to the labels. The device 100 may comprise a data interface 120 for accessing data 040 representing the set of parameters of the ordinal regression model. For example, the number of parameters can be at most or at least 20, at most or at least 40, or at most or atleast 60. For example, the number of weights can be at most or at least 10, at most or atleast 20, or at most or at least 30. The number of intercepts can be equal to the number of possible labels minus one. For example, the number of intercepts can be at most or at least 2, at most or at least 5, or at most or at least 10. Instead or in addition, the data interface may be for accessing one or more records 030. For example, the number of records can be at most or at least 1000, at most or at least 10000, or at most or at least 100000. The feature values of a record may be numeric values, e.g., represented integers or reals, represented e.g. as fixed-point numbers or floating-point numbers. The labels can for example be represented as integers, as indicator vectors, or the like. Typically, the values of the dataset 030 and / or parameters 040 are secret values of the multi-party computation. Such secret values may be stored in various ways, as also discussed elsewhere; e.g., as secret shares or the like. Values representing real numbers can be stored in a fixed-point representation, for instance. For example, as also illustrated in Fig.1, the input interface may be constituted by a data storage interface 120 which may access the data 030, 040 from a data storage 021. For example, the data storage interface 120 may be a memory interface or a persistent storage interface, e.g., a hard disk or an SSD interface, but also a personal, local or wide area network interface such as a Bluetooth, ZigBee or Wi-Fi interface or an ethernet or fibreoptic interface. The data storage 021 may be an internal data storage of the system 100, such as a hard drive or SSD, but also an external data storage, e.g., a network- accessible data storage. In some embodiments, the data 030, 040 may each be accessed from or distributed across different data storages, e.g., via a different subsystem of the data storage interface 120. Each subsystem may be of a type as is described above for data storage interface 120. The device 100 may further comprise a processor subsystem 140 which may be configured to, during operation of the system 100, compute the likelihood gradient. The computation of the likelihood gradient may comprise, using the multi-party computation, determining a lower intercept and a higher intercept by obliviously selecting the lower intercept and / or the higher intercept from the one or more intercepts based on the label ofthe record. The computation of the likelihood gradient may further comprise, using the multi-party computation, computing a lower exponential and a higher exponential by applying a multi-party exponentiation protocol based on the lower intercept and the higher intercept, respectively. The computation of the likelihood gradient may further comprise, using the multi-party computation, determining gradients of one or more weights and / or one or more intercepts based on the lower exponential and the higher exponential. For example, processor subsystem 140 may be configured fit the ordinal regression model parametrized by parameters 040 to multiple records 030, and as part of that compute the likelihood gradient for a record 030, e.g., iteratively and / or for multiple records. As also discussed with respect to Fig.3, the device 100 may be further configured to provide inputs to the multi-party computation, e.g., to obtain part of the dataset 030 in plain form and input this part to the multi-party computation, e.g. by secret-sharing, encrypting, or otherwise masking them. Instead or in addition, the device 100 may be further configured to obtain outputs from the multi-party computation, e.g., to obtain the optimized parameters 040 and / or a result of applying a model using the optimized parameters 040 to further inputs. The system 100 may also comprise a communication interface 180 configured for communication 126 with at least one further cryptographic device of the cryptographic system. Communication interface 180 may internally communicate with processor subsystem 140 via data communication 125. Communication interface 180 may be arranged for direct communication with the other devices, e.g., using USB, IEEE 1394, or similar interfaces. As illustrated in the figure, communication interface 180 may also communicate over a computer network 099, for example, a wireless personal area network, an internet, an intranet, a LAN, a WLAN, etc. For instance, communication interface 180 may comprise a connector, e.g., a wireless connector, an Ethernet connector, a Wi-Fi, 4G or 4G antenna, a ZigBee chip, etc., as appropriate for the computer network. Communication interface 180 may be an internal communication interface, e.g., a bus, an API, a storage interface, etc. Fig. 2 shows a cryptographic device 200 for use in a cryptographic system asdescribed herein, e.g., in Fig.3. The cryptographic system may be for performing a privacy- preserving computation on secret data. The privacy-preserving computation may beperformed as a cryptographic secure multi-party computation between multiplecryptographic devices including device 200. The privacy-preserving computation may comprise computing, using the multi-party computation, a class probability and / or a classification of an ordinal regression model for a record 050. For example, the ordinal logistic regression model may have been fitted as described herein, e.g., may correspond to model 040 of Fig.1. The record 050 may be as described in Fig.1. In general, the device 200 may be as described for device 100 of Fig.1. In particular, device 200 may comprise a data interface 220, a processor subsystem 240, and / or a communication interface 280, for which the same options are available as discussed for cryptographic device 100 of Fig.1. It is also possible to combine devices 100 and 200, e.g., a single cryptographic device may be configured to perform both regression model fitting and use. In particular, the cryptographic device 200 may comprise a data interface 220. The data interface can for example be for accessing data 040 representing a fitted, in other words trained, model, e.g., as fitted by device 100 of Fig.1. The data interface 220 may be further for accessing one or more records 050 representing an input to which the model is applied. Here, the model parameters 040 of the model, and / or the input 050 to which the model is applied, may be secret values of the multi-party computation. Similarly to device 100 of Fig.1., processor subsystem 240 may be further configured to provide input to and / or receive output from the multi-party computation. In general, each device described in this specification, including but not limited to the system 100 of Fig.1 and the system 200 of Fig.2, may be embodied as, or in, a single device or apparatus, such as a workstation or a server. The device may be an embedded device. The device or apparatus may comprise one or more microprocessors which execute appropriate software. For example, the processor subsystem of the respective system may be embodied by a single Central Processing Unit (CPU), but also by a combination or system of such CPUs and / or other types of processing units. The software may have been downloaded and / or stored in a corresponding memory, e.g., a volatile memory such as RAM or a non-volatile memory such as Flash. Alternatively, the processor subsystem of the respective system may be implemented in the device or apparatus in the form of programmable logic, e.g., as a Field-Programmable Gate Array (FPGA). In general, each functional unit of the respective system may be implemented in the form of a circuit. The respective system may also be implemented in a distributed manner, e.g., involving different devices or apparatuses, such as distributed local or cloud-based servers. Fig. 3 shows a cryptographic system 010 for performing a privacy-preservingcomputation as a cryptographic secure multi-party computation. The computation may comprise fitting an ordinal regression model and / or using a fitted ordinal logistic regression model, as described in more detail elsewhere. The cryptographic system 010 may in general comprise multiple input devices, multiple different cryptographic devices, and at least one result device, where the sets of input, cryptographic, and result devices may overlap with each other. As illustrated, the devices typically communicate over a computer network 099, e.g., the internet or a local network. In particular, shown in the figure are three cryptographic devices CP1, 201; CP2, 202; and CP3, 203. The cryptographic devices may be based on cryptographic device 100 of Fig.1 or cryptographic device 200 of Fig.2. The number of cryptographic devices that is used can vary depending on the particular technique used for the multi-party computation and the security properties which are desired. For example, the number of cryptographicdevices CPi can be two, three, or more.The cryptographic devices CPi may be configured to perform a secure multi-party computation (also known per se as multi-party computation, secure computation, or MPC). Generally, a multi-party computation may be a distributed protocol between thecryptographic devices for performing a computation in a privacy-preserving way. Dependingon the specific technique used, MPC may ensure privacy and / or correctness of the computation against an attacker that eavesdrops or controls one or more (but typically notall) of the cryptographic devices. As known per se, any computation can be performed as amulti-party computation (in other words, “under the multi-party computation”), but concrete computational and communication efficiency can in general greatly depend on how exactly the computation is performed. In particular, the multi-party computation can be performed based on one of the following techniques: -based on secret sharing, in particular arithmetic secret sharing such asShamir secret sharing, replicated secret sharing, or additive secret sharing. For example, the multi-party computation can be based on the techniques described in Shamir, “How to Share a Secret”, Communications ACM, 1979; Ben-Or, Goldwasser, Wigderson, “Completeness Theorems for Non-Cryptographic Fault-Tolerant Distributed Computation (Extended Abstract)”, Proceedings of the 20th Annual ACM Symposium on Theory of Computing, 1988; Chaum, Crepeau, Damgaard, “Multiparty Unconditionally Secure Protocols (Extended Abstract)”, Proceedings of the 20th Annual ACM Symposium on Theory of Computing, 1988; Ito, Saito, Nishizeki, “Secret sharing scheme realizing general access structure”, Electronics and Communications in Japan (Part III: Fundamental Electronic Science), 1989; Damgaard, Pastro, Smart, Zakarias, “Multiparty Computation from Somewhat Homomorphic Encryption”, proceedings CRYPTO 2012; -based on garbled circuits, e.g., see Yao, “Protocols for SecureComputations (Extended Abstract)”, 23rd Annual Symposium on Foundations of Computer Science, Chicago, 1982; -based on oblivious transfer, e.g., see Goldreich, Micali, Wigderson, “How toPlay any Mental Game or A Completeness Theorem for Protocols with Honest Majority”, Proceedings of the 19th Annual ACM Symposium on Theory of Computing, 1987; -based on threshold homomorphic encryption, e.g., see Cramer, Damgaard,Nielsen, “Multiparty Computation from Threshold Homomorphic Encryption”, proceedings EUROCRYPT 2001; -based on any combination of the above, e.g., see Demmler, Schneider,Zohner, “ABY - A Framework for Efficient Mixed-Protocol Secure Two-Party Computation”,proceedings NDSS 2015. Various higher-level operations such as sorting and integer comparison can be performed based on such basic multi-party computation protocols as discussed e.g. in M. Keller, "MP-SPDZ: A Versatile Framework for Multi-Party Computation", proceedings ACM CCS 2020; or as implemented in MPyC, see https: / / github.com / lschoe / mpyc. The multi-party computation may be configured to perform operations on so called sharings, or secret shares, of values. A secret share may be a distributed representation of an input, intermediate, or output value of the MPC. A limited number ofshares, up to a certain threshold ^, may not allow to derive the represented value. Thethreshold may be configurable, with different techniques supporting different possible threshold. For example, the multi-party computation may be an honest majority MPC, wherethe threshold ^ is strictly smaller than half the number of parties ^, e.g., 1⁄ 2 (^ − 1). Or, themulti-party computation can be a full-threshold MPC, where the threshold can be higher,e.g., ^ − 1. Examples of sharings are arithmetic sharing, such as Shamir secret sharing orreplicated secret sharing; XOR sharing; or Yao sharing. It is stressed that the term secret sharing in this specification also includes Yao sharings, e.g., secret values of an MPCcomputation performed using garbled circuits, as also done in “ABY - A Framework forEfficient Mixed-Protocol Secure Two-Party Computation”. A value that is computed on by the MPC but that is represented among the parties in such a way that no single party, more generally no unqualified set of parties, can derive the value from that representation, is referred to as a secret value, or private value, of the MPC. A secret value can be a secret sharing, but it is also possible e.g. to use a threshold encryption. For example, a secret value can be a secret input, a secret output, or a secret intermediate value. Here, a secret input may be known in the plain by the party inputting it, and known only in a secret representation by the cryptographic devices CPi; and similarly, a secret output may be learned in the plain by the party receiving it as output, but may be known only in a secret representation by the cryptographic devices CPi. A private intermediate value may be known only to the cryptographic devices CPi, and only as a secret representation. By processing values using secret representations, the data can be kept secret, at least as long as the underlying assumptions of the multi-party computations (e.g., a number and / or type of corruptions of the cryptographic devices) are satisfied. Also shown in the figure are a number of input devices INP1, 101; INP2, 102; up to INPk, 103. The input devices may input values occurring in the fitting or use of an ordinal logistic regression model, e.g., the input devices may together input a dataset of records on which the model is fitted and / or applied, and / or initial values for the parameters of the model to be fitted. For example, respective input devices may input respective sets of records with a common set of features, or may input respective sets of features for a common set of records. As another example, an input device may input a record to which a fitted model may be applied. The input devices 101-103 may use the hardware configuration discussed in Fig.1. The number of input devices can be two, at most or at least three, or at most or atleast five, for example. In many cases, the sets of inputs devices INPi and cryptographicdevices CPi may wholly or partially overlap. For example, the set of input devices may be asubset or a superset of the set of cryptographic devices, or may be exactly the same. Further shown is a result device RES. The result device RES may obtain aresult of the MPC based on the performed privacy-preserving computation. For example, theresult device RES may obtain the fitted parameters of the ordinal regression model; or avalue derived from the fitted parameters, e.g., a result of applying the fitted model to an input. It is also possible for multiple respective result devices to obtain multiple respective results of the multi-party computation. Although illustrated as a separate device in the figure,the result device(s) RES can be the same devices as an input device INPi and / or cryptographic device CPi. Generally, the result device may be implemented using the hardware configuration discussed w.r.t. Fig.1. Many known multi-party computation techniques are defined per se for thecase where the input and result devices INPi and RES form a subset of the set ofcryptographic devices CPi that perform the MPC. To use such techniques in a setting where an input and / or result device does not perform the MPC itself, an input device can for example determine a secret representation, e.g., a secret sharing, and distribute it among the computation devices. Similarly, a result device can for example receive a secret representation, e.g., respective secret shares, of an output from the computation devices and derive the output from the secret representation. It is also possible to use specific techniques for letting an external party provide inputs to and / or obtain outputs from a multi- party computation. For example, the techniques from the following reference can be used: T. P. Jakobsen, J. B. Nielsen, and C. Orlandi. “A framework for outsourcing of secure computation”, proceedings CCSW’14. Some general information on ordinal regression is now presented. For further information about ordinal regression, reference is made to chapter 13 of Harrell, F.E. (2001) Regression Modeling Strategies: With Applications to Linear Models, Logistic Regression, and Survival Analysis. Springer-Verlag, New York. http: / / dx.doi.org / 10.1007 / 978-1-4757- 3462-1 (incorporated herein inasafar as formulas for ordinal regression are concerned). General, an ordinal regression may relate to records comprising values for anumber of ^ features ^^, ... , ^^. A respective variable can for example be a continuousvariable or a binary variables, e.g., a normalized continuous variable. Different types can be combined. Features may also be referred to as predictor variables. A labeled record may further comprise a label, also referred to as a response variable, ^. The label may take on a categorical yet ordered value. For instance, the response might represent a movie rating, where the possible options are ’terrible’, ’bad’, ’okay’, ’good’, ’great’. Using an ordinalregression model model, the relation between the features and the response variable maybe investigated, and predictions may be made about further records for which no response is observed. Mathematically, an ordinal regression model may be represented as follows.Let ^ denote the number of classes. The classes may be represented by numbers 1, ... , ^.The probability that a response is at most class ^ ∈ {1, ... , ^ − 1} given the predictors ^ maybe modeled as: where ^(^,^^^) ∈ ^ denotes the intercept term corresponding to classes ^ and ^ + 1, and ^ ∈^^ denotes the model weights that correspond to the ^ features. Overall, there may be ^ −1 intercept terms, corresponding to respective boundaries between classes. It may beobserved in the expression above that ^(^ ≤ ^|^) = 1.Moreover, it may be noted that:^(^ = ^|^) = ^(^ ≤ ^|^) − ^(^ ≤ ^ − 1|^)for ^ ∈ {2, ... , ^}. Overall, the probability mass function may be characterized as follows: ^(^ = ^|^) = ^(^ ≤ 1|^), ^^^ = 1.Given records ^^and corresponding observed responses ^^, the likelihood function of the model may be defined as: where the probability is above. In many cases, the log-likelihood may be used, in which the above expression may correspond to a sum: ^^)^ For the sake of fitting a model to a given set of known predictor values- responses, the gradient of the log-likelihood function may be used. First, the partialderivatives w.r.t. ^^, ... , ^^ are discussed. In general: Three cases for ^ may be distinguished. If ^^ = 1, then it may be found that: denotes the ^-th component of the feature vector ^^. Next, if ^^ = ^, then it maybe found that: Finally, if 1 < ^^ < ^, then it may be found that: Now the partial derivatives with respect to the intercept terms^(^,^), ^(^,^), ... , ^(^^^,^) are discussed: As for ^, the cases for the class ^ may be considered. If ^^ = 1, then theremay only be contribution towards the partial derivative w.r.t. ^(^,^), with the others being zero: If ^^ = ^, then the only non-zero contribution may be towards the partialderivative w.r.t. ^(^^^,^): If 1 < ^^ < ^, the contribution may be to the partial derivatives w.r.t. ^(^^,^^^^) Fig. 4 shows a detailed, yet non-limiting, example of how to fit and / or applyan ordinal regression model using multi-party computation. Generally, data shown in this figure may be stored in secret form, e.g., as secret shares, by the cryptographic devices performing the multi-party computation, e.g., as discussed w.r.t. Fig.3. The figure shows a record comprising a feature vector FV, 431, and a label LB, 432. The feature vector may comprise one or more feature values. The label may becomprised in a set of multiple ordered labels, e.g., may be represented as a number 1, ... , ^.The figure further shows parameters of an ordinal regression model. The ordinal regression model may be an ordinal logistic regression model. The parameters may comprise one or more weights WT, 442, corresponding to the features. The parameters mayfurther comprise one or more intercepts IC, 441, corresponding to the labels. The figure further shows a gradient operation Grad, 450, wherein a likelihoodgradient ∇F, 410, of an ordinal regression model may be computed for record FV, LB, withrespect to the parameters. As illustrated in this figure, the gradient operation Grad may be performed aspart of fitting the ordinal logistic regression model to a set of multiple record, including the record FV, LB. In general, such fitting may be performed in various ways. The fitting may be an iterative process wherein an iteration may comprise an update operation Upd, 490, inwhich the parameters IC, WT are updated based on the determined gradient ∇F. Ways ofperforming an iterative optimization based on a gradient under multi-party computation are known per se and can be applied. A particularly advantageous way of performing the optimization is discussed elsewhere in this specification. As initial values for the iterative optimization, it is for example possible to use zero for the weights and to use the logitfunction of the cumulative counts for the intercepts, e.g., as computed under MPC.The figure further illustrates the use of the ordinal regression model parameterized by parameters IC, WT, by applying the model to a record FV. The record to which the model is applied can be different from the records to which the model is fitted. Moreover, the fitting and applying do not need to be performed by the same set of cryptographic devices. As illustrated in this figure, using the multi-party computation, a classification operation Class, 497, may be performed, in which a classification output CL, 498 may be determined for the record FV. The classification output may comprise a class probabilityand / or a classification, for example. In the illustrated example, the classification output CL iscomputed by computing Cum, 495, a cumulative class probability CP for a class based onthe set of parameters IC, WT. In particular, the following pseudocode illustrates how operations Cum, Class may be implemented. Square brackets denote values that are in this example computed on as secret values of the multi-party computation, e.g., are secret-shared:def predict_row( [ip] = <[β*], [x*]> / / inner productfor c=1,...,C-1: [cpc] = 1 * Rec(1+Exp([-[θc]-[ip])) / / cumulative class probabilitiesfor c=1,...,C-1: [pc] = [cpc] - [cpc-1] / / by convention, [cp0]=0[pC] = 1 - [p1] - ... - [pC-1]return [p1], ..., [pC] / / class probabilitiesdef classify_row([p1], ..., [pC]): [c] = Argmax([p1], ..., [pC]) / / by recursively comparing return [c] / / classificationIn the above example, Rec can be implemented by a multi-party reciprocalprotocol, as is known per se. Exp can be implemented by a multi-party exponentiationprotocol, as also discussed elsewhere. Argmax can be implemented for example byiteratively computing a maximum under the multi-party computation, and, in the iteration, updating the maximum index obliviously under the if the current index is maximal. The values that are being computed upon may be represented in the multi-party computation as fixed-point numbers, for instance. When performing fitting, the following pseudocode can for example be used to implement the initialization of the parameters IC, WT: function [c1,...,cC] = bincount([y1], ..., [yN]) [Σc1], ..., [ΣcC] = [c1], [c1]+[c2], ..., [c1]+...+[cC] for c = 1, ..., C-1: [θc] = logit([Σcc]) for j = 1, ..., K: [βj] = 0 return [θ1], ..., [θc- In this example, bincount refers to a multi-party computation protocol for computing the number of times the class values occur in the list of labels. This can be implemented for example by using an index to unit vector protocol as discussed elsewherein this specification, and summing the results. Moreover, logit refers to a multi-partycomputation protocol for computing the logit function. This can be implemented using techniques that are known per se, such as a Taylor expansion. It is also possible to evaluate the logit function based on a multi-party computation implementation of the log function. In this way, a particularly accurate and efficient approximation may be obtained, as also discussed elsewhere in this specification. Fig. 5 shows a detailed, yet non-limiting, example of how to compute alikelihood gradient of an ordinal regression model, e.g., an ordinal logistic regression model, using multi-party computation. The discussed techniques can e.g. be used in combinationwith Fig.1, with Fig.3, and / or to implement the Grad operation in Fig.4. As in the abovefigure, some or all of the data shown in this figure may remain secret to the cryptographic devices performing the multi-party computation, e.g., may be secret-shared. Values can be represented as integer and / or fixed point values, for example, as appropriate. The likelihood gradient may be computed for a record. As illustrated in the figure, the record may comprise one or more feature values FV, 531. The feature values may be as discussed for Fig.4. Moreover, the record may comprise a label LB, 532. The label may be comprised in a set of multiple ordered labels, e.g., as discussed for Fig.4. The figure further shows a set of parameters of the model. The set of parameters may comprise one or more weights WT, 542 corresponding to the features FV. The set of parameters may further comprise one or more intercepts IC, 541, correspondingto the labels LB. The weights WT and intercepts IC may be as discussed for Fig. 4.As shown in the figure, an intercept selection operation ISel, 520, may be performed in which, using the multi-party computation, a lower intercept LI, 524, and a higher intercept HI, 525, may be determined by obliviously selecting the lower intercept LI and / or the higher intercept HI from the one or more intercepts IC based on the label LB of the record. The lower intercept LI may be the intercept corresponding to the boundary preceding the label LB. If the label is the lowest label, no lower intercept may be selected, e.g., the lower intercept LI may be set to zero. The higher intercept HI may be the intercept corresponding to the boundary succeeding the label LB. If the label is the highest label, no higher intercept may be selected, e.g., the higher intercept may be set to zero. At least one of the lower and the higher intercept terms may be set to an intercept parameter IC. Specifically, as shown in the figure, the lower and higher intercepts may be determined based on a unit vector representation UV, 522, of the label LB. The unit vector representation may comprise respective indicator values for the respective possible labels. Optionally, an indicator value for one label can be left implicit by defining it based on theother indicator values, e.g., as one minus the other indicator bits. To determine the unitvector representation UV, a cryptographic index-to-unit-vector protocol I2Uv, 521, may be applied. Such protocols are known per se, and are available e.g. in the mpyc and MP-SPDZ implementations of multi-party computation, e.g., as demux_array(bit_decompose(value, C)) in MP-SPDZ. As illustrated in the figure, the lower and higher intercept LI, HI, may bedetermined from the set of intercept parameters IC and the unit vector representation UV byapplying a multi-party inner product protocol InPr, 523, implementations of which are knownper se. Interestingly, the unit vector representation UV may be re-used between multiplecomputations of the gradient for the same record. Based on the lower and higher intercepts LI, HI, a multi-party exponentiation protocol Exp, 560, may be applied to determine lower and higher exponentials EL, 562, EH, 563. The inputs to the exponentials may be defined as needed to compute the likelihood gradient in respective cases depending on the label LB, e.g., as the sum of the lower andthe higher intercept respectively with an inner product of the weights WT with the featurevalues FV. A detailed example of a particularly efficient and accurate multi-party exponential protocol is discussed elsewhere, although generally a regular Taylor expansion can also be used. As shown in the figure, in addition to the lower and higher intercept, theexponentiation protocol Exp may also be used to compute, using the multi-partycomputation, a reciprocal ELI, 561, of the lower exponential. Also this reciprocal may beused in some cases of the formulas for the likelihood gradient. The input may be the additive inverse of the input for the lower exponential EL, and exponentiation Exp may thus result ina reciprocal of the exponential. Interestingly, by using exponentiation protocol Exp instead ofdetermining reciprocal ELI from the lower exponential EL itself, both round complexity andoverall complexity may be reduced given that, especially with the techniques discussed herein, exponentiation can be implemented relatively efficiently. In general, outputs ELI, EL,EH of the exponentiation protocol Exp may be computed at least in part in parallel.The computation of the gradient may also use an exponent of the sum of the inputs for the lower and higher exponentials. Interestingly, this exponent may be determinedby computing, using the multi-party computation, a product of the lower exponential EL andthe higher exponential HE. In this case, the low cost of computing the product compared to the exponential, may lead to this option being preferred to using the exponential protocolExp to compute this exponential.The computation of the gradient may further comprise determining a gradientdenominator GD, 571, based on whether the label LB is the first label, the last label, orneither label according to the ordering of the labels. For example, respective candidate denominators may be determined according to the respective cases, and the gradient denominator GD may be determined by making an oblivious selection DSel, 570, among the candidate denominators. The oblivious selection can for example be implemented by applying a multi-party inner product protocol to values derived from the exponentials ELI, EL, EH on the onhand, and indicators derived from the determined unit vector UV, on the other hand. Forexample, a first indicator bit may be determined, indicating whether the label LB is the firstlabel; and / or a last indicator bit may be determined, indicating whether the label LB is thelast label; and / or a middle indicator bit may be determined, indicating whether the label LB is neither the first label nor the last label. A multi-party reciprocal protocol Rec, 580, may be applied to the gradientdenominator GD to obtain a gradient reciprocal GR, 581. Thus, a reciprocal may be obtained that can be used in the further computation DGrad, 510, of one or more intercept gradients IG, 511 and / or one or more weight gradients WG, 512, instead of computing respective reciprocals for the respective cases. In particular, the gradient denominator may contributeto (in the sense of changing the computed value) gradients WG of one or more weightsand / or to a gradient IG of the intercept for the first label and / or a IG gradient of the interceptfor the last label. A detailed pseudocode example is now given of determining a likelihood gradient. As in other examples, values between square brackets may be secret, e.g., secret- shared among the cryptographic devices performing the multi-party computation. The values may be represented as integers and / or fixed point numbers, for example.def precompute([X*,*],[y*]): / / compute data for use in multiple evaluations for the recordsfor i = 1,..., N: [Yi,*] = to_unit_vector([yi], C) / / compute unit vector rep'n of label, C possible values[fi] = [yi] == 1 / / compute first indicator bit[li] = [yi] == C / / compute last indicator bit[mi] = 1 - [fi] - [li] / / compute middle indicator bitreturn [Y*,*], [f*], [l*], [m*] ([θ*], [β*], [Xi,*], [Yi,*], [fi], [mi], [li]): [θL] = [θ1][Yi,2] + ... + [θC-1][Yi,C] / / obliviously select lower intercept[θH] = [θ1][Yi,1] + ... + [θC-1][Yi,C-1] / / obliviously select higher intersect[eL] = exp([θL] + [ip]) ; [eH] = exp([θH] + [ip]) ; [eLI] = exp(-[θL] - [ip]) / / exponentiation[d] = [mi] ? ([eH] + 1) * ([eL] + 1) : ([fi] ? [eH] + 1 : -[eLI] - 1) / / comp. gradient denominator[n] = 1 - ([mi] ? [eL] * [eH] : 0)[t] = [n] * rec([d]) / / compute reciprocal of gradient denominator[∂ / ∂β1], ..., [∂ / ∂βk] = [t] * [Xi,1], ..., [fx] * [Xi,k] / / partial derivatives w.r.t. weights β[eθL] = exp([θL]) ; [eθH] = exp([θH]) [∂ / ∂θ1] += [fi] * [t] for c = 1, ..., C-2:[∂ / ∂θc] += [mi] * [Yi,c+1] * [eθL] * (1 + [eH]) * rec(([eθL] - [eθH]) * (1 + [eL]))for c = 2, ..., C-1: [∂ / ∂θc] += [mi] * [Yi,c] * [eθH] * (1 + [eL]) * rec(([eθH] - [eθL]) * (1 + [eH]))[∂ / ∂θC-1] += [li] * [t] / / partial derivatives w.r.t. intercepts θ return In the above example, to_unit_vector denotes a multi-party index-to-unit-vector protocol, as also discussed elsewhere. Moreover, rec denotes a multi-party reciprocalprotocol, as is known per se in the art. Further, exp denotes a multi-party exponentiationprotocol. In general, such a protocol may be implemented using techniques that are known per se in the art, e.g., using by a multi-party computation evaluation of a Taylor expansion. Interestingly, however the inventors also envisaged a particularly efficient and accurate technique for evaluating the exponential function under multi-party computation, as discussed in more detail elsewhere. Various embodiments involve the evaluation of a logarithm under multi-party computation. In particular, applying an ordinal logistic regression model to a record, as discussed with respect to Fig.4, may be based on evaluating a logarithm. Interestingly, the inventors devised techniques by which this evaluation can be performed particularly efficiently and accurately. In particular, evaluation of a logarithm with base 2 is discussed. Logarithms with other bases can be performed by combining such a base 2 logarithm with an approximate multiplication. To evaluate the logarithm, the value may first be normalized to the interval [0.5, 1]. An evaluation of the base-2 logarithm of the normalized value may be determined. The estimate may be converted to an evaluation of the base-2 logarithm of the original input. The normalization can for example be performed as described in O. Catrina and S. Amitabh, "Secure computation with fixed-point numbers.", proceedings Financial Crypto 2010. The normalization may return c, v such that c = x * v, as well, as a prefix OR of the bit decomposition of the input. The first transition from 0 to 1 in the prefix or may correspond to the bit length of the input. The evaluation of the base-2 logarithm of the normalized value may be performed by performing an evaluation of the natural logarithm, and multiplying the result by 1 / log(2). The evaluation of the natural logarithm can be performed by performing a Taylor expansion, for example, using Horner's Method. A Taylor expansion around the point 0.75 was found to work well. The base-2 logarithm of the input can be determined form the base-2 logarithm of the normalized input using known techniques. In particular, if the normalization results in c, v such that c = x * v, abd letting c' = xv / 2^e, the computed logarithm may belog2(c') = log2(x) + log2(v) - e. From this, log2(x) may be computed, for example, by usingthe returned prefix OR of the bit decomposition of v.Returning to Fig.4, the fitting of the ordinal regression model may beperformed as a numeric optimization of an objective function that is based on the likelihood gradient, e.g., that minimizes the log-likelihood of the ordinal regression model. Although techniques are known per se for performing numeric optimizations under multi-party computation, the inventors envisaged a particularly efficient way of performing the optimization, with good convergence properties. Namely, as also illustrated in the figure, the optimization may be performed by updating, using the multi-partycomputation, the set of parameters IC, WT based on an optimization direction w, 470 and astep size α, 480. Here, the the step size α may be determined using the multi-partycomputation from the optimization direction w and from a curvature value. This step size isalso often referred to as the learning rate. In particular, the provided optimization may allow to obtain accurate results with a decreased the number of iterations, and accordingly, a decreased number of evaluations Grad of the gradient. In many cases, in known MPC works, a step size is used that is constant and known to the parties. Thereby, it is avoided to compute the step size under MPC. In particular, MPC techniques thereby avoid adapting to the MPC setting techniques todetermine the step size dynamically e.g. using line search. In line search, the step size isdetermined by one or more evaluations of the objective function. In MPC, this is undesirable since this typically involves a large number of numeric operations. However, using a constant step size, especially in the setting of multi-party computation, has the disadvantage that, if if chosen too large, the method may not converge, or may run into numerical issues, e.g. due to limited range of fixed point arithmetic as is often used in secure multi-party computation. If chosen too small, convergence is generally slow, leading to poor performance. Interestingly, however, the inventors realized that it is still possible to use a dynamic step size when performing a numeric optimization using MPC. Namely, the inventors propose to compute the step size as a secret value of the MPC from two other secret values of the MPC, namely, from the secret optimization direction and from a secret curvature value for the objective function. This is also referred to as a curvature-adaptive step size. Accordingly, the step size may be determined dynamically without line search. The curvature value may be indicative of the curvature of the objective function in the direction of the optimization direction, e.g., may be computed as an inner product between the optimization direction and the gradient of the objective function. In the specific setting of MPC, the use of a curvature-adaptive step size leads to significant performance improvements. Compared to using a step size that is constant, or that is determined independently from the optimization direction and / or the curvature, e.g., according to a fixed schedule, the curvature-adaptive step size leads to faster convergence, while it avoids the overhead that attempting to use line search in MPC would be expected give. It is noted that the trade-off between a fixed step size; a curvature- adaptive step size; and a step size determined by line search is different in the MPC setting than when performing optimization on plain data, where the overhead of using line search can be much more acceptable. Specifically, the inventors realized that, in the MPC setting, the use of a curvature-based step size as described herein, makes it possible to efficiently apply a quasi- Newton optimization algorithm on secret data. For example, BFGS, or especially L-BFGS, can be used. Generally, in a quasi-Newton optimization algorithm, an optimization direction may be used that is determined by computing a product of an approximation of an inverse Hessian of the objective function, with a gradient of the objective function. Interestingly, quasi-Newton methods typically use fewer iterations than for example gradient descent. However, when applying quasi-Newton methods, it typically does not work to use a constant or otherwise fixed step size. Namely, in many cases, this does not lead to convergence. Such convergence problems can arise for example when using quasi-Newton to fit a regression model such as logistic regression. Accordingly, when applying L-BFGS outside of the MPC setting, line search is typically used. This is in contrast to the MPC setting where, as discussed, it is known in the literature to use gradient descent with a fixed step size, and where attempting to use line search would incur a great performance penalty. Moreover, in the MPC setting, selecting a good step size is particularly problematic, since the data and parameters need to be kept secret and so cannot be inspected to manually correct the learning rate. Interestingly, the use in MPC of a secret step size determined based on a secret curvature value, as described herein, makes it possible to efficiently use quasi- Newton algorithms in the MPC setting. In particular, by using a quasi-Newton method, compared to existing MPC-based optimization techniques that use e.g. gradient descent, the number of iterations needed to reach convergence can be greatly reduced, greatly improving efficiency of the overall optimization. When using a fixed step size, in practice, quasi-Newton methods can have numeric stability and convergence issues, so when using a fixed step size, gradient descent may be preferred. Due to the adaptive step size, however, convergence problems due to the use of a fixed step size are avoided. In particular, super- linear convergence of the overall method can be attained, in particular for logistic regression but also for other optimization problems. Moreover, the computation of the step size can be implemented relatively efficiently under MPC, e.g., it is avoided to evaluate the objective function as part of determining the step size, as would be needed when attempting to use line search in MPC. Optionally, the secret step size may be stored, by a cryptographic device carrying out the multi-party computation, as a secret fixed-point value of the multi-party computation. The secret step size may be computed by applying one or more privacy- preserving fixed point arithmetic protocols. Cryptographic protocols for carrying out fixed- point computations are known per se and can be applied relatively efficiently. In particular,especially in combination with performing a normalization of the input data and / or including aregularization term in the objective function of the numeric optimization, fixed-point numbers may be used to combine beneficial performance with sufficient accuracy. It is also possible to use secret floating-point values, however. Optionally, the curvature value may be determined under the multi-party computation by determining a secret inner product between the secret optimization direction and a secret gradient of the objective function. Interestingly, such an inner product can be implemented relatively efficiently using MPC, in particular when using MPC based on linear secret sharing and / or when using MPC-based fixed-point representations, as discussed in more detail elsewhere. The inner product may represent the curvature in that a small value may indicate a relatively flat curvature of the objective function in the optimization direction, and a large value may indicate a relatively steep curvature. The step size may be determined as a decreasing function in the amount of curvature, e.g., the more curvature, the smaller the step may be. Detailed examples are provided herein. Optionally, the step size may be determined from the gradient and from the optimization direction by applying a privacy-preserving inner product protocol; a privacy- preserving square root protocol; and a privacy-preserving reciprocal protocol. In particular, the inner product may be used to determine the curvature value, as discussed; and the square root and the reciprocal may be used to determine the step size as a decreasing function in the amount of curvature. As also discussed elsewhere, MPC protocols for thesenumeric operations are known per se in the art and can be implemented relatively efficiently.Optionally, the numeric optimization may be implemented by using L-BFGS. In this case, a two-loop recursion may be used to determine the secret optimization direction. Interestingly, it is possible to implement this two-loop recursion efficiently under MPC by using numeric operations for which implementations are available per se, e.g., using a fixed-point inner product, a fixed-point reciprocal, and a fixed-point multiplication; or alternatively using their floating-point equivalents. This way, the low computation overhead, low memory requirements, and accurate results of L-BFGS can be attained efficiently in the MPC setting, especially when combined with the curvature-based step size computation. As illustrated in Fig.4, the optimization may be be performed by determining, using the multi-party computation, a secret optimization direction w, 470; and determining, using the multi-party computation and based on the optimization direction w, a secret step size α, 460. The step size is also referred to as the learning rate. Based on the secret optimization direction w and the secret step size α, and using the multi-party computation, the secret parameters IC, WT of the objective function may be updated. In particular, thedetermining of w and α, and the updating of IC, WT may be performed iteratively, e.g., for afixed number of iterations or until a stopping criterion is reached. It is noted that various optimization techniques that are known per se follow the pattern of iteratively determining anoptimization direction w and updating the parameters IC, WT to be optimized using a stepsize. For example, the numerical optimization used can be gradient descent (e.g., damped gradient descent); Newton-Rhapson (e.g., damped Newton-Rhapson); or a quasi-Newton method such as BFGS or L-BFGS. In particular, in many cases, determining the optimization direction w maycomprise applying a gradient evaluation operation Grad, 450, in which a gradient ∇F, 410, of the objective function may be evaluated under the multi-party computation. This is the case for gradient descent, where the gradient ∇F may be used as optimization direction w;but also for quasi-Newton methods, where the gradient ∇F may be used as an input to afurther optimization direction determining operation Dir, 420 that outputs the direction w. Asillustrated in the figure, gradient evaluation Grad in many cases uses a dataset FV, LB onwhich the parameters IC, WT are fitted, as illustrated elsewhere in this specification forlogistic regression. Interestingly, the inventors realized that it is possible to use L-BFGS as numeric optimization technique under multi-party computation. L-BFGS is beneficial because it is efficient numerical optimization method with a high convergence rate that does not need many evaluations of either the gradient Grad or the objective function, e.g., the log- likelihood function, itself. This is particularly beneficial in combination with multi-party computation, because in this setting the numeric operations involved in evaluating the objective function or its gradient, e.g., secret fixed point operations, are relatively expensive. In particular, L-BFGS may be considered to be based on the Newton-Rhapson method. Given a smooth objective function ^ and a point ^^, Newton-Rhapsondetermines a new point ^^^^which further minimizes the function: ^^^^ = ^^ − ^^^^^(^^)^^(^^).Here, ^^( ) denotes the gradient of ^ , ^^^ denotes the inverse of the hessian of ^, and ^denotes the step size for this iteration. While, in the non-MPC scenario, the inverse of the Hessian can in many cases be computed directly, it is preferred in MPC to avoid this because of thecomputational costs. For example, for logistic regression, the Hessian may include ^ × ^partial derivatives, computed based on the dataset FV, LB. Moreover, also inverting theHessian using is typically relatively expensive in MPC, much more so than in the non-MPC setting. These aspects make Newton-Rhapson expensive under MPC. As the inventors realized, many of these efficiency problems can to a large degree be remedied by using a quasi-Newton method, such as in particular L-BFGS. Instead of computing the Hessian directly and then inverting it, L-BFGS may effectively determine anapproximation to the inverse Hessian ^^^(^^) by keeping a history of results from previousiterations (for example, of at most or at least five previous iterations), and may use thisapproximation to determine the optimization direction w as^ = ^^^(^^)^^(^^).Here, the inverse Hessian may be explicitly computed under multi-party computation, but it is also possible to avoid this explicit computation by applying two-loop recursion under the multi-party computation. In particular, using two-loop recursion, the determination Dir of theoptimization direction w from the gradient ∇F under multi-party computation may beimplemented as illustrated by the following pseudo-code.Protocol. Multi-party computation implementation of Two-Loop-RecursionInput: History Output: An estimate of ^^^^^^(^^) [^] ← [^] + ([^^] − [^])[^^]end return [^] In the above pseudo-code, the notation [^] is used to denote that ^ is a secretvalue of the multi-party computation. The above pseudo-code can be implemented using known techniques for multi-party arithmetic, e.g., fixed-point arithmetic. Specifically, thecomputation of [^^] can be implemented by an oblivious inner product followed by anoblivious reciprocal of fixed point numbers. Similarly, other ^. , . computations shown in thealgorithm can be implemented by an oblivious inner product, and products by and byoblivious multiplication. The computation of [^] can be implemented by computing thenumerator and denominator and then applying an oblivious fixed point division. Determining the optimization direction for L-BFGS may in particular comprise checking whether a history of sufficient previous iterations is already available, e.g., if five iterations have already been performed. If this is not the case, the optimization direction can for example be determined by using gradient descent. As part of the L-BFGS history, also previous values of ^^may be kept. This is advantageous when using multi-party computation, since this value is relatively costly to compute. The figure further shows a step size determining operation SS, 460. Interestingly, using such an operation, the step size α may be determined adaptively per iteration, instead of using a fixed step size or a fixed schedule of step sizes, for example. Generally, using an adaptive step size can speed up convergence and thus reduce the number of iterations that is needed. This is an important advantage in the MPC setting, since iterations are typically relatively costly to perform. Such faster convergence can be attained when using gradient descent, for example. For some optimization techniques, in particular for quasi-Newton methods such as L-BFGS, using an adaptive step size not only speeds up convergence, but also helps ensure that the optimization converges at all. In other words, when using a fixed step size as is typically done in work on optimization under MPC, such methods may in many cases not converge, for example for some types of regression. By using the step sizedetermining operation SS proposed herein, such convergence problems can be solved whilestill having an efficient implementation under MPC. For example, the implementation is much more efficient than if it would be attempted to computed the step size through line search, as is normally done in the non-MPC setting. Such a line search may evaluate the objective function at several different potential candidates for the step size until one is found that provides an optimal improvement. In the MPC, such a solution may be costly since it involves evaluating the objective function multiple times. Interestingly, the proposed stepsize determining operation SS can avoid this. In particular, operation SS may not use thedataset FV, LB, e.g., may use the gradient ∇F and optimization direction w only.In particular, step size determining operation SS may determine the secretstep size α from the secret optimization direction w and from a secret curvature value of theobjective function using the multi-party computation. Specifically, the curvature value may be determined by determining a secret inner product between the secret optimization direction w and a secret gradient of the objective function. The step size α may be determined such that, the larger the curvature, e.g., the larger the value of the inner product, the smaller the step size. In particular, the step size can be determined by evaluating the following formula under multi-party computation: Interestingly, the above formula can be efficiently evaluated under multi-party computation using techniques that are known per se, e.g., using a privacy-preserving inner productprotocol; a privacy-preserving square root protocol; and a privacy-preserving reciprocalprotocol. In particular, this formula may be used in combination with an objective function that is self-concordant, for example the objective function for fitting a logistic or other regression model. Mathematically, it is known from the optimization literature that this step size formula can be used in combination with a variety of optimization methods; in particular, in combination with L-BFGS it is known to provide super-linear convergence. See W. Gao et al., "Quasi-Newton Methods: Superlinear Convergence Without Line Searches for Self-Concordant Functions", arXiv:1612.06965v3 (incorporated herein by reference, specifically the definition of self-concordance and of the curvature-adaptive step). The figure also shows an updating operation Upd, 490, that is configured toupdate, using the multi-party computation, the secret parameters IC, WT of the objectivefunction based on the secret optimization direction w and the secret step size α. Theupdating may comprise adding to the values of the secret parameters a scaling of theoptimization direction according to the secret step size, e.g., [^^^^] ← [^^] − [^][^].The following pseudo-code demonstrates determining the iterative numeric optimization discussed with respect to this figure, in the example where L-BFGS is used as a numeric optimizer, and the step size is computed according to the formula above:Protocol. MPC implementation of L-BFGS; curvature-adaptive step size; two-loop recursionInput: Smooth objective function ^, initial starting point [^^] ∈ ^(^), memory size ^, numberof iterations ^Output: Point Compute [^^(^^)] / / operation Gradfor ^ ← 1 to ^ do [^^^^] ← [^^] − [^][^] / / operation Updend return [^^]In the example, [. ] is used to denote secret values of the multi-party computation. Thesecret values may be represented as secret fixed-point values of the multi-party computation, and the operations shown in the example may be implemented by privacy- preserving fixed point arithmetic protocols, as also discussed elsewhere. The determined parameters IC, WT may represent the fitting of the ordinalregression model to a dataset FV, LB that is partly or fully secret. Although not shown in thisfigure, the parameters may be further used to apply the fitted model on inputs under themulti-party computation. In this setting, the parameters IC, WT are typically secret values ofthe MPC. The inputs are typically secret as well, but that is not needed per se. To apply the model, for example, a log-likelihood may be computed and optionally compared to a threshold value. Specifically, the secret step size may be computed SS from from a secret optimization direction w and from a secret curvature value of the objective function being optimized. Such a step size computed based on a curvature value, may be referred to as a curvature-adaptive step size. In particular, the curvature may be determined by determining a secret inner product of the secret gradient of the objective function, with the secret optimization direction. The step size may be computed to be decreasing in the amount of curvature: the more curvature, the smaller the step size. Interestingly, the provided techniques provide a step size for the numeric optimization that can improve the convergence of the numeric optimization, while being itself efficient to compute under multi-party computation. This is in contrast for example to line search methods, which typically comprise multiple evaluations of the objective function and are thus costly to perform under multi-party computation, especially for relatively expensive objective functions such as the objective function for fitting a logistic regression model. In particular, the step size may be computed according to the following formula: Preferably, the objective function is non-concordant, in which case this step size formula is known from the scientific literature to provide strong mathematical convergence properties. In particular, L-BFGS may be used as numeric optimization algorithm, in which case this choice of step size may provide super-linear convergence of the numeric optimization. Interestingly, the step size may be computed using relatively few invocations of multi-party secure arithmetic protocols. In particular, the step size computation may be implemented by an oblivious inner product; an oblivious fixed point square root; and anoblivious reciprocal. In particular, these operations FxIp, FxSqr, FxRc may be implementedto operate on secret fixed point numbers as known from the multi-party computation literature per se, for example in Design of large scale applications of secure multiparty computation: secure linear programming", S.J.A. de Hoogh. A technique for efficiently and accurately evaluating an exponential function under multi-party computation is now discussed. The inventors realized that, in the specific computational setting of multi-party computation, it is advantageous to evaluate the exponential function for a value x by using the approximation (1+x / K)^K. Here, K can be at most or at least 8, at most or at least 16, or at most or at least 32, for example. Preferably, the exponential function is applied to an input from a limited domain, e.g., to values of at most 2, at most 1, or at most 0.5. A lower bound may not be needed. As is known per se, the value (1+x / K)^K converges to the exponential as K approaches infinity. For larger values of x, also a relatively large value of K may be needed. The inventors realized, however, that in many practical cases where the sigmoid is used, the value x is typically relatively small. In this case, also for small values of K, (1+x / K)^K mayaccurately represent the exponential of x. This is in particular the case for the sigmoid that isused in regression, and especially when the input features are normalized and / or when regularization is applied to the model parameters. In such cases, a relatively small value of K suffices. Moreover, the inventors realized that the value (1+x / K)^K can be computed efficiently as a secret value of a multi-party computation. Only a relatively small amount of secure multi-party numeric operations, e.g., fixed-point operations, may be needed. Despite this, still, an accurate evaluation of the exponential function can be obtained. This performance / accuracy trade-off is different in the MPC setting than in the plain data setting, where different approximations are typically used. Optionally, the evaluation of the exponential function may be stored as a secret fixed-point number of the secure multi-party computation. The exponential function may be evaluated by applying secure multi-party computation protocols for fixed-point computation, in particular for fixed-point multiplication, as known per se. For example, a secret fixed-point value may be initialized to (1+x / K), and (1+x / K)^K may be computed from (1+x / K) by square-and-multiply under the multi-party computation. Interestingly, K may be set to a power of two, in which case the square-and-multiply can be implemented by repeated squaring, making computation under multi-party computation particularly efficient. Specifically, the exponential function may be evaluated by iteratively updating an estimate EXP of the value of the exponential function at the secret point x. The estimate can be stored as a secret fixed-point number, for example, as is known for example from the reference "Design of large scale applications of secure multiparty computation: secure linear programming", S.J.A. de Hoogh. The iterative updates may be performed by using a square-and-multiply operation Sqm. In particular, the exponential function may be estimated based on: for all ^ ∈ ℝ, and accordingly, an estimate of ^^^(^) may be obtained by computing the lefthand side of the above formula for up to ^ terms for a given value ^, e.g.,^ ^(^) = ^1 +^ ^ ^ . Here, the value ^ may be selected based on the desired accuracy and / or abound on the input value. Namely, for ^ > 0 the error of the estimation may be boundedfrom above by (^^ ∗ ^^)⁄ 2 ^. A value for ^ that provides the desired efficiency may bederived from this bound. In practice, a value of at most or at least 32, at most or at least 64, or at most or at least 128 can be selected, for example. In particular, having selected a value for ^, the estimate EXP may first beinitialized to 1 + ^⁄ ^ . This value may be computed by multiplying (e.g., fixed-pointrepresentations of) ^ by ^^^. Interestingly, since ^ is public, no division of secret numbersunder multi-party computation is needed, improving efficiency. The square-and-multiply operation Sqm may update the estimate EXP byperforming a square-and-multiply algorithm. It is preferred to select the parameter ^ as apower of 2, such that the square-and-multiply can be implemented by repeated squarings. Inthis case, the square-and-multiply Sqm may be implemented by ^^^(^) squarings, allowinga particularly efficient implementation under multi-party computation.Fig. 6 shows a block-diagram of a cryptographic method 1000 of performinga privacy-preserving computation. The privacy-preserving computation may be performed by a cryptographic device as a secure multi-party computation between multiple cryptographic devices comprising the cryptographic device. The privacy-preserving computation may comprise a computation of a likelihood gradient of an ordinal regression model for a record with respect to a set of parameters. The record may comprise one or more feature values and a label. The label may be comprised in a set of multiple ordered labels. The set of parameters may comprise one or more weights corresponding to the features and one or more intercepts corresponding to the labels. For example, the cryptographic device can be device 100 of Fig.1 or device 200 of Fig.2. However, this is not a limitation, in that the method 1000 may also be performed using another system, apparatus or device. The method 1000 may further comprise the carrying out of the secure multi-party computation by the other cryptographic devices. For example, the method 1000 may be carried out by a cryptographic system, e.g., cryptographic system 010 of Fig.3. The method 1000 may be computer-implemented. The cryptographic method 1000 may comprise, in an operation titled "OBLIVIOUSLY SELECT INTERCEPTS", using the multi-party computation, determining 1010 a lower intercept and a higher intercept by obliviously selecting the lower intercept and / or the higher intercept from the one or more intercepts based on the label of the record. The cryptographic method 1000 may comprise, in an operation titled "EXPONENTIATE FOR INTERCEPTS", using the multi-party computation, computing 1020 a lower exponential and a higher exponential by applying a multi-party exponentiation protocol based on the lower intercept and the higher intercept, respectively. The cryptographic method 1000 may comprise, in an operation titled "USE EXPONENTIALS IN GRADIENTS", using the multi-party computation, determining 1030 gradients of one or more weights and / or one or more intercepts based on the lower exponential and the higher exponential. It will be appreciated that, in general, the operations of method 1000 of Fig. 10 may be performed in any suitable order, e.g., consecutively, simultaneously, or a combination thereof, subject to, where applicable, a particular order being necessitated, e.g., by input / output relations. The method(s) may be implemented on a computer as a computer implemented method, as dedicated hardware, or as a combination of both. As also illustrated in Fig.7, instructions for the computer, e.g., executable code, may be stored on a computer readable medium 1100, e.g., in the form of a series 1110 of machine-readable physical marks and / or as a series of elements having different electrical, e.g., magnetic, or optical properties or values. The medium 1100 may be transitory or non-transitory. Examples of computer readable mediums include memory devices, optical storage devices, integrated circuits, servers, online software, etc. Fig.11 shows an optical disc 1100. The instructions may be instructions for one or more particular devices of the cryptographic system. In particular, the instructions may comprise instructions for a cryptographic device to perform a computation of a likelihood gradient for an ordinal regression model and / or to apply a fitted ordinal regression model to a record. Examples, embodiments or optional features, whether indicated as non- limiting or not, are not to be understood as limiting the invention as claimed. It should be noted that the above-mentioned embodiments illustrate rather than limit the invention, and that those skilled in the art will be able to design many alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. Use of the verb "comprise" and its conjugations does not exclude the presence of elements or stages other than those stated in a claim. The article "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. Expressions such as “at least one of” when preceding a list or group of elements represent a selection of all or of any subset of elements from the list or group. For example, the expression, “at least one of A, B, and C” should be understood as including only A, only B, only C, both A and B, both A and C, both B and C, or all of A, B, and C. The invention may be implemented by means of hardware comprising several distinct elements, and by means of a suitably programmed computer. In the device claim enumerating several means, several of these means may be embodied by one and the same item of hardware. The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage.

Claims

CLAIMS1. A cryptographic system (010) for performing a privacy-preservingcomputation on secret data,- wherein the cryptographic system comprises multiple cryptographic devices(321-323), wherein the multiple cryptographic devices are configured to perform the computation as a cryptographic secure multi-party computation between the multiple cryptographic devices,- wherein the privacy-preserving computation comprises computing a likelihoodgradient of an ordinal regression model for a record with respect to a set of parameters, wherein the record comprises one or more feature values and a label, wherein the label iscomprised in a set of multiple ordered labels, wherein the set of parameters comprises oneor more weights corresponding to the features and one or more intercepts corresponding to the labels,- wherein a cryptographic device (321-323) of the multiple cryptographicdevices is configured to compute the likelihood gradient by: -using the multi-party computation, determining a lower intercept and ahigher intercept by obliviously selecting the lower intercept and / or the higher intercept from the one or more intercepts based on the label of the record; -using the multi-party computation, computing a lower exponential anda higher exponential by applying a multi-party exponentiation protocol based on the lower intercept and the higher intercept, respectively; -using the multi-party computation, determining gradients of one ormore weights and / or one or more intercepts based on the lower exponential and the higher exponential.

2. The system (010) of claim 1, wherein the cryptographic device is configuredto further compute, using the multi-party computation, a reciprocal of the lower exponentialby applying a multi-party exponentiation protocol based on the lower intercept, and to usethe reciprocal of the lower exponential to compute gradients of one or more weights.

3. The system (010) of any preceding claim, wherein determining the gradientsof the weights and / or intercepts comprises computing, using the multi-party computation, a product of the lower exponential and the higher exponential.

4. The system (010) of any preceding claim, wherein the cryptographic device isconfigured to determine a unit vector representation of the label using a multi-party index-to- unit-vector protocol, and to obliviously select the lower and / or higher intercept using the unit vector representation.

5. The system (010) of claim 4, wherein the cryptographic device is configuredto use the unit vector representation in multiple computations of the likelihood gradient for the record.

6. The system (010) of any preceding claim, wherein the lower exponential andthe higher exponential are computed at least in part in parallel.

7. The system (010) of any preceding claim, wherein the cryptographic device isconfigured to determine gradients of one or more parameters of the set of parameters by:- determining, using the multi-party computation, a gradient denominator,based on whether the label of the record is the first label, the last label, or neither label of the set of ordered labels;- applying a multi-party reciprocal protocol to compute a reciprocal of thegradient denominator, and computing the gradients based on the computed reciprocal.

8. The system (010) of claim 7, wherein the gradients of the one or moreparameters comprise a gradient of a weight and / or a gradient of the intercept for the first label and / or a gradient of the intercept for the last label.

9. The system (010) of any preceding claim, wherein the cryptographic device isfurther configured to compute, using the multi-party computation, a class probability and / or a classification of the ordinal regression model for a further record.

10. The system (010) of claim 9, wherein computing the class probability and / orthe classification comprises computing a cumulative class probability for a class based on the set of parameters.

11. The system (010) of any preceding claim, wherein the cryptographic device isconfigured to fit the ordinal regression model to multiple records comprising the record, wherein the fitting comprises the computation of the likelihood gradient.

12. The system (010) of claim 11, wherein the fitting is performed as a numericoptimization, wherein the numeric optimization is performed by updating, using the multi-party computation, the set of parameters based on an optimization direction and a step size,wherein the cryptographic device is configured to determine the step size using the multi- party computation from the optimization direction and from a curvature value.

13. A cryptographic device (100, 321-323) for use in the cryptographic system(010) comprising multiple cryptographic devices according to any one of claims 1-12, wherein the cryptographic device is for performing a privacy-preserving computation as a cryptographic secure multi-party computation between the multiple cryptographic devices,- wherein the privacy-preserving computation comprises computing a likelihoodgradient of an ordinal regression model for a record with respect to a set of parameters, wherein the record comprises one or more feature values and a label, wherein the label is comprised in a set of multiple ordered labels, wherein the set of parameters comprises one or more weights corresponding to the features and one or more intercepts corresponding to the labels, wherein the cryptographic device comprises:- a communication interface (180) configured for communication with at leastone further cryptographic device of the cryptographic system;- a data interface (120) for accessing data (040) representing the set ofparameters (030) of the ordinal regression model;- a processor subsystem (140) configured to compute the likelihood gradientby: -using the multi-party computation, determining a lower intercept and ahigher intercept by obliviously selecting the lower intercept and / or the higher intercept from the one or more intercepts based on the label of the record; -using the multi-party computation, computing a lower exponential anda higher exponential by applying a multi-party exponentiation protocol based on the lower intercept and the higher intercept, respectively; -using the multi-party computation, determining gradients of one ormore weights and / or one or more intercepts based on the lower exponential and the higher exponential.

14. A cryptographic method (1000) of performing a privacy-preservingcomputation, wherein the privacy-preserving computation is performed by a cryptographicdevice as a secure multi-party computation between multiple cryptographic devices comprising the cryptographic device,- wherein the privacy-preserving computation comprises computing a likelihoodgradient of an ordinal regression model for a record with respect to a set of parameters, wherein the record comprises one or more feature values and a label, wherein the label is comprised in a set of multiple ordered labels, wherein the set of parameters comprises one or more weights corresponding to the features and one or more intercepts corresponding to the labels, wherein the cryptographic method comprises:- using the multi-party computation, determining (1010) a lower intercept and ahigher intercept by obliviously selecting the lower intercept and / or the higher intercept from the one or more intercepts based on the label of the record;- using the multi-party computation, computing (1020) a lower exponential anda higher exponential by applying a multi-party exponentiation protocol based on the lower intercept and the higher intercept, respectively;- using the multi-party computation, determining (1030) gradients of one ormore weights and / or one or more intercepts based on the lower exponential and the higher exponential.

15. A transitory or non-transitory computer-readable medium (1100) comprising data (1110) representing instructions which, when executed by a processor system, cause the processor system to perform the cryptographic method of claim 14.

Citation Information

Patent Citations

  • Privacy-preserving machine learning

    US20200242466A1

  • High-Precision Privacy-Preserving Real-Valued Function Evaluation

    US20200358601A1