System, method, and apparatus for estimating life-cycle impact while preserving privacy

A privacy-preserving federated learning system addresses data confidentiality issues in life-cycle assessments by allowing local training and secure aggregation of impact scores, enhancing accuracy and sustainability insights.

WO2025262125A1PCT designated stage Publication Date: 2025-12-26MERCK PATENT GMBH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/067078
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-19
Filing Date
2025-06-18
Publication Date
2025-12-26

AI Technical Summary

Technical Problem

Current life-cycle assessment (LCA) systems face challenges due to confidentiality concerns and data variability, leading to limited data sharing and inaccurate sustainability reporting, as companies are reluctant to contribute proprietary data to public databases.

Method used

A privacy-preserving federated learning system that allows local clients to train models on private data and generate impact scores, using techniques like stochastic gradient descent, secure multiparty computation, and homomorphic encryption to aggregate updates while maintaining data confidentiality.

Benefits of technology

Enables accurate and secure estimation of life-cycle impact scores across various dimensions, balancing data privacy with improved sustainability insights and model accuracy through collaborative learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000037_0001
    Figure IMGF000037_0001
Patent Text Reader

Abstract

A system, method, and apparatus for estimating life-cycle impact while preserving data privacy uses federated learning to train machine learning models across multiple local clients without exposing private data. Local clients train local models on private datasets. Privatized update data from these local models may be transmitted to a central server which aggregates this data to enhance a global model. The global model is redistributed to clients, thereby improving its accuracy through successive federated learning rounds, without accessing any raw private data. Local clients can query their local models to obtain impact scores for products or processes based on the aggregated learning. A range of privacy-preserving computation protocols limit exposure of sensitive data. This cooperative framework enables collective model refinement reflective of broad data access without compromising confidentiality. The system provides accurate impact values alongside data security assurances for clients.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEM, METHOD, AND APPARATUS FOR ESTIMATING LIFE-CYCLE IMPACT WHILE PRESERVING PRIVACYBACKGROUNDRelevant Field

[0001] The present disclosure relates to assessments, such as life-cycle impact. More particularly, the present disclosure relates to a system, method and apparatus for estimating an impact, such as a life-cycle impact, while preserving privacy of data contributors.

[0002] Description of Related Art

[0003] Life cycle assessments (LCAs) are methods for measuring the environmental impacts of products, processes or services over their entire life cycle. LCAs involve quantifying raw material and energy inputs, emissions, resource usage, and other releases across all stages from cradle-to-grave.

[0004] Variants of LCA, including cradle-to-grave, cradle-to-gate, and cradle-to-cradle, address different portions of the product lifecycle. Each variant adjusts the system boundaries and the scope of analysis to suit specific informational needs or stakeholder interests. For instance, cradle-to-gate assessments focus on the product lifecycle up to the point of sale, excluding usage and end-of-life stages, while cradle-to-cradle emphasizes a closed-loop approach where the end-of- life disposal phase is a recycling process. While LCAs including its variants provide valuable sustainability insights, companies have been reluctant to contribute proprietary data to public LCA databases due to confidentiality concerns.

[0005] Two key challenges with current LCA data are scale and variability. There are millions of potential materials and products to assess, yet databases have limited entries that may not represent site-specific conditions. Values for the same chemical can also vary widely across databases. These data issues, along with confidentiality concerns, disincentivize companies from sharing data and undermine the accuracy of sustainability reporting.

[0006] A solution that lowers confidentiality barriers to sustainability data sharing could accelerate LCA model development. Enabling companies to contribute data securely would lead to better insights on lower-impact alternatives across industries. There is a need for privacy -preserving methods tailored to the unique data and modeling requirements of LCAs.SUMMARY

[0007] In some embodiments, the product impact value generated by the system corresponds to a production activity of the product across its lifecycle. The production activity quantified through the impact value may involve material transport such as movement of raw materials or finished products between supply chain stages. It may additionally encompass material extraction activities like mining of metal ores or drilling for oil and gas. Furthermore, the impact value can reflect material handling operations including packaging, storage, or loading and unloading. Material use during manufacturing steps such as machining, casting, or chemical processing can also contribute to the product impact value. Additionally, end-of-life material disposal through approaches like landfilling, incineration, composting or recycling may play a role in determining the overall impact value for the product. In essence, the flexibility of the system allows product sustainability metrics to be tailored to specific production activities of relevance, ranging from resource extraction to disposal.

[0008] In some embodiments, the impact value estimated by the local model may directly be an impact score itself, without needing additional conversion. The impact score is a standardized, dimensionless value reflecting an environmental sustainability metric such as greenhouse gas emissions, energy consumption, water utilization, waste generation, or economic impacts associated with a product or process. For example, the impact value returned in response to a query could be a carbon footprint score on a universal scale of 1 to 100. The score enables local clients to readily evaluate and compare the sustainability of products, processes or services against internal and external benchmarks. By having the local model output actionable scores, organizations can seamlessly use the system to make decisions towards meeting their environmental performance goals while maintaining data privacy through federated learning.

[0009] In some embodiments, the system's impact value may be an optional value selected from a broad range of sustainability-related indicators across environmental, social, and economic dimensions. These optional impact values that may be estimated by the system include but are not limited to climate change, greenhouse gas emissions, carbon dioxide emissions, carbon dioxide equivalent emissions, ozone depletion, CFC-11 equivalent emissions, toxicity, human toxicity, particulate matter formation, ionizing radiation, photochemical ozone creation potential, acidification, water acidification, ocean acidification, sulfur dioxide emissions, nitrogen oxide emissions, ammonia emissions, eutrophication, terrestrial eutrophication, freshwater eutrophication, marine eutrophication, nitrogen runoff levels, phosphorus runoff levels, ecotoxicity, water utilization, resource depletion, mineral resource depletion, fossil fuel depletion, land use change, habitat loss, biodiversity loss, renewable energy use, non-renewable energy use, energy consumption, waste generation rates, total cost of ownership, market dependency, technological obsolescence rates, adaptability to climate change, resilience to disasters, crisis response capabilities, and many other environmental, social, or economic indicators across the product, process, or service life cycle. In this manner, the system provides flexibility in estimating a range of optional sustainability impact values while protecting data privacy.

[0010] In some embodiments, the central server further comprises an additional score component configured to generate one or more types of scores using the impact value estimated by the aggregated model. These scores may optionally include a health score quantifying impacts on human health, an ecosystem score indicating effects on wildlife and habitats, and a natural resources score reflecting the consumption of materials like minerals, metals, fuels, and water. By leveraging the impact value produced through privacy -preserving federated learning, the score component can translate the quantitative impact data into more readily interpretable sustainability scores across these various dimensions. Generating multiple scores allows assessment of environmental performance from different stakeholder perspectives. The health, ecosystem, and natural resource scores complement the aggregated learning of the distributed model while avoiding exposure of sensitive local data.

[0011] In some embodiments, the central server further comprises a score component configured to convert the impact value to an interpretable score using conversion factors obtained from a database. The score component translates the quantitative impact value into a dimensionless score along a standardized scale to indicate sustainability performance. The conversion factors that facilitate generating these readable scores are retrieved from a database accessible to the central server. Keeping a localized scoring capability allows each client to rapidly quantify impact data into ratings reflecting environmental benchmarks without needing to externally transmit scoring requests. However, the underlying logic and benchmarks derive from the centralized database to ensure uniformity. Enabling clients to efficiently contextualize raw impact values into readily understood scores complements the privacy-preserving impact valuation functionality offered by the system.

[0012] In some embodiments, each local client of the plurality of local clients further comprises a training component configured to train the local model using the local dataset to generate update data. The training component executes training algorithms, such as stochastic gradient descent, on the local datasets to update parameters of the local model. The training process generates update data corresponding to the changes in parameters after running multiple training steps on the local private dataset. Through this localized training process, each local client is able to train its local model using its own private dataset without needing to share that data externally. The update data reflects the learning achieved from this localized training.

[0013] In some embodiments, each local client further comprises a training component configured to train the local model using the local dataset to generate update data. The training component may fine tune the local model in order to train it using the local dataset. This fine tuning of the local model by the training component generates the update data that is later privatized and transmitted to the central server. Optionally, the fine tuning may involve additional training epochs focused on the local client's private data in order to enhance the performance of the local model on estimating impacts relevant to that client. Thus, the training component enables further customization of each local model through fine tuning tailored to each client's particular data. This fine tuning supplements the broader learning accumulated across clients into the aggregated global model, allowing both global trends and local specifics to be reflected in the local model’s impact estimations. The fine tuned training still generates update data to be privatized and transmitted to enhance the aggregated model while preserving privacy.

[0014] In some embodiments, each local client may further comprise a privatization component configured to privatize the update data generated from training the local model. This privatization component privatizes the update data by adding noise, such as Gaussian noise or Laplacian noise, to the update data before transmitting it. Adding noise assists in obscuring the update data to prevent reconstruction or inference of the private data while still allowing the update data to improve the aggregated model. Thus, the privatization component generates privatized update data from the initial update data through noise addition techniques in order to preserve privacy of the private data while enabling collaborative learning.

[0015] In some embodiments, each local client may have a privatization component configured to privatize the update data generated from training the local model. One technique this privatization component may utilize is clipping, in which updates exceeding a predefined threshold are scaled down to meet a set limit. By clipping the update data before transmission to the central server, this adds an additional layer of privacy preservation that prevents inference of the original private data from the update values. Thus, the privatization component on each client enables clipping of model updates as one method to transform sensitive training data into privatized versions that protect confidentiality while still allowing the central server to improve the aggregated global model.

[0016] The system of some embodiments may optionally involve each local client having a privatization component configured to preserve privacy. This privatization component clips the update data from the local model to be within a threshold and adds noise, such as Gaussian or Laplacian noise, to the clipped update data. By clipping and adding noise in this manner, the privatization component generates privatized update data to be transmitted to the central server. Thus, the privatization component provides additional privacy protection before the local client sends its update data through obscuring the updates while still allowing the central server to improve the aggregated model for better collective accuracy over successive learning rounds. This approach balances data privacy for each local client with enabling collaboration to construct a robust shared model reflecting insights from private data across participants. The privatization protocols followedby each local client privatization component mitigate inference risks and uphold stringent confidentiality assurances.

[0017] Embodiments of the system may involve the use of optional added noise by the privatization component of the local clients to generate the privatized update data. This added noise may optionally comprise Gaussian noise or Laplacian noise in some system implementations. The application of such statistical noise distributions can serve to further obscure the local update data to enhance privacy preservation. Thus, the privatization process performed by local clients may optionally make use of Gaussian noise or Laplacian noise insertion as a supplemental technique to protect sensitive private data within the local dataset while still allowing collaborative learning through aggregated model updates.

[0018] In some embodiments, the update data that is privatized and transmitted from each local client to the central server corresponds only to the differences between the current state of that client's local model and the previous aggregated model provided by the central server. Rather than transmitting entire model parameters or weights with each update, only the changes in the local model relative to the preceding global model are communicated. By restricting the updates to reflect merely the delta between successive versions, less information is exposed with each transmission from local clients to the central server. This enhancement in privacy and security is achieved by computing model differentials based on the present local model parameters and those from the most recently distributed aggregated model. Transmitting just the model differences ultimately allows the central server to reconstruct the new state of each local model and aggregate the updates while limiting the amount of information shared from local clients.

[0019] In some embodiments, the central server may further comprise an optimization component configured to minimize an error in the aggregated model. This can be achieved by adjusting a plurality of model parameters of the aggregated model, such as the weights and biases of a neural network model. The optimization component may leverage techniques such as stochastic gradient descent to iteratively tune these parameters in order to reduce the discrepancy between the model's predictions and the actual ground truth data. This has the effect of enhancing the overall accuracy of the aggregated model in estimating a diverse range of product or process impact metrics while preserving privacy across the network of clients. The optimization takes place solely on the server side using the privatized update data without accessing the raw private data from any individual client.

[0020] In some embodiments, the central server is configured to receive privatized update data from multiple local clients. An optimization component within the central server may apply weighting factors to the privatized update data received from each client prior to aggregating this data. These weighting factors are based on defined metrics specific to each local client that contribute data. As an example, metrics such as the quantity of data, variability of data, industry segment of the client, or historical accuracy may guide the weighting. By incorporating customizable weighting factors to the contributed data during aggregation, the subsequent updates to the global model can reflect a tailored impact from each local client in a personalized manner while still fully preserving privacy of the underlying data.

[0021] In some embodiments, the central server is configured to receive multiple privatized update data from the plurality of local clients. The central server further comprises an optimization component configured to apply a weighting factor to each of the received privatized update data based on a metric of the respective local client. One such weighting metric that may optionally be used corresponds to the size of the company associated with each local client. Larger companies may be weighted more heavily than smaller companies. Applying a weighting to the privatized update data based on company size allows the aggregated model to reflect the impact of company size on the learning process, while still preserving privacy across clients.

[0022] In some embodiments, the central server receives a plurality of privatized update data from the plurality of local clients. The central server is configured to apply a weighting factor to each client's privatized update data based on a metric of the respective local client. The metric by which the update data is weighted can correspond to the amount of data in the local client's local dataset, the variability of data in the local dataset, the industry segment of the local client, or the historical accuracy of update data from the local client. By applying weighting factors to each local client's update data based on these metrics, the server can tailor the impact of each client's data in the aggregated model while maintaining client data privacy.

[0023] In some embodiments, the central server may be configured to further enhance privacy protections by adding noise directly to the privatized update data received from each local client before aggregating it. This additional noise insertion into the already obscured updates provides an extra layer of security to prevent any reconstruction of the original private data. The noise added may follow statistical distributions such as Gaussian or Laplacian to mathematically guarantee differential privacy. By enabling the central server to insert randomized noise into the aggregated dataset, the system allows for tunable privacy while still accumulating collaborative learnings through federated optimization of the global model. The level of noise induced balances model accuracy with confidentiality goals. This optional privacy-preserving feature expands upon other protocols within the system to restrict any identifiable leakage of sensitive data through the learning process.

[0024] In some embodiments, the central server may be configured to add noise directly to the weights of the aggregated model as an additional privacy-enhancing measure. This noise addition technique can obscure the specific parameter values in the aggregated model, making it more difficult to isolate any individual client's contribution while still preserving the overall machine learning accuracy. The noise added to the model weights may be derived from statistical distributions such as Gaussian or Laplacian noise. By inserting noise into the aggregated model weights, the system provides an extra layer of security that builds on the privatized update data already transmitted from each client. Adding noise to the model weights before optimization is an optional functionality that provides enhanced protection in case the privatized client updates are compromised. This embodiment offers supplementary data protection to strengthen privacy while enabling federated learning for impact estimation.

[0025] In some embodiments, the central server is configured to communicate the aggregated model back to each of the local clients after updating it based on the privatized update data. Each of the local clients may then use this received aggregated model to update or replace its own local model for the next round of localized training. This allows the enhanced aggregated model reflecting collective learning across clients to be incorporated into the local models for further refinement. Through this distribution process, the central server enables all clients to benefit from the accumulated knowledge gained from aggregating the privatized updates provided by each client, while maintaining the privacy guarantees of the federated learning paradigm. Over successive rounds, this circular process of updating local models, generating and aggregating privatized updates centrally, and disseminating the improved global model back to clients, allows the system to iteratively improve all models to estimate life cycle impact values more accurately without external exposure of any raw private data.

[0026] In some embodiments, the central server may be configured to employ secure multiparty computation (SMPC) during model updates to allow for privacy -preserving computation and aggregation of the privatized update data. SMPC techniques allow collective computation on sensitive data without any node revealing its raw private data. Specifically, each local client may partition its local private dataset and create 'shares' of the data using cryptographic secret sharing schemes. These shares may then be transmitted to other clients, where computations occur directlyon the encrypted shares rather than the actual data. The shares are mathematically blinded in a manner that no single party can view the raw data belonging to another party. In the described system, the central server may act primarily as a coordinator node that facilitates the initialization of the SMPC protocols and guides the execution flow but cannot access any client's data during computations. The SMPC protocols obscure the individual contributions of each client during training, thereby converting their private updates into anonymized shares that immunize the sensitive information from reconstruction by any other party. The collective learning accumulated across training rounds continuously enhances the accuracy of the aggregated model while simultaneously guaranteeing privacy to all data owners.

[0027] In some embodiments, the system utilizes secure multi-party computation (SMPC) techniques during model updates to preserve privacy. SMPC allows collective computation on sensitive data without nodes revealing their individual private data points. Specifically, SMPC enables the clients in the system to compute an average across private inputs contributed from multiple clients without any one client exposing their raw private data to other clients or the central server. The secure multi-party computation protocols facilitate nodes performing useful analytics on combined datasets in a distributed environment with an untrusted central server while still achieving information-theoretic security for the clients' confidential data. Thus, the system leverages SMPC to enable an aggregation of insights from distinct data sources without compromising the privacy assurances for any individual data contributor.

[0028] In some additional embodiments, the central server utilizes homomorphic encryption (HE) for the transmission and aggregation of the privatized update data from each local client. Homomorphic encryption enables the local clients to decrypt the aggregated model after the central server aggregates the privatized update data, without the central server actually having access to the raw model updates themselves. In particular, the homomorphic encryption may employ the Paillier algorithm to facilitate the encryption process between the central server and the local clients regarding the transmission and aggregation of the privatized update data. The use of homomorphic encryption further enhances the overall privacy protections and reduces the risk of any sensitive private data being exposed during these communications. This framework allows the aggregated learning to be shared across clients while upholding strong guarantees around maintaining the confidentiality of any contributor's raw private data within the system.

[0029] In some embodiments, the central server utilizes homomorphic encryption (HE) for the transmission and aggregation of the privatized update data. Homomorphic encryption enables clients to decrypt the aggregated model without the aggregator accessing the actual model updates. The homomorphic encryption may employ the Paillier algorithm, which facilitates an encryption process allowing computations to be carried out on ciphertext to generate an encrypted result which, when decrypted, matches the result of operations performed on the plaintext. Using homomorphic encryption ensures that even after aggregation of the privatized local updates, no single client's updates can be isolated or identified, thereby preserving privacy while still improving the global aggregated model.

[0030] In some embodiments, the central server applies a differential privacy algorithm to ensure privacy guarantees during the aggregation process of privatized update data from the local clients. This algorithm utilizes a set gradient norm bound (C) and noise scale. The differential privacy algorithm introduces calibrated statistical noise to mask any individual privatized update data contribution, thereby limiting exposure risks related to sensitive information in any one local client's private dataset. By configuring appropriate hyperparameters around clipping thresholds and noise levels, the system can balance accuracy considerations against policy and legal privacy thresholds. The overall framework remains focused on enabling broad collective model improvements through federated learning while upholding confidentiality assurances.

[0031] In some embodiments, the central server applies a differential privacy algorithm with set parameters for managing privacy during the aggregation process. This is intended to provide privacy guarantees for the sensitive data from local clients. The differential privacy algorithm may utilize a privacy accounting method to compute an overall privacy cost for the system using the epsilon and delta parameters. The epsilon parameter determines the maximum privacy loss tolerated, while the delta parameter indicates the probability that the privacy loss will be higher than epsilon. By tracking these privacy costs, administrators can configure the system to maintain rigorous confidentiality standards for data contributors. Thus, the differential privacy algorithm may help quantify and bound the privacy risks to local client data during federated learning.

[0032] In some embodiments, to further enhance data security during transmission, model updates between the local clients and the central server are communicated by transmitting only the differences between the current local model and a previously aggregated model provided by the central server, rather than transmitting entire model states. By only transferring the changes in model parameters, such as differences in weight matrices or activation thresholds, the risk of exposing private data is reduced as opposed to transmitting full model snapshots, while still allowing the central server to accumulate the learning from various local clients into the aggregated model. This approach limits the volume of update data exchanged across potential insecure channels and obscures the specific elements being tuned during training rounds. Consequently, communicating only inter-model deltas bolsters privacy protections around sensitive local datasets utilized to refine local models as part of the overall federated learning paradigm. The computation of these difference updates may be handled by software components within both the central server and the individual local clients.

[0033] In some embodiments, model updates between the local clients and the central server are transmitted by communicating only the differences between models, thereby enhancing data security during transmission. Specifically, the differences between models that are communicated are computed based on the current local model on each local client and a previously aggregated model provided by the central server from a recent round of federated learning. By transmitting only model differences, less information is exposed during data transfer across the network, providing an additional layer of privacy protection. This allows the central server to update the global aggregated model to reflect the latest changes in each local model without needing to access the raw private data from any local client.

[0034] In some embodiments, the central server employs a stochastic gradient descent (SGD) algorithm with differentially private updates to minimize an error in the aggregated model while preserving privacy. The stochastic gradient descent algorithm applied by the central server may include steps for sampling the private data, computing gradients based on the loss function, clipping or bounding gradients if they exceed a threshold to preserve privacy, adding noise sampled from distributions like Laplace or Gaussian to further obscure the gradients and prevent reconstruction of the private data, and updating the parameters of the aggregated model by descending along the privatized gradients. By iterating through multiple rounds of local private dataset sampling to compute clipped, noised gradients that guide model updates in aggregate optimization, the stochastic gradient descent method allows the central server to continuously enhance the global aggregated model while simultaneously preserving the privacy of individual raw private data points from each local client. The additional application of differential privacy techniques shields the model improvement process, minimizing the ability to trace updates back to any local private data via reconstruction or inference attacks. Thus the system is able to minimize global model error and boost accuracy of life cycle impact scores over federated learning rounds without ever directly accessing local private datasets.

[0035] In some embodiments, the central server employs a stochastic gradient descent (SGD) algorithm with differentially private updates to minimize an error in the aggregated model while preserving privacy. This SGD algorithm may involve several key steps. First, the algorithm samples mini-batches of private data from the local datasets. Next, gradients of the loss function with respect to the model parameters are computed for each data sample. The gradients are then clipped by their norm to limit the influence of any single data point. Noise drawn from distributions like Gaussian or Laplacian is added to the clipped gradients to obscure the data. Finally, the noised gradients are used to update the model parameters through gradient descent. By iterating through sampling, gradient computation, clipping, noise addition, and descent, the SGD algorithm enables collaborative improvement of the aggregated model while maintaining privacy guarantees.

[0036] In some embodiments, the central server may be configured to compute an aggregate mathematical function over the private data from the local clients using secure transmission methods, without revealing any of the individual private data points. For example, the central server could calculate an average value across private data entries from multiple clients. To preserve privacy, this computation of an aggregate statistic can leverage secure multiparty computation protocols, homomorphic encryption schemes, or other cryptographic techniques to obscure individual data entries while still allowing useful aggregated outputs. Through these secure transmission mechanisms that keep individual private data obscured, the central server is able to derive collective insights reflective of broader learning across clients, without being able to view or infer any single client's confidential data.

[0037] In some embodiments, the central server may adjust the weighting factor applied to each local client's privatized update data based on a reliability metric associated with that client. When aggregating the privatized updates to enhance the global model, the central server can vary the influence of each client's contributions based on an assigned reliability score. This score rates the trustworthiness of the data or accuracy of the predictions from that client. Clients providing higher quality, more consistent updates may receive a higher weighting in the aggregated model, while those with sparse, noisy data may have a lower weighting. By tuning each local model's weighting factor according to reliability metrics tied to that specific client, the system allows for a more nuanced integration of the privatized updates into the global model. This enables tailored impact reflecting the credibility of each data source, while still preserving privacy across all participants.

[0038] In some embodiments, the central server and the local clients may utilize Secure Multi-Party Computation (SMPC) protocols to prepare local model updates for transmission, thereby preserving privacy by obscuring individual updates. These SMPC protocols can allow the local clients to add cryptographic noise to their local model updates before transmitting them. The noise obscures the updates to prevent inference of the underlying private data while still retaining overall patterns to improve the aggregated model. Additionally, the central server can aggregate the privatized updates from the local clients using SMPC techniques. This allows computing collective functions like averages without revealing any individual client's raw private data points. The SMPC protocols enable privacy-preserving distributed computation and transmission of local model updates as well as secure aggregation of these updates to train the global model, without exposing the proprietary data from any single local client.

[0039] In some embodiments, the system utilizes Secure Multi-Party Computation (SMPC) protocols to prepare local model updates for transmission while preserving privacy. As part of these protocols, each local client performs operations to add cryptographic noise to their local model updates before transmitting them. Adding noise obscures the actual update values, making it difficult to isolate or identify any individual client's contributions. This cryptographic noise injection prevents external inference of the private data while still allowing the central server to aggregate theobscured updates to improve the global model. The cryptographic perturbations induced by the SMPC protocols immunize sensitive information from reconstruction by any party other than the originating client. Thus, the system leverages noise injection techniques as one method within the suite of SMPC protocols to enable decentralized learning while protecting the privacy of each participant's local dataset.

[0040] The system may further utilize Secure Multi-Party Computation (SMPC) protocols to preserve privacy during the federated learning process. As part of these protocols, in some embodiments each local client may perform secure transformations on the local model updates prior to transmission. These secure transformations act as an additional privacy-preserving measure by generating secure, shareable versions of the local update data. By obscuring the raw updates through cryptographic techniques, the resulting privatized updates prevent any external inference of the private local data while still allowing the central server to aggregate the updates to improve the global model. Thus, the SMPC protocols enable a layer of security at the local client level before transmission to bolster the overall privacy assurances of the system.

[0041] The system may involve the central server distributing the aggregated global model back to the local clients in an encrypted manner such that each local client can only decrypt the portions of the aggregated model relevant to update their own local model. This cryptographic process maintains the overall confidentiality of the global aggregated model while allowing the local clients to access the relevant updated information to enhance their local models. By enabling selective decryption of only the required portions of the aggregated model, the privacy and security of the sensitive model is preserved throughout the process of federated learning between the central server and local clients.

[0042] In some embodiments, the privatized update data from a subset of local clients may be aggregated together and transmitted to the central server as a batch. By aggregating the privatized update data from multiple local clients into a batch before transmission, this helps obscure the contribution of any individual local client to the central server. This batch transmission of aggregated privatized updates from subsets of clients provides an additional layer of privacy protection by avoiding sending individual updates, which could potentially allow inference of private data. Instead, grouping updates anonymizes individual contributions to further preserve confidentiality of sensitive information in the local datasets.

[0043] In some embodiments, each local client's processor is configured to execute a training process on the local model using the local dataset to generate update data corresponding to the private data. The local client's processor then performs computations in accordance with the Secure Multi-Party Computation (SMPC) protocol to generate a secure, shareable version of the update data, thereby defining the privatized update data. These computations involve adding cryptographic noise or performing secure transformations to obscure the update data to prevent inference of the private data. The central server's processor is configured to aggregate the privatized update data from the multiple local clients using SMPC protocols, ensuring the aggregation process does not reveal any local update data from any of the clients. The central server then updates the aggregated model to form an updated version reflecting the combined learning from all the local clients, without accessing the private data of any client. Finally, the central server distributes the new version of the aggregated model to the local clients for updating their respective local models.

[0044] In some embodiments, each local client may further comprise a data preprocessing component configured to transform the local dataset before using it to train the local model. This data preprocessing component can carry out various operations such as data normalization, feature extraction, handling of missing values, and dimensionality reduction. Techniques such as principal component analysis, independent component analysis, and t-distributed stochastic neighbor embedding may optionally be used by the data preprocessing component to reduce thedimensionality of the local dataset while preserving patterns that are valuable for training. The goal of this data preprocessing is to manage the complexity of the environmental impact data and transform the local dataset into a suitable format that allows effective training of the local model without compromising privacy. By extracting insights from the raw data while respecting confidentiality, the data preprocessing facilitates building an accurate local model.

[0045] In some embodiments, each local client further comprises a data preprocessing component configured to transform the local dataset prior to training the local model. The data preprocessing component may be configured to perform common data processing techniques such as data normalization to rescale values, feature extraction to derive informative variables, dimensionality reduction to simplify complex data, and handling of missing values in the dataset. For example, dimensionality reduction may be achieved via standard methods like principal component analysis, independent component analysis, or t-distributed stochastic neighbor embedding. These preprocessing steps help manage data complexity and formatting inconsistencies, ensuring effective model training without exposing private data.

[0046] In some embodiments, the data preprocessing component may utilize dimensionality reduction techniques on the local dataset prior to training the local model. Dimensionality reduction serves to simplify complex high-dimensional data into more manageable lower-dimensional representations. Specifically, the dimensionality reduction may be performed using principal component analysis, independent component analysis, or t-distributed stochastic neighbor embedding. Principal component analysis employs an orthogonal linear transformation to convert a set of possibly correlated variables into linearly uncorrelated principal components that capture maximal variance. Independent component analysis statistically separates a multivariate signal into additive subcomponents based on the assumption that the subcomponents are non-Gaussian and independent from each other. T-distributed stochastic neighbor embedding focuses on preserving similarities between nearby points and differences between distant points when embedding highdimensional data into a two or three-dimensional space for visualization. By transforming the raw data into insightful reduced features, these techniques facilitate more efficient and effective model training while also obscuring potentially sensitive details within the original high-dimensional dataset.

[0047] In some embodiments, the central server further comprises a validation component that is configured to assess the accuracy of the aggregated model. This validation component evaluates the aggregated model using a held-out test dataset that was not included in the original training data. By testing the aggregated model's predictions against this unseen holdout data, the validation component can quantify the model's ability to provide accurate impact value estimates for new data samples outside of the original training distribution. This validation process serves to gauge the real-world generalization performance and reliability of the aggregated model in providing sound impact assessments across diverse situations while preserving privacy. The validation component's evaluations based on an impartial test dataset also help safeguard against the model being overfitted to the training data in a way that could compromise its efficacy for practical applications.

[0048] In some embodiments, the central server's validation component may also be configured to quantify privacy risk of the aggregated model. This would allow assessing the degree to which the aggregated model might enable inference about individual clients' private data. By evaluating privacy exposure alongside accuracy, additional assurances regarding sensitive data confidentiality could be provided. Potential techniques for quantifying privacy risk include computing differential privacy budgets, tracking membership inference susceptibility, and evaluating model explanations. Incorporating privacy risk quantification into the validation process enables balancing utility and protection goals.

[0049] In some embodiments, the system may contain a monitoring component within the central server to evaluate fairness and bias metrics for the aggregated model. Specifically, the central server may optionally incorporate a monitoring component designed to determine whether the aggregated model behaves in a fair and unbiased manner when queried for impact value estimations. Through assessing mathematical properties of the model and analyzing outputs across various subgroups, the monitoring component can quantify biases and imbalances in predictions to ensure just access to capabilities.

[0050] Some embodiments of the system may provide an explanation of the reasoning behind an impact value prediction generated through a query component to enhance interpretability. Specifically, the query component that interfaces with either the local model or the global aggregated model to extract impact value estimates may also be configured to return an explanation detailing the rationale for the prediction provided. This functionality within the query component allows the system to describe the contributing factors, data provenance, model logic, and other justifications that substantiate a particular impact value forecast by the model. By elucidating the why behind impact predictions, the system enables users to not only obtain sustainability metrics through queries but also gain transparency into the models decision-making processes and build appropriate trust in the estimations. The explanations complement the accuracy of the model with increased intelligibility.

[0051] In some embodiments, the central server further comprises a user interface configured to visually depict life cycle stages corresponding to an impact value. This user interface may allow users to see a graphical representation of the various life cycle stages, such as raw material extraction, manufacturing, distribution, product use, and end-of-life, that contribute to the overall impact value associated with a product or process. The interface can illustrate the breakdown of impacts across these different life cycle phases to provide further insights. For example, the interface may show that manufacturing accounts for the majority of carbon dioxide emissions whereas product disposal represents only a small fraction. This ability to visualize the distribution of an impact value across the life cycle can assist stakeholders in identifying areas to target for sustainability initiatives. By leveraging data visualizations, the user interface transforms complex, multidimensional impact data into intuitive graphics to support data-driven decision making.

[0052] In some embodiments, the private update data that is transmitted from each local client to the central server is comprised of differentially private stochastic gradient updates. This means that before the local model update data is communicated externally, differential privacy techniques are applied to introduce carefully calibrated noise to the model gradients to prevent reconstruction of the original private data while still allowing useful learning. Random noise sampled from distributions like Gaussian or Laplace is added to the local stochastic gradient updates during training based on a select privacy budget. This privatization step enables sharing the broad patterns from local learning with the central server without compromising privacy. The aggregated model thus benefits from the common signal present across noisy model updates from several local clients without being able to actually discern individual private data points, thereby enhancing utility for all participants simultaneously with strong privacy protections for each.

[0053] The system may employ federated learning protocols for decentralized training between the transmitter and receiver. In some embodiments, the transmitter in each local client and the receiver in the central server utilize protocols designed specifically to enable decentralized training of machine learning models across multiple local nodes. This decentralized approach facilitates the collaborative development of an aggregated global model through successive rounds of localized training on distributed private datasets, centralized model aggregation, and redistribution of the enhanced model back to clients. The federated learning protocols enable thisiterative transfer learning process in a peer-to-peer manner without requiring access to raw private data in individual clients.

[0054] In some embodiments, the system involves a method for estimating impact while preserving privacy. The method may involve training a local model using a local dataset containing private data at each of a plurality of local clients. The method may then involve generating privatized update data from the local model at each of the plurality of local clients. The method may also involve transmitting the privatized update data from each of the plurality of local clients to a central server. Additionally, the method may involve receiving the privatized update data from each of the plurality of local clients at the central server. Finally, the method may involve updating an aggregated model at the central server by aggregating the privatized update data from the plurality of local clients without accessing the private data in the local dataset of any of the plurality of local clients.

[0055] In some embodiments, the method may further involve querying the aggregated model at the central server in order to receive an impact value. This querying of the aggregated model preserves the privacy of the private data from all of the plurality of local clients. By querying the aggregated model that reflects the collective learning of the local clients rather than accessing the raw private data itself, the central server can generate useful impact values without compromising the confidentiality of any individual client's private data. Thus, querying the global model that has been updated through multiple rounds of federated learning enables valuable analytics while still upholding stringent privacy protections for all data contributors.

[0056] In some embodiments, the method may further involve the central server distributing the updated aggregated model back to the plurality of local clients after aggregating the privatized update data. By sending the enhanced aggregated model to each local client, this allows the local clients to incorporate the improved global model into their local versions. The central server is configured to communicate the most current aggregated model through the network to each participating client. Upon receiving this updated aggregated model, the individual local clients have the option of utilizing it to replace their existing local model. This distribution and replacement process provides a mechanism for the centralized learning to be propagated across all clients, thereby enhancing the accuracy of impact value predictions for every participant in a collaborative manner while preserving privacy. Over successive rounds of federated learning, whereby privatized updates are aggregated centrally then redistributed locally, the federated network collectively builds an increasingly sophisticated model for life cycle impact assessments without exposing raw private data.

[0057] In some embodiments, the method may further involve distributing the updated aggregated model back to the plurality of local clients after it has been enhanced through the federated learning process. Upon receiving the improved aggregated model, each local client may optionally replace its existing local model with this updated aggregated model. This replacement of the local model enables the local client to leverage the latest version of the model, which has aggregated learning from across the network without exposing any raw private data. By continually updating the local models with the refined aggregated global model, the overall accuracy continues to increase over successive rounds as more collective knowledge is incorporated. Allowing each local client to swap its local model for the most recently updated aggregated model is an additional step that may be included as part of the federated learning method in order to propagate the benefits of accumulated collaborative learning back to individual local nodes. This model replacement process aligns with the system's objectives to balance data privacy preservation alongside enhancing utility for sustainability impact estimation.

[0058] In some embodiments, the method may involve querying the aggregated model at the central server for the purpose of updating the aggregated model. During the process of updatingthe aggregated model by aggregating the privatized update data from multiple local clients, the central server may perform queries on the aggregated model. These queries by the central server on the aggregated model could be used to test and validate the performance of the model during the training process, in order to ensure that the accuracy continues to improve with each update to the model. Allowing querying of the aggregated model at the central server optionally provides a mechanism to monitor the modeling progress in terms of predictive capability.

[0059] In some embodiments, the method may involve querying the local model before training the local model using the local dataset. This allows each local client to leverage the learning already accumulated within its local model to estimate impacts without needing to externally transmit requests or data. By enabling low-latency queries directly through the local model, the method can facilitate rapid sustainability assessments while preserving the privacy of the local private data. The local model that is queried prior to the training may be identical to the aggregated model previously distributed by the central server. Thus, querying the local model prior to additional localized training provides a means to extract impact insights reflective of the global learning while maintaining data confidentiality.

[0060] In some embodiments, the method involves querying the local model prior to training it using the local dataset. In these embodiments where the local model is queried before the training process, the local model may be an identical copy of the aggregated model that was distributed to the local client from the central server. Using an identical copy allows the local client to evaluate the performance of the aggregated model on its local private dataset before updating the model. This enables a comparison between the localized private data predictions versus the aggregated model's predictions, serving as a diagnostic check while still preserving privacy as sensitive data remains entirely on the local client. The identical versions facilitate this assessment of the generalized aggregated model's particular applicability to the individual local client. After checking the unmodified aggregated model's accuracy, the local client then proceeds to train the local model on its dataset to customize and tune the model to its private data.

[0061] In some embodiments, the impact value estimated by the aggregated model is one of several sustainability metrics. These may include, but are not limited to, a greenhouse gas emission value denoting the amount of greenhouse gases emitted across the life cycle of a product or process. It could also be an energy consumption value quantifying the total energy demand including renewable and non-renewable sources. Additionally, the impact value may represent the water utilization across the supply chain and usage lifecycle stages measured as a water use value. Another example is a waste generation value reflecting the quantity of solid or liquid wastes created. Furthermore, the impact value can be an economic impact value encapsulating costs and externalities such as potential liabilities associated with the life cycle.

[0062] In some embodiments, the private data used to train the local models corresponds to product data. The product data may contain proprietary or commercially sensitive details about materials, manufacturing processes, supply chains, or other aspects that relate to specific products. As the local models train on this private product data in a federated learning approach, the estimated impact value they output reflects product-level environmental sustainability metrics. For instance, the impact value could quantify energy consumption, waste generation, or emissions associated with raw material extraction, fabrication, usage, and disposal stages of a manufactured good. Thus, by leveraging private local datasets containing granular confidential product details, the system can estimate product impact values through collaborative training of an aggregated model, without compromising the privacy of any individual contributor's sensitive commercial product information.

[0063] In some embodiments, the impact value estimated by the aggregated model corresponds specifically to a product impact value. This product impact value represents an environmental or social sustainability metric associated with a particular product across its lifecycle.The product categories can span various industries, such as food, apparel, electronics, chemicals, and more. The private data used in the localized training process to enhance the aggregated model may consist of confidential product information proprietary to a company, such as detailed bills of materials, manufacturing specifications, usage profiles, or end-of-life disposal guidelines. The product impact value generated upon querying the trained aggregated model quantifies sustainability aspects related to raw material extraction, production processes, distribution channels, consumer use phase, and eventual disposal or recycling. In essence, the queried impact value may characterize the holistic environmental footprint attributable to a product across its existence. The improved accuracy of this product impact valuation stems from the collective learning enabled by the federated learning paradigm, while simultaneously preserving data privacy and confidentiality.

[0064] In some embodiments, the impact value estimated by the system corresponds specifically to a product impact value. This product impact value may be associated with a production activity related to the product over its lifecycle. For example, the product impact value could quantify sustainability metrics affiliated with activities such as material transport, extraction, handling, usage, or disposal across the various stages from raw material sourcing through manufacturing, distribution, consumer use and end-of-life management. By enabling accurate assessment tailored to different production activities involved in a product system, the system provides granular visibility into environmental performance that can empower targeted impact mitigation efforts.

[0065] In some embodiments, the product impact value estimated by the system corresponds specifically to environmental impacts associated with production activities of the product across its lifecycle. These production activities, which contribute to the overall impact value computed, may include but are not limited to: material transport such as movement of raw inputs between supply chain nodes, material extraction activities like mining or harvesting of resources, material handling operations including storage and packaging, material usage during manufacturing or consumer utilization phases, as well as material disposal via techniques like landfilling, incineration, or recycling.

[0066] In some embodiments, the method further comprises where the impact value generated from querying the model is an impact score. The impact score serves as an interpretable rating that encapsulates the environmental, social, or economic impact into a standardized value on a defined scale. Conversion factors or scoring formulas may be applied to the raw impact value to generate a correlative impact score that facilitates contextualization, comparison to benchmarks, and comprehension of the sustainability metric. By translating the quantitative impact value into a more readable score, users and stakeholders can readily understand the performance and progress in relation to targets. Thus, the impact value, when processed into an actionable score, enables straightforward evaluation and data-driven decisions regarding products, processes or services analyzed through the privacy-preserving distributed learning system.

[0067] In some embodiments, the method may further comprise generating a life-cycle score at the central server using multiple queries to the query component that receives multiple impact values, including the previously mentioned impact value. The query component can receive the series of impact values through conducting a plurality of queries. The central server then utilizes these multiple quantified impact values to calculate an aggregated life-cycle score. This score serves as an overall sustainability indicator across various impact categories reflected in the individual impact values. By performing successive queries, each returning measurements of distinct environmental or social effects, the system can produce a holistic score encapsulating the full spectrum of influencers across the product or process life cycle. The algorithm to derive the composite score from the impact values may rely on weightings or conversion factors stored in anaccessible database. Thus, through leveraging the locally trained models to rapidly obtain multiple impact valuations, the central server can integrate these into interpretable overall life cycle scores.

[0068] In some embodiments, the method may involve generating a life-cycle score at the central server using multiple queries to a query component that is configured to receive multiple impact values, including the previously mentioned impact value estimated by the aggregated model. Each individual query of the multiple queries made to the query component may correspond to extracting only a single impact value of the multiple impact values provided by the aggregated model. In this approach, each distinct impact value generated by the aggregated model in response to a query reflects a separate sustainability dimension, enabling a multidimensional life-cycle sustainability assessment through the multiplicity of queries while preserving privacy across all data contributors.

[0069] The central server may generate additional interpretable scores based on the impact value, beyond the environmental sustainability score. In some embodiments, the central server uses the impact value that was produced by querying the aggregated model to generate scores that provide insights into health, ecosystem quality, and natural resource usage associated with a product or process. For example, the impact value may indicate greenhouse gas emissions, and the central server may use this value to compute a health score related to human toxicity or respiratory impacts from those emissions. Similarly, an ecosystem score could capture biodiversity loss or habitat damage associated with raw material extraction reflected in the impact value. Additionally, a natural resources score may reflect factors such as water depletion or mineral resource consumption related to a product lifecycle. By processing the impact value into more domain-specific scores, the central server enables easy understanding of how a product or process affects people, ecosystems, and resource availability, in addition to environmental sustainability. The scores distill the impact values into actionable metrics for multiple stakeholders.

[0070] In some embodiments, the method involves the central server converting the impact value to a score using a conversion factor obtained from a database. More specifically, after an impact value is received at the central server through querying the aggregated model, the central server may utilize its score component to generate an interpretable score based on the impact value. The score component applies a conversion factor or formula retrieved from the central database to translate the quantitative impact value into a standardized score along a predefined scoring system, thereby enabling contextualization of the impact value as a performance rating. For instance, a carbon footprint impact value denoted in kilograms of carbon dioxide equivalent emissions may be converted into a score between 1 and 100 reflecting environmental sustainability. This conversion of the raw impact value into a readable score assists stakeholders in evaluating products, processes or services against targets or alternatives and in making informed decisions. The database that provides the conversion factors and scoring logic to the central server may be synchronized with or replicated from databases present on the individual local clients.

[0071] In some embodiments, the method may further involve each local client of the plurality of local clients training the local model using the local dataset to generate update data. This additional local model training process may use the processor of each local client to execute a training component. The training component performs computations like gradient descent to update parameters of the local model in order to minimize a loss function and improve predictions on the local private data. By running multiple training iterations, each local model generates update data reflecting the new state of the model parameters after optimization on the local dataset.

[0072] In some embodiments, the method may further involve fine-tuning the local model at each local client to thereby train the local model using the local dataset to generate the update data. Specifically, each local client may utilize techniques to fine-tune its local model parameters using the private data in order to enhance the training process and yield improved update data. Thisfine-tuning acts to specialize the local model to the characteristics and patterns found within the specific local dataset, allowing it to make better predictions on the local data. The fine-tuned model then undergoes further training epochs where gradients are computed over mini-batches of the local private data, ultimately producing refined update data. By adapting the local model to its local data context through fine-tuning, each client is able to generate higher quality updates which better capture the uniqueness of its data while still preventing any exposure of the raw private information. The specialized updates may then contribute to improving the global aggregated model.

[0073] In some embodiments, the method may involve each local client privatizing the update data to generate the privatized updated data by adding noise to the update data. Adding noise may be performed at each local client of the plurality of local clients. The noise addition serves to obscure the original update data values to prevent inference of the private data while still retaining overall trends and patterns to improve the aggregated model's performance. In this manner, the privatization step acts as an additional privacy-preserving layer following the generation of update data from training of the local models. The specific noise added may be Gaussian noise, Laplacian noise, or other forms of randomly generated noise. By introducing noise to the model updates before transmission to the central server, the confidentiality of the raw private data at the local clients can be maintained throughout the federated learning process.

[0074] In some embodiments, privatizing the update data at each local client of the plurality of local clients to generate the privatized updated data may involve clipping the update data. The clipping process sets an upper limit to the magnitude of the update values that can be transmitted from each local client. This ensures that the updates do not reveal excessive detail about the private training data while still allowing the central server to improve the aggregated model. By clipping the gradients or weight updates during training at each local client prior to transmission to the central server, an additional layer of privacy and security is introduced to prevent reconstruction of sensitive local data. Thus, the privatization techniques utilized in the system and method may optionally include clipping or capping the locally generated update data before it is anonymized into privatized form and communicated externally.

[0075] In some embodiments, the method may further involve each local client clipping the update data and then adding noise to the clipped update data in order to generate the privatized update data that is transmitted. Clipping the update data refers to scaling down any updates that exceed a predefined threshold value. Adding noise, such as randomly generated numbers from a Gaussian distribution, helps to obscure the original update data to prevent reconstruction or inference of the private data while still retaining overall trends and patterns that can improve the aggregated model. By clipping and adding noise in this manner within each local client prior to transmission, an additional layer of security and privacy preservation is introduced into the federated learning process described herein.

[0076] In some embodiments, the method involves adding noise to the update data to generate the privatized update data. The added noise may be Gaussian noise or Laplacian noise. Specifically, each local client clips and adds noise to the update data to generate the privatized update data that gets transmitted to the central server. By adding noise, such as Gaussian or Laplacian noise, to the clipped update data, privacy of the private data in the local datasets can be preserved while still allowing the central server to aggregate the updates to enhance the aggregated model. The addition of carefully calibrated noise ensures that the privatized updates reflect broader learning trends without allowing inference of any local client's raw private data.

[0077] In some embodiments, the privatized update data that is transmitted from each local client to the central server corresponds only to the differences between the current state of the local model on that client and the most recently distributed aggregated global model. By communicating just the changes in the local model parameters rather than the full set of parameters, the amount ofdata exchanged is reduced. This serves to enhance the efficiency of transmission across the network. Additionally, by limiting the transmitted data to reflect only model differences, the exposure of information from the local private datasets is constrained even in the event of data interception. The computation of differences may rely on version histories of models to align the relevant parameters for comparison. The application of differential privacy techniques on these model differences provides further privacy protection before transmission to the central server. Thus, the system design choice to transmit only model parameter differences allows efficiency and defensive privacy advantages.

[0078] In some embodiments, the central server may incorporate an optimization component configured to minimize an error in the aggregated model. This can be achieved by adjusting a plurality of model parameters within the aggregated model based on the received privatized update data from the multiple local clients. The optimization component may leverage optimization algorithms such as stochastic gradient descent to tune the weights and biases of the aggregated model to enhance its accuracy in estimating impact values reflective of the cumulative learning across clients. Optional noise insertion directly into the aggregated model weights or privatized client updates before optimization provides an additional layer of privacy protection alongside utility gains from the optimization process. Through successive rounds of federated learning, the accuracy improvements from ongoing optimization combine with strong privacy guarantees to advance a collaborative analytical framework benefitting stakeholders.

[0079] In some embodiments of the method, the central server receives multiple privatized update data transmissions from the plurality of local clients, including the privatized update data from each local client. Upon receiving this collection of privatized update data, the central server may apply a weighting factor to the privatized update data from each local client based on a metric or measure specific to that local client. This metric used for determining the weighting could for example correspond to a size of a company associated with the local client. Alternatively, the metric may involve the amount of data, variability of data, industry segment, or historical accuracy of update data from the respective local client's dataset. By weighting the contributions from each local client differently based on such reliability metrics in this manner, the subsequent aggregation and updates to the global model can reflect a more customized, differential impact from each local client, while still preserving the overall privacy assurances.

[0080] In some embodiments, the central server may apply a weighting factor to each client's privatized update data based on certain metrics associated with the respective client. One such metric that may be used is the size of the company represented by the local client, such that clients from larger companies may be assigned a higher weighting factor. This allows the aggregated model to reflect the potentially greater quantity, diversity, or reliability of data from larger corporations. Specifically, the central server is configured to receive multiple privatized update data from the plurality of local clients and apply differential weighting factors to each client's updates based in part on company size or other corresponding metrics of relevance. By incorporating customizable weighting schemes during the federated aggregation process, the resulting global model can better represent the learning from the distributed clients in a customized manner while still preserving local data privacy.

[0081] In some embodiments, the weighting factor applied to each local client's privatized update data by the central server may be based on various metrics associated with that client. These metrics, which influence the impact of each local client on the aggregated model, can include the amount of data in the local dataset, the variability of data in the local dataset, the industry segment of the local client, or the historical accuracy of update data from that local client. For example, local clients with larger or higher-quality datasets may be assigned a higher weighting by the central server when aggregating privatized updates from multiple clients, thereby giving greater effect tothe learning from clients with more useful data while preserving privacy. The tailored influence aims to produce a more accurate aggregated model reflective of broad insights without directly accessing any raw private data from the individual clients.

[0082] In some embodiments, the central server may add additional noise to the privatized update data received from each local client as part of the aggregation process. This extra layer of noise insertion serves to provide enhanced privacy protection and further obscure any patterns in the update data that could potentially allow inference of the original private data. The noise added by the central server is configured to be cryptographically secure and prevent any leakage of sensitive information through the aggregated model while still retaining overall accuracy. By inserting randomized noise drawn from statistical distributions, the central server can immunize the aggregated model against reconstruction attacks. The additional noise supplements the local differential privacy techniques applied by each client, collectively minimizing exposure risk through a defense-in-depth approach. This multifaceted privacy assurance allows more participants to securely contribute their data toward the aggregated model.

[0083] In some embodiments, the central server may add noise directly to the weights of the aggregated model as an additional privacy-enhancing technique. By inserting randomized perturbations into the parameters of the aggregated model, any patterns that could allow inference of individual clients' private data are obscured. The noise attributes can include using distributions like Gaussian or Laplacian noise with zero mean and calibrated standard deviations to limit the magnitude of the introduced randomness. This direct injection of noise into the aggregated model weights provides supplemental privacy assurances alongside existing protocols like secure multiparty computation and differential privacy used when aggregating the privatized update data from clients. The approach balances preserving privacy alongside retaining sufficient accuracy in the shared model. Introducing noise to the aggregated model weights represents one method that may be optionally used by the central server to further anonymize the contributions of individual clients to the model updates in a randomized fashion, thereby limiting potential reverse engineering while still allowing the model to benefit from collective learning accumulated across clients.

[0084] In some embodiments, the updated aggregated model generated by the central server is shared back with each of the local clients. Specifically, the central server may communicate the most current version of the aggregated model to each of the plurality of local clients across the network. Upon receiving this globally-enhanced aggregated model, each local client then utilizes it to replace its own current local model. In this way, all local clients obtain the benefit and improved accuracy derived from the federated learning process and collective contributions made across clients, while still preserving the privacy of their own local private data. This ensures the local models at each client site evolve concurrently alongside the global aggregated model as it accumulates more anonymized learnings through iterative optimization rounds involving many clients. By continuously synchronizing and replacing local models with the most recently updated aggregated model from the server, the overall estimation capabilities improve in parallel for both localized analysis by individual clients and global analytics by the central server on behalf of multiple clients combined.

[0085] In some embodiments, the method may involve employing secure multi-party computation (SMPC) during model updates to allow for privacy-preserving computation and aggregation of the privatized update data at the central server. SMPC techniques enable collective computation on sensitive data without any node revealing its raw private data. Specifically, each local client may partition its local private dataset and create 'shares' of the data using cryptographic secret sharing schemes. These shares are then transmitted to other clients, where computations occur directly on the encrypted shares rather than the actual data. The shares are mathematically blinded in a manner that no single party can view the raw data belonging to another party. SMPC protocolslike additive secret sharing, Yao's garbled circuits, and homomorphic encryption may be utilized to facilitate nodes performing useful analytics on combined datasets in a distributed environment with an untrusted central server while achieving information-theoretic security for the clients' confidential data.

[0086] In some embodiments, the system facilitates secure multi-party computation (SMPC) to enable clients to collaboratively compute an aggregate function over their collective private data without any client revealing its individual private data points. SMPC employs cryptographic techniques to allow a group of clients to jointly compute an average across their inputs while keeping each client's contributed values hidden from the other participants. This privacy-preserving distributed calculation relies on existing methods within SMPC that obscure the individual data samples by combining them into anonymized intermediate results. The averaged aggregate that emerges reflects the summary trends among the clients' private data while immunizing the participants from exposing or reconstructing the raw inputs. This approach thereby satisfies the system’s dual requirements for utility and confidentiality during cooperative analytics on decentralized private datasets. Through SMPC, clients can contribute to an aggregated learning model without surrendering control over their sensitive data.

[0087] In some embodiments, the method may involve utilizing homomorphic encryption (HE) for the transmission and aggregation of the privatized update data at the central server. Homomorphic encryption enables clients to decrypt the aggregated model without the aggregator being able to access the actual model updates. In particular, homomorphic encryption employs the Paillier algorithm to facilitate the encryption process. This allows for privacy-preserving transmission and aggregation of the privatized update data from the local clients. Through the homomorphic encryption technique, clients can decrypt the aggregated global model distributed back from the central server without exposing their original private data.

[0088] In some embodiments, the central server utilizes homomorphic encryption (HE) for the transmission and aggregation of the privatized update data. Homomorphic encryption enables clients to decrypt the aggregated model without the aggregator accessing the actual model updates. One implementation of homomorphic encryption that may be used is the Paillier algorithm. The Paillier algorithm facilitates the encryption process within the homomorphic encryption framework employed by the system. Using this cryptographic technique allows for secure transmission and aggregation of privatized updates while still enabling the clients to recover meaningful information from the aggregated model. Thus, homomorphic encryption like Paillier provides additional privacy assurances alongside the core federated learning approach.

[0089] In some embodiments, the method involves the central server applying a differential privacy algorithm during the aggregation process. This algorithm has defined parameters including a set gradient norm bound (C) and a noise scale. By applying this differential privacy algorithm with the specified norm bound and noise scale values when aggregating the privatized update data, the central server can provide mathematical privacy guarantees for the local clients' private data. The differential privacy ensures that any single local client's unique data cannot be isolated or identified after aggregation, thereby upholding confidentiality while still allowing the aggregated model to benefit from the collective learning across local clients. In some embodiments, the central server applies a differential privacy algorithm with predefined parameters during the aggregation process to ensure privacy guarantees. Specifically, the differential privacy algorithm uses a gradient norm bound (C) and a noise scale as hyperparameters. To quantify the overall privacy risk, the algorithm further utilizes a privacy accounting technique to compute a cumulative privacy cost encapsulated by the parameters epsilon and delta. This privacy cost reflects the cumulative information leakage over the course of the federated learning process. By capping epsilon and delta at acceptable thresholds per industry standards or benchmarks, the system can provide mathematical assurancesregarding the privacy protections. For instance, an epsilon value may be maintained around 10 to 40 depending on the application. The differential privacy algorithm and the privacy accounting method work in conjunction to facilitate accurate aggregated model updates while still upholding stringent privacy guarantees for the local client data.

[0090] In some embodiments, the method involves transmitting model updates between the local clients and the central server by communicating only the differences between the current local model and a previous version of the aggregated model provided by the central server, rather than transmitting entire model states. By only sharing model differences, the system aims to enhance data security during transmission and further obscure the individual contributions of each client. Specifically, this approach computes model deltas based on the changes that have occurred in the local model parameters since the last time it was synchronized with the global model. These model deltas capture the new learning from the latest round of local client training, contain less information than full models, and help prevent reconstruction of the underlying private training data. Transmitting model deltas rather than complete model copies thus adds an additional layer of security alongside the other privacy-preserving protocols used within the system. The incremental model changes can subsequently be aggregated and applied by the central server to improve the global model. This framework facilitates recurring model updates and convergence to higher accuracy through federated learning while upholding strong data privacy assurances.

[0091] In some embodiments, model updates between the local clients and the central server may be transmitted by communicating only the differences between the current local model and a previous version of the aggregated model provided by the central server. This approach of transmitting just the changes in the model parameters aims to enhance data security during transmission. Specifically, rather than sending entire model files which could potentially reveal more information, only the minimal set of updates denoting how the local model has shifted relative to the earlier aggregated model distributed by the central server are shared. By computing and minimizing the extent of data exchanged, the exposure risk during model update transfers is reduced while still allowing the local learning to be aggregated into the global model. Thus, the central server can efficiently receive localized model updates from each client and update the aggregated model to advance the overall system's accuracy, without needing access to full model snapshots from each client at every iteration. This targeted transmission of model differences enables collaborative learning while strengthening confidentiality protections around what data is shared externally.

[0092] In some embodiments, the central server employs a stochastic gradient descent (SGD) algorithm with differentially private updates to minimize an error in the aggregated model while preserving privacy. The SGD algorithm allows the central server to iteratively optimize the aggregated model by descending along gradients that reflect the collective learning from the local clients' privatized updates. However, instead of directly using the raw gradient updates, the central server first processes them through privacy-enhancing techniques such as clipping and noise addition to generate differentially private gradients. These privatized gradients obscure the contributions of individual clients while still capturing overall optimization trends. By aggregating the learning accumulated across multiple stochastic steps and batches drawn from the entire client population, the differentially private SGD algorithm continually improves the accuracy of the aggregated model for impact estimation while simultaneously ensuring that the privacy of each client's raw private data is preserved. The SGD methodology provides an efficient mechanism for the central server to refine the aggregated model in a privacy-conscious manner on behalf of all participants.

[0093] In some embodiments, the central server employs a stochastic gradient descent (SGD) algorithm with differentially private updates to minimize an error in the aggregated model while preserving privacy. This SGD algorithm involves several key steps. First, it samples mini-batches of training data. For each sampled data batch, it computes the gradients of the loss function with respect to the model parameters. Next, it clips the gradients if they exceed a predefined threshold value. The SGD algorithm then adds noise, such as Gaussian or Laplace noise, to the clipped gradients. Finally, it performs a descent update to modify the model parameters using the noisy gradients. By iterating through multiple rounds of this sampling, gradient computation, clipping, noise addition and descent, the SGD algorithm enables iterative improvement of the aggregated model through privatized updates from local clients while protecting the privacy of any individual client's raw private data.

[0094] In some embodiments, the central server is configured to compute an aggregate function over the private data from the local clients. This is accomplished using secure transmission methods that do not reveal the individual private data points from any given local client. For example, techniques like secure multi-party computation or homomorphic encryption allow the central server to calculate summary statistics or aggregate metrics across the local clients' private datasets without actually accessing or observing any individual data entries. This enables derivation of collective insights and trends while still preserving the privacy of sensitive information from each participant. Through these cryptographic protocols, the central server can determine overall distributions, averages, variability, correlations and other aggregate functions over the union of local private data sets without being able to isolate or access raw data from any single client. Therefore, the system facilitates useful high-level analytics that leverage collective learning across clients while offering rigorous protections around individual data confidentiality.

[0095] In some embodiments, the method may involve adjusting the weighting factor of the aggregated model based on a reliability metric associated with each local client at the central server. More specifically, each local client's respective privatized update data may be assigned a weighting factor when aggregated into the global model. This weighting factor can correspond to metrics indicative of the reliability of each local client's data contributions, such as the historical accuracy of a client's model updates. Local clients providing higher-quality updates may have their updates more heavily weighted during aggregation, thereby granting them greater influence in enhancing the global model. This ability to tune each client's impact on the model update process based on reliability indicators can improve the accuracy and robustness of the final aggregated model. In effect, this weighting approach amounts to a credibility scoring system that helps determine the trustworthiness of each participant's model improvements so as to integrate their contributions judiciously. The central server can compute and apply these dynamic weighting factors during federated learning rounds without accessing the raw data from any client.

[0096] In some embodiments, the method utilizes Secure Multi-Party Computation (SMPC) protocols to prepare local model updates for transmission at both the central server and the local clients. This preserves privacy by obscuring individual updates through cryptographic techniques. Specifically, before transmitting local model updates from a local client to the central server, the SMPC protocols obscure the individual contributions of each client. This is done by adding cryptographic noise or performing other secure transformations on the updates. As a result, the privatized update data transmitted to the central server is a secure, shareable version of the original update data that prevents inference of any local client's private data. The central server can then aggregate this privatized data to improve the global model without revealing any single client's data. SMPC thereby enables a privacy-preserving collaborative approach to enhancing the aggregated model's accuracy.

[0097] In some embodiments, the system utilizes Secure Multi-Party Computation (SMPC) protocols to preserve privacy when preparing local model updates for transmission between the local clients and the central server. As part of these protocols, each local client may perform operations to add cryptographic noise to their local model updates prior to sending them to thecentral server. Adding this type of noise serves to obscure the raw update data, making it difficult to isolate or identify any individual client's contributions. This adds an extra layer of security that prevents the inference or reconstruction of the sensitive private data used in training the local models. By leveraging built-in cryptographic protections, the SMPC protocols allow the local model updates to be transmitted and aggregated in a privacy-preserving manner without revealing any private data from individual local clients.

[0098] In some embodiments, as part of the Secure Multi-Party Computation (SMPC) protocol utilized to preserve privacy, each local client may perform secure transformations on the updates from its local model before transmission. These transformations serve to generate secure, shareable versions of the local model updates. Through techniques such as adding cryptographic noise or employing other encryption schemes, the updates are obscured to prevent inference of the private data while still allowing model improvements based on aggregate trends. This approach enables each local client to transform sensitive model updates into anonymized representations that protect confidential information belonging to that client. The collective learning from these privatized updates of many clients facilitates enhancement of the centralized aggregated model in a privacy -respecting manner. Thus, the SMPC protocol provides computations to convert private local data into a shareable format to achieve both collaboration and privacy goals.

[0099] In some embodiments, the central server distributes the aggregated global model back to the local clients in such a way that each client can only decrypt the specific updates that pertain to their own local model. This selective decryption ability maintains the overall confidentiality of the global aggregated model at the central server level. Specifically, the central server utilizes cryptographic techniques to allow each local client to isolate and access the relevant parameters and weights during model transmission while keeping all other details of the global model private. This ensures that clients can receive the improvements from collective learning that relate to their local dataset and models without exposing the sensitive details of other participants. Thus, the central server enables a level of selective decryption when sharing back the aggregated model so that local clients can benefit from the shared intelligence while critical elements remain encrypted.

[0100] In some embodiments, the privatized update data from a subset of the local clients may be aggregated together prior to transmission to the central server. Specifically, the updates from multiple local clients can be combined into a batch that is then sent to the central server. By aggregating the updates from several local clients into a single transmission, this approach helps obscure the contribution of any one individual local client. This batch transmission of aggregated privatized updates from groups of local clients enhances privacy by preventing the isolation of any singular local client's update data. Thus, the central server receives transmissions corresponding to aggregated batches from subsets of the collective group of local clients rather than individual updates that could be traced back to a particular local client. This batched approach for sending aggregated blocks of updates from subsets of local clients increases anonymity when transmitting data to the central server.

[0101] In some embodiments, the method involves each local client executing a training process on its local model using its local dataset to generate update data corresponding to the private data. Computations are then performed in accordance with the Secure Multi-Party Computation (SMPC) protocol to generate a secure, shareable version of the update data, thereby defining the privatized update data. These computations involve adding cryptographic noise or performing secure transformations to obscure the update data to prevent inference of the private data. The central server then aggregates the privatized update data from the plurality of local clients using SMPC protocols, ensuring the aggregation process does not reveal any local update data. The central server updates the aggregated model to form an updated version, reflecting combined learning fromthe multiple local clients without accessing any private data. Finally, the central server distributes the new version of the aggregated model back to the local clients for updating their respective local models.

[0102] In some embodiments, the system for estimating impact while preserving privacy is implemented in a distributed manner without a central server. In these embodiments, the system comprises a plurality of local clients. Each local client has a database containing private data, a processor to train a local model using the private data and generate privatized update data, and a transmitter to transmit the privatized update data. The plurality of local clients are configured to collectively update an aggregated model in a decentralized fashion without any individual local client accessing another local client's private dataset. This distributed approach allows the local clients to jointly improve the aggregated model through multiple rounds of federated learning without compromising the privacy of their local private data.BRIEF DESCRIPTION OF THE DRAWINGS

[0103] These and other aspects will become more apparent from the following detailed description of the various embodiments of the present disclosure with reference to the drawings wherein:

[0104] Fig. 1 shows a block diagram illustration of a cloud-based system to estimate a lifecycle impact in accordance with an embodiment of the present disclosure;

[0105] Fig. 2 shows a block diagram illustration of federated learning in accordance with an embodiment of the present disclosure;

[0106] Fig. 3 shows a block diagram illustrating gradient descent in accordance with an embodiment of the present disclosure; and

[0107] Fig. 4 shows a flow chart diagram of a method of using federated learning to generate an aggregated model that preserves privacy in accordance with an embodiment of the present disclosure.DETAILED DESCRIPTION

[0108] Fig. 1 shows a block diagram of a system 100 to estimate an impact value 178 (such as life-cycle impact value) while preserving privacy of sensitive data, such as private data 158 of a local client 104a, in accordance with an embodiment of the present disclosure. The system 100 includes a cloud-service provider 102 and multiple local clients 104a-c that communicate over a network 108. A life-cycle estimator 178 implements federated learning to generate, train, and / or update an aggregated model 114 while preserving the privacy of the private data 160 in each of the local clients 104a-c.

[0109] Within the cloud service provider 102, a resource dispatcher 110 is utilized to dynamically allocate computing resources based on demand. This allocation includes the provisioning of a virtual server 122, which is tailored to deliver scalable computing solutions. The virtual server 122 is a composite of several virtual components, including but not limited to a virtual processor 124, virtual memory 126, and virtual disk space 128. These virtual components are configurable and can be tailored to the specific computational needs of the system 100.

[0110] The database 132 housed within the cloud service provider 102 serves as a repository for various data types and structures, ranging from serverless databases to traditional cloud databases. This flexibility allows the database 132 to store and manage a diverse range of data sets, including the conversion factors lookup table 146, training data 148, user accounts 150, and aggregate privatized data 152. The training data 148 can be in addition to the training data foundwithin the private data 160. In some embodiments of the present disclosure, all training data resides in a local client 104a. In other embodiments, the private data 160 is in addition to the training data 148 found on the cloud-service provider 102. The nature of the database 132 facilitates efficient data retrieval and manipulation.

[0111] In addition to the virtualized resources, the cloud service provider 102 incorporates a server farm 119, which may comprise an array of servers, designated as servers #1 to #n 121. These servers 121 can be physical or virtual machines configured to handle the computational load distributed by the resource dispatcher 110. The server farm 119 offers robust computing power and can be scaled to accommodate varying loads, ensuring that the system 100 maintains high performance and availability.

[0112] In some embodiments, the system 100 can alternatively be implemented on a local server setup within a single machine. In this configuration, the local server may house the necessary hardware to perform the system's functionalities, similar to those provided by the cloud service provider 102. The single machine can encompass a dedicated processor, memory, and disk space, which are allocated specifically for the system's operations. The database, equivalent to database 132, can be hosted locally, providing the system 100 with the required data storage and management capabilities on-premises. This local server approach allows for a contained environment where the system 100 can operate independently of external cloud resources, potentially offering enhancements in terms of data security and response times due to the proximity of the hardware resources.

[0113] Thus, the cloud-service provider 102 contains servers 121 and other infrastructure to receive privatized update data 174 from the local clients 104 and to aggregate this privatized update data 174 in a database 132 as aggregated privatized data 152. The aggregated privatized data 152 may be used to train and update an aggregated model 114 via an optimization component 115. The aggregated model 114 may be shared among all of the local clients 104 to allow any local client 104 to use the query component 120 to query the aggregated model 114 to estimate an impact value 178. In some embodiments, the query component 120 is only used to testing and / or training the aggregated model 140 such that it is required that the local clients 104 perform queries only on a respective local model 164. The privatized update data 174 protects the original private data 160 on the local clients 104. Each local client 104 has a database 154 containing a local dataset 156 that includes private data 160 as well as public data 158. The local client 104c trains a local model 164 using the private data 160 to generated privatized update 174 which is sent to the life-cycle estimator 184.

[0114] The GUI component 186 residing within the local client 104a acts as a graphical user interface, facilitating interaction between the system's human users and other components of the local client 104a. It enables users to view, manage, and extract insights from the system in an intuitive visual format tailored to varying access levels and roles.

[0115] The GUI component 186 includes multiple graphical components and visualizations for querying impact estimates and life cycle scores from the local model 164, configuring system parameters, managing user permissions, tracking analytics, and coordinating distributed learning rounds.

[0116] Through dynamic and customizable menus, forms, charts, graphics and widgets powered by up-to-date web frameworks, the GUI component 186 grants authorized users a consolidated portal to handle operational and analytical tasks within the system 100. Real-time visual feedback and explanations justify predictions, enhancing interpretability of model behavior. Search filters, comparison tools, sharing options and interactive what-if scenarios support data- driven decision making.

[0117] The GUI component 186 may present life cycle impact estimates at product, process or company levels depending on the user's scope. Thus, a local user may use the GUI component186 to query only the local model 164, which may be identical to the aggregated model 114 as received during a latest update round. Interactive geospatial depictions of supply chain hotspots enable targeted strategies. Custom reports integrating external ecosystem datasets provide context for sustainability initiatives. Users can visualize product and process optimizations in relation to reduction targets across emissions, waste, water and more.

[0118] Configuration screens give administrators control over federated learning parameters like communication protocols, privacy algorithms used, and access privileges. User account dashboards centralize profile management, API keys, model accuracy metrics, data uploads and impact queries. The interface includes in-line tutorials, tooltips and chatbots to simplify usage for non-technical users. Role-based views ensure important actions require explicit permissions, logged for auditing. Input validation and encryption protect data security. Version histories prevent overwriting of parameter changes.

[0119] While providing analytical flexibility to users, the GUI component 186 maintains consistent and intuitive navigation standards across views. Shortcuts, notifications, and accessibility accommodations assist usage. Integrations with complementary systems enable smooth data import / export and automation. By aligning human-centered design with sustainability objectives, the GUI component 186 delivers a platform for organizations to make measurable progress in minimizing their life cycle impacts.

[0120] As previously mentioned, a local user may use the GUI component 186 to query only the local model 164, which may be identical to the aggregated model 114 as received during a latest update round. The query component 187 is a software module within the local client 104a that interfaces with the local model 164 to query and extract predictions on life cycle impact values. The query component 187 enables the local client 104a to leverage the learning accumulated within its local model 164 to estimate sustainability metrics without having to externally transmit its private data 160.

[0121] In some embodiments, the query component 187 queries an exact copy of the aggregated model 114 that resides within the local client 104a, while in other cases, it queries a version of the aggregated model that has been updated through additional local training based on the private data 160. The query component 187 extracts impact value predictions from the local model 164 in response to user-submitted inputs, and may potentially convert these values into interpretable sustainability scores via the score component 189.

[0122] The query component 187 may facilitate querying the local model 164 while training is in progress. For example, tt can query after each epoch to monitor impact value predictions as model accuracy improves during the training process. In some embodiments, the query component187 obtains a single impact value per query, while in others, it may extract multiple values corresponding to different sustainability indicators through a batch of queries. The specific nature of the query and the categories of impact values returned depend on user requirements and interface selections.

[0123] By enabling low-latency queries directly through the local model 164, the query component 187 allows local clients 104a to make rapid sustainability assessments without needing to externally transmit requests or data through the network 108. This preserves the privacy of the local private data 160 while still allowing impact analysis. The query component 187 is engineered to interface with the local model architecture and feature space to extract relevant and accurate impact valuations. It may thus play a role in maximizing the utility of the federated learning paradigm for decentralized, privacy -preserving life cycle assessments.

[0124] The impact value 188 represents an estimated measure of environmental, social, or economic impact corresponding to a product, process, or service across its life cycle stages. This value is generated by the local client 104a querying its local model 164 using a query component 187.

[0125] In one embodiment, the impact value 188 quantified by the local model 164 reflects product-related impacts across raw material extraction, manufacturing, distribution, usage, and end- of-life stages. For example, the impact value 188 for a smartphone could encapsulate cumulative energy consumption, carbon dioxide emissions, water utilization, and waste generation from material mining, component fabrication, phone assembly, packaging, transportation, consumer charging and disposal.

[0126] The impact value 188 may correspond to a diverse range of impact categories assessed through life cycle assessment, including but not limited to greenhouse gas emissions, energy use, water depletion, toxicity, eutrophication, acidification, ozone depletion, and more. Specific examples of impact values 188 estimated by the local model 164 include total carbon dioxide equivalent emissions measured in kilograms, cumulative non-renewable primary energy demand measured in megajoules, blue water consumption measured in liters, and hazardous waste disposed measured in kilograms.

[0127] In certain embodiments, the impact value 188 may represent an aggregated indicator that combines multiple impact categories into a single metric. For instance, the ReCiPe or TRACI methods can be applied to condense distinct environmental impacts into a singular eco-point score. Economic impacts such as manufacturing costs and end-user expenses can also be incorporated to estimate the overall ownership costs. The impact value 188 may additionally reflect social impacts across the life cycle, including labor wages, job creation, working conditions, and community health.

[0128] The level of specificity of the impact value 188 can vary based on the data resolution and model complexity. Granular bill of material data can enable component-level impact analyses while generic industry benchmarks may yield higher-level product category impact estimates. Assumptions regarding usage profiles, lifetime, and disposal scenarios can also influence the impact value fidelity.

[0129] As the local models 164 trained on local private data 160 go through successive rounds of federated learning orchestrated by the central server 102, their accuracy and sophistication in estimating useful impact values 188 continue to improve. The corresponding impact scores 190 derived from the impact values 188 thus become more reliable indicators of environmental sustainability over time.

[0130] The impact value 188, when used alongside other values by the score component 189, can yield impact scores 190 reflecting a wide spectrum of sustainability metrics while protecting data privacy. The open-ended nature of the impact value 188 allows flexibility in the specific economic, environmental and social dimensions quantified, tailored to the estimation requirements for different applications.

[0131] The score component 189 is a component within the local client 104a. The score component 189 generates a score 190 using the impact value 188 produced by querying the local model 164. The impact value 188 is received from the query component 187, which queries the local model 164. Thus, the score component 189 may receive the impact value 188 from the query component 187.

[0132] In some embodiments, the score component 189 may be a local version of the score component 182 found in the cloud-service provider 102. Thus, the score component 189 in the local client 104a performs a similar function to the score component 182 in the cloud-service provider 102.

[0133] The score component 189 converts the impact value 188 into the score 190. The impact value 188 may be a quantitative value reflecting an environmental impact, such as greenhouse gas emissions, energy use, water consumption, waste generation, etc. The score 190 may be a dimensionless value on a standardized scale that indicates the sustainability performance based on the impact value 188.

[0134] To generate the score 190, the score component 189 may use conversion factors obtained from a database, such as the database 154 in the local client 104a or the database 132 in the cloud-service provider 102. The conversion factors translate the impact value 188 into comparable score 190 based on baselines or benchmarks. This enables the score 190 toContextualize the impact value 188 as a performance rating.

[0135] In some implementations, the score component 189 may generate multiple scores 190 corresponding to different impact categories, such as scores for emissions, water use, and waste, providing a sustainability profile. The score component 189 may also produce aggregate scores reflecting overall environmental performance.

[0136] The score 190 enables local clients 104a to evaluate products, processes, or services for meeting sustainability goals. It facilitates comparison against targets, between alternatives, or over time periods. By quantifying impact data into readable scores 190, the score component 189 assists stakeholders in making decisions and tracking improvements. Its functionality complements the privacy-preserving impact valuation offered by the system 100.

[0137] The score 190 is generated by the score component 189 residing within the local client 104a. The score component 189 functions to convert an impact value 188 into an interpretable score 190 that can be easily understood by users. The impact value 188 is produced by the local model 164 in response to a query submitted via the query component 187.

[0138] In one embodiment, the score 190 is a numerical value representing an environmental sustainability metric such as greenhouse gas emissions, energy consumption, water utilization, waste generation, or economic impacts associated with a product or process. The score component 189 applies a conversion factor or formula to translate the raw impact value 188 into a standardized score 190 along an established scoring system that allows for comparison against benchmarks. For example, the impact value 188 which denotes the carbon dioxide equivalent emissions in kilograms may be converted into a carbon footprint score 190 on a scale of 1 to 100.

[0139] The score component 189 can produce multiple scores 190 corresponding to different dimensions of environmental performance. These scores can include metrics for health impacts, ecosystem quality impacts, natural resource consumption, carbon footprint, water usage footprint, and many others. Weightings may be applied when synthesizing multiple impact values 188 into a composite sustainability score 190.

[0140] In some implementations, the score component 189 retrieves conversation factors, weightings, and scoring logic from a database 154 accessible to the local client 104a. This database 154 may be synchronized with or replicated from the centralized database 132 managed by the cloud service provider 102. Keeping the score component 189 localized allows each client 104a to generate scores 190 rapidly without needing to transmit requests to external servers and also to maintain the privacy of the queries. However, the scoring rules and benchmarks are uniform as they derive from the centralized database 132.

[0141] The score component 189 may be configurable to meet the specific needs of different users and use cases. It can output scores 190 reflecting proprietary scoring systems employed within a private company. For external reporting and benchmarking, it can apply standardized templates like the GRI sustainability standards. The component 189 can also produce detailed reports to accompany the scores 190, providing breakdowns of performance by life cycle stages and recommendations for improvement opportunities.

[0142] In some implementations, the score component 189 may apply more advanced techniques when calculating scores 190. For instance, multi-criteria decision analysis (MCDA) methods can be incorporated to account for conflicting priorities between environmental, social, and economic dimensions. The scoring models can also encompass probabilistic methods to capture uncertainty ranges. And explainable artificial intelligence techniques may be used to provide transparency into the reasoning and data provenance behind a score 190.

[0143] By aggregating many privatized update data 174 from many local clients 104, the lice-cycle estimator 184 constructs a robust aggregated model 114 for estimating impact values 178, such as life-cycle impacts, without accessing the original private data 160 of any of the local clients 104. The system 100 provides impact values 178 while preserving data privacy across the network 108 of local clients. Impact values 178 can be further processed into scores 180 reflecting environmental sustainability metrics. Thus, this federated learning within the system 100 enables collaborative training of the aggregated model 114 without sharing the raw private data 160. The local dataset 156 remains protected on each local client 104 while allowing accurate analytics in the cloud. This balances both data privacy and utility across the distributed network.

[0144] In more detail of one specific embodiment, the system 100 implements federated learning to enable collaborative training of the aggregated model 114 by multiple local clients 104 without sharing their raw private data 160 with the cloud service provider 102 or other entities across the network 108, thereby preserving data privacy. Each local client 104a-c has a local dataset 156 containing private data 160 as well as public data 158. The private data 160 resides only on the local client 104 and is not shared externally. The local dataset 156 is stored on a database 154 on each local client 104.

[0145] Both private data 160 and public data 158 may be used in training and refining the local models 164 and the aggregated model 114. The training process may leverage the unique characteristics of each data type to enhance the accuracy and reliability of the impact estimations while preserving the confidentiality of the private data. The private data 160 may contain sensitive information proprietary to the local client, such as detailed manufacturing processes, specific material compositions, or confidential supply chain information. Public data 158, on the other hand, may include broadly available information such as industry-standard emission factors, generic material properties, or widely recognized benchmarks for environmental impacts. The public data 158 may supplement the training by providing a broader context and filling in gaps where the private data 160 may be sparse or non-specific. For instance, where a local client's private data 160 might lack certain information on upstream processes, public data 158 from public databases may provide generic lifecycle data that can be integrated into learning batches create a more complete local model 164.

[0146] In federated learning, each local client 104 has a local model 164 that is trained on the local private dataset 156 consisting of private data 160, using a processor 162 and a training component 166. The training component 166 executes training algorithms, such as stochastic gradient descent, on the local datasets to update parameters of the local model 164. The local model 164 is initially set to or seeded by the aggregated model 114 from the previous round of federated learning. The training process generates update data 168 corresponding to the changes in parameters after running multiple training steps on the local private dataset 156.

[0147] To train machine learning models without directly accessing the private data 160, each local client 104a-c utilizes its local dataset 156 to train a local model 164 using the training component 166 and the processor 162. The training component 166 performs computations like gradient descent on the local model 164 parameters to minimize a loss function and make predictions on the local data. Multiple iterations of training are performed to update the local model 164 - this generates update data 168 reflecting the new state of the model.

[0148] The local model 164 may be initialized using parameters from a previous version of the aggregated neural network model 114 obtained from the cloud service provider 102. The training component 166 on each client 104 then performs multiple iterations of gradient descent and backpropagation to update the weights and biases across the layers of the local neural network model 164 to minimize a loss function and make accurate predictions on the local private data 160.

[0149] Specifically, the training component 166 samples mini -batches of private data 160 from the local dataset 156. Each data sample is forward propagated through the neural network model 164 by applying matrix multiplications between the weight matrices and input values at each layer and then passing the output through an activation function to generate a prediction. The prediction is compared to the actual label via a loss function like mean-squared error, mean absolute error, hubber loss, etc. The loss gradients for each weight are then computed by backpropagating the errors across the network using the chain rule of derivatives. Gradients indicate how much small changes to each weight impact the overall loss - this determines the update direction during optimization. The training component 166 sums the gradients over the minibatch before updating the weights through gradient descent. By iteratively sampling training data batches to calculate gradients and tune the weights by descent, the local neural network model 164 becomes more accurate on the local private data 160 over multiple epochs of the training process. Between epochs, batch updates are compiled to generate the update data 168 reflecting new parameter values or the suggested deltas needed to change the model’s weights or biases.

[0150] Although transmitting model updates back to the cloud without transmitting any training data already obscures some data, clients can add additional layers to preserve privacy. To preserve privacy, each local client 104 has an additional, optional privatization component 170 that performs computations to transform the update data 168 into privatized update data 174 before transmitting it to the cloud service provider 102. The privatization component 170 utilizes techniques such as differential privacy by clipping gradients and adding random noise from distributions like Gaussian or Laplacian. Cryptographic approaches like homomorphic encryption, secure multi-party computation protocols, and Paillier encryption may be also applied by the cryptographic component 172 within the privatization component 170. These methods obscure the original update data 168 values to prevent inference of the private data 160 while still retaining overall trends and patterns to improve the aggregated model 114 performance. This privatized update data may be configured to be difficult for anyone maliciously attempting to intercept the update to infer what parts of the model were being updated (e.g., a man-in-the-middle attack), because many or all weights could have small perturbations introduced by the privatization module as described herein.

[0151] The privatized update data 174 from each local client 104 is transmitted via the transmitter 176 over the network 108 to the receiver 116 located on the cloud service provider 102. The transmitter 176 can securely transmit data across a network 108. To ensure data integrity and security, transmitter 176 may employ a multitude of protocols and techniques tailored to the requirements of the system. One commonly used data format is the Extensible Markup Language (XML), which facilitates structured data interchange among diverse systems. XML is advantageous for its flexibility and wide support across platforms and programming environments. For safeguarding the data in transit, various encryption methodologies may be used, including symmetric-key encryption for speed and efficiency, and asymmetric (public-private key) cryptography for securely exchanging keys over an untrusted network. Complementing encryption, compression algorithms can be applied to reduce the data size for efficient transmission, which is particularly beneficial when dealing with large datasets or limited bandwidth scenarios. In addition, transmitter 176 might implement secure communication protocols such as Transport Layer Security (TLS) or its predecessor, Secure Sockets Layer (SSL), to create a secure channel. For ensuring thatthe data is not tampered with, cryptographic hash functions and digital signatures may be utilized, providing a means to verify the integrity and authenticity of the data. Additionally, the transmitter 176 may use secure file transfer protocols like SFTP or SCP when the transmission involves files, rather than data streams. Furthermore, when employing public-private cryptography, transmitter 176 can integrate with established Public Key Infrastructure (PKI) systems to manage key distribution and trust verification. These protocols and techniques collectively ensure that transmitter 176 can facilitate secure, efficient, and reliable data transmission within the described system when communicating the privatized update data 174 through the network 108 to the receiver 116.

[0152] The receiver 116 collects the privatized updates from many local clients 104 and passes it to the optimization component 115. The data may be stored in the database 132 as aggregated privatized data 152. The optimization component 115 utilizes all the privatized update data 174, such as stored as aggregate privatized data 152, using a differentially private stochastic gradient descent algorithm to generate an updated aggregated model 114. The aggregation provides model improvements reflecting the broad learning from the collective local private data 160 across clients 104, without the cloud 102 actually accessing the raw private data 160.

[0153] The updated aggregated model 114 does not contain any of the original private data 160 from clients 104, only the cumulative learning. This global model can then be distributed back to all clients 104a-c via the network 108. Each local client 104 incorporates the enhanced aggregated model as its new local model 164. Further local training refines the model, and newly generated privatized updates 174 are sent to the cloud 102 - thus completing another round of federated learning.

[0154] In some embodiments, each local privatized update 174 is weighted differently by the optimization component 115 based on assigned metrics for the respective client 104, such as data quantity, variability, historical accuracy, company size, industry segment, etc. Weighting factors enable tailored impact for each local client 104 in the updated aggregated model 114 while preserving privacy. Additional noise may also be inserted directly into the aggregated model 114 weights or privatized update data 174 before optimizing to further enhance privacy. The updated aggregated model 114 is then transmitted back to all clients 104 to be used as the starting point for the next round of localized training, thereby continuously enhancing model accuracy through successive federated learning cycles.

[0155] In alternate and optional privacy-preserving approaches, the cloud service provider 102 can utilize secure multiparty computation techniques instead of or in addition to differential privacy when aggregating the privatized updates 174 from each local client 104. Secure multiparty computation allows collective computation of an aggregate function like an average without revealing any one client's individual private data points. Encryption protocols are applied to privatize and secure each local update 174 before transmission such that even post-aggregation, no single client's updates can be isolated or identified, thereby preserving privacy while still improving the global aggregated model 114.

[0156] As the federated learning process cycles through successive rounds of localized client training, cloud aggregation of privatized updates, and distribution back to clients of enhanced models, model accuracy continues improving to estimate life cycle impact values 178 without compromising data privacy - thereby achieving both utility and confidentiality goals of the system 100. The life cycle estimator 184 queries the aggregated model to generate anonymized impact scores 180 while preventing any exposure of sensitive private data 160 from individual local clients 104.

[0157] The aggregated model 114 may be trained and queried using a comprehensive array of features that span across various dimensions including products, supply chains, manufacturers,production processes, and sustainability impacts. One or more of these features may be excluded or deidentified for privacy reasons. In the context of products, the model can incorporate an array of features. These may encompass a wide variety of product categories such as food, apparel, electronics, furniture, and chemicals. Specific product identifiers that can be used include names, brands, model numbers, and GTIN codes. Detailed information about product components and bills of materials is also considered, which outlines raw materials, intermediate products, and subassemblies. Further product-related features include attributes like weight, dimensions, performance specifications, and certifications. The intended markets for the products, their expected lifetime, usage patterns, and guidelines regarding usage, maintenance, and disposal are also factored into the model.

[0158] When it comes to supply chain data, the model can process a range of features that include supplier identifiers and attributes like names, locations, industries, and sizes. It can also handle geographic coordinates of supply chain stages, energy mix in those regions, modes of transport between supply chain nodes such as truck, rail, ocean, or air, distances traveled between these nodes, and the types of packaging used for material transport.

[0159] Regarding the input on manufacturers, the aggregated model can utilize comprehensive details such as company names, locations, contact information, industry classifications, and production technologies and techniques. The model can also integrate data about the consumption of energy sources, water usage, waste generation profiles, emissions monitoring, and sustainability practices, including circular production and environment or social performance certifications.

[0160] As for production processes, potential inputs could range from process flow diagrams that clearly delineate unit operations to process simulation results showing mass and energy balances. Life cycle diagrams that map out supply chain stages, unit operation parameters like temperature, pressure, and flow rates, and details on chemicals or reagents consumed are also included. Furthermore, the model takes into account energy demand and greenhouse gas emissions for each unit process, waste creation, treatment, disposal practices, worker hours, and capital equipment used in production steps.

[0161] Additionally, in addressing sustainability impacts, the model is capable of incorporating features such as the life cycle stage of a product or process, categories of impact including emissions, resource usage, and ecological damage, as well as the magnitude of these impacts over a relevant timescale. It can also factor in the geographical location of impacts, associated costs or externalities, impact scoring against baselines or thresholds, and sustainability ratings across various indicators such as greenhouse gas emissions, water utilization, waste generation, and toxicity.

[0162] Additionally, the aggregated model 114 can ingest or utilize a wide range of other product, supply chain, manufacturer, process or sustainability related properties, statistics, attributes, indicators, credentials, benchmarks, regulations, guidelines, totals, sums, counts, volumes, scores, figures, measurements, trends, projections, dimensions, coordinates, visualizations, diagrams, models, simulations, policies, practices, constraints, relationships or alternative archetypes or frameworks thereof that may pertain to or influence environmental or social impacts across product life cycles. All variations, additions, features, descriptors, categorizations, or metadata thereof which may correspond to, emulate, simulate, estimate, infer, associate with, or relate to life cycle impacts and preservation of data privacy can serve as viable input to train or query the aggregated model 114.

[0163] In some embodiments, the process of preparing data to train the aggregated model 114 for estimating life-cycle impacts may include various preprocessing techniques. One such preprocessing technique that may be utilized is Principal Component Analysis (PCA), a statisticalprocedure that utilizes an orthogonal transformation to convert a set of observations of possibly correlated variables into a set of values of linearly uncorrelated variables called principal components. The process begins by calculating the covariance matrix of the data to understand how the variables in the dataset vary from the mean with respect to each other. Following this, eigenvalues and eigenvectors are derived from this covariance matrix, which are indicative of the directions of maximum variance in the data. These eigenvectors form the new feature space, and by projecting the original data along these new axes, PC A achieves dimensionality reduction. The first principal component captures the maximum variance, with each succeeding component having the highest variance possible under the constraint that it is orthogonal to the preceding components. The number of principal components retained is determined based on the amount of total variance they capture, often chosen to include those that add up to a significant proportion of the total variance, such as 95%. This method is particularly advantageous when dealing with datasets where many variables are involved, and the relationships between them are complex and interdependent.

[0164] In another embodiment, another algorithm that may be used is Independent Component Analysis (ICA), which is a computational method for separating a multivariate signal into additive, independent non-Gaussian signals. It is based on the assumption that the observed data variables are linear mixtures of some unknown latent variables, and the mixing system is also unknown. Here, the goal is to recover the latent variables, which are assumed to be non-Gaussian and statistically independent from each other. ICA begins by centering the data, followed by whitening it to ensure that the signals are uncorrelated and their variances equal unity. The next step involves estimating the mixing matrix by maximizing the statistical independence of the estimated components. This is typically achieved through an iterative process, using optimization algorithms that may include gradient descent or Newton's method, and by leveraging non-Gaussianity measures such as negentropy or kurtosis. ICA is particularly useful in fields such as signal processing and complex datasets where the underlying signals need to be identified and segregated.

[0165] In yet another embodiment, Linear Discriminant Analysis (LDA) may be utilized which is a method used to find a linear combination of features that characterizes or separates two or more classes of objects or events. It is a supervised technique, meaning it takes into account the known class labels during the transformation process. The mechanics of LDA involve computing the means of each class, calculating the scatter within and between the classes, and then finding the linear axis that maximizes the separation between the class means while minimizing the variance within each class. This axis, or these axes in the case of multiple discriminants, serve as the new feature space for classification purposes. LDA is particularly adept at preparing datasets for machine learning tasks where the goal is to predict the category to which a new observation belongs.

[0166] In yet another embodiment, a t-Distributed Stochastic Neighbor Embedding (t-SNE) is a non-linear dimensionality reduction technique well suited for embedding high-dimensional data into a space of two or three dimensions, thereby enabling visualization in a scatter plot. t-SNE differs from PCA and LDA in that it is a probabilistic technique that focuses on keeping similar instances close and dissimilar instances apart in the low-dimensional space. It starts by converting the Euclidean distances between points in the high-dimensional space into conditional probabilities that represent similarities. The similarity of datapoint xj to datapoint x_i is the conditional probability, pj|i, that x_i would pick xj as its neighbor if neighbors were picked in proportion to their probability density under a Gaussian centered at x_i. It then aims to minimize the Kullback-Leibler divergence between these conditional probabilities and the corresponding probabilities in the lowdimensional space, which are represented by a t-distribution. This optimization is usually performed using gradient descent. The 't' in t-SNE refers to the t-distribution that is used in the low-dimensional space which allows for a better handling of outliers and the formation of clear clusters, as opposed to the Gaussian distribution used in the high-dimensional space.

[0167] In yet an additional embodiment, autoencoders can be used, which is a specialized type of neural network, used to learn a compressed representation of the input data. An autoencoder consists of two parts: an encoder that maps the input to a hidden representation, and a decoder that reconstructs the input data from the hidden representation. The encoder and decoder are trained together to minimize the reconstruction error, often using backpropagation and an optimization algorithm such as stochastic gradient descent. The hidden layer of the autoencoder acts as a bottleneck in the network, forcing the autoencoder to learn a compact and often more useful representation of the input data. This compressed representation captures the most important aspects of the data necessary for reconstruction or other downstream tasks, and can be used for dimensionality reduction, denoising, or feature learning purposes. Autoencoders can be designed with various architectures, including convolutional layers for image data or recurrent layers for sequence data, and can include regularization terms such as sparsity constraints or denoising criteria to improve the robustness and generalizability of the learned features.

[0168] These preprocessing techniques may optionally be used to facilitate the management of complexity of the environmental impact data and / or for ensuring that the aggregated model 114 can be trained effectively without compromising the privacy of the data contributors. The techniques may be combined for use in transforming raw data into a form that is both insightful for analysis and respectful of confidentiality, ensuring that the integrity of the data is maintained while extracting valuable patterns and trends.

[0169] Within this system, the GUI component 118 may be used for enabling user interaction with the cloud-based system, specifically designed to support users as stored in the user accounts 150 within the database 132. The GUI component 118 is a user interface that serves as a secure portal through which users can query the aggregated model 114. This interface is engineered to provide an intuitive and user-friendly environment for users to input queries and receive results from the aggregated model. The GUI component 118 is constructed with multiple layers of security to ensure that all interactions with the aggregated model are conducted in a manner that upholds the stringent privacy requirements of the system.

[0170] Through this interface, users can perform a variety of tasks, including but not limited to, querying the life-cycle impact values estimated by the aggregated model, visualizing the impact scores, and downloading reports. The GUI component 118 is designed to provide real-time feedback and visualizations that help users understand the results of their queries, thereby assisting them in making data-driven decisions.

[0171] The GUI component 118 can also serve as a coordination or configuration webpage for the local clients 104. This functionality allows system administrators or authorized users to configure the operational parameters of the local clients, set up security protocols, and manage the federated learning process. Users can access this webpage to download encryption keys, manage their accounts, and perform other administrative tasks that are critical to the secure and efficient operation of the system.

[0172] In some embodiments, the GUI component 118 may complement an Application Programming Interface (API) that allows for programmatic access to the aggregated model 114. While the GUI component 118 provides a visual and interactive experience for users, the API enables automated systems, third-party applications, and scripts to communicate with the model, thereby facilitating integration with other systems and enabling batch processing or automated querying.

[0173] The GUI component 118 may be tailored to accommodate the varying levels of expertise and access privileges of different users. It includes features such as role-based access control, which ensures that users only have access to the functions and data that are pertinent to their role within the organization. For instance, a system administrator may have access to configurationsettings that are not available to a standard user, who may only have the capability to query the model and view results.

[0174] Security settings within the GUI component 118 may leveraging the latest encryption standards and authentication protocols to protect sensitive data. The interface may require multifactor authentication for access, ensuring that only authorized personnel can perform certain actions within the system.

[0175] In one embodiment, the system 100 can be implemented in a distributed manner across multiple entities using secure multi-party computation (SMPC) protocols to preserve data privacy. In this approach, the local clients 104a act may act as distinct nodes in the network 108 that perform decentralized, collaborative training of machine learning models under cryptography -based privacy preservation schemes. The nodes interact with each other and optionally using the central server 102 through SMPC protocols that allow collective computation on sensitive data without any node revealing its raw private data.

[0176] Specifically, each local client 104a may partition its local private dataset 160 and creates 'shares' of the data using cryptographic secret sharing schemes. These shares are then transmitted to other clients, where computations occur directly on the encrypted shares rather than the actual data. The shares are mathematically blinded in a manner that no single party can view the raw data belonging to another party.

[0177] SMPC techniques like additive secret sharing, Yao's garbled circuits, and homomorphic encryption may be utilized alongside protocols like secure aggregation, anonymous messaging, and oblivious transfer. They facilitate nodes performing useful analytics on combined datasets in a distributed environment with an untrusted central server while achieving information- theoretic security for the clients' confidential data.

[0178] In some embodiments, the central server 102 may act primarily as a coordinator node that facilitates the initialization of the SMPC protocols and guides the execution flow but cannot access any client's data during computations. The central server 102 may also constructs the global aggregated model 114 for estimation based on encrypted model updates from clients. It oversees iterating rounds of distributed training but relies entirely on cryptographic mechanisms rather than trust assumptions to maintain clients' privacy.

[0179] The SMPC protocols obscure the individual contributions of each client during training, thereby converting their private updates into anonymized shares that immunize the sensitive information from reconstruction by any other party. The collective learning accumulated across training rounds continuously enhances the accuracy of the aggregated model 114 while simultaneously guaranteeing privacy to all data owners. Fig. 2 shows a block diagram 200 illustration of federated learning in accordance with an embodiment of the present disclosure. In some embodiments, the process 200 depicted in Fig. 2 provides a comprehensive overview of the federated learning system for estimating life-cycle impacts while preserving the privacy of data contributors, as introduced in Fig. 1 and elaborated upon in the system 100 of Fig. 1. The process 200 encompasses a plurality of components which interact with each other to facilitate the collaborative training of machine learning models while ensuring that the private data of the individual local clients remains confidential.

[0180] The result of the process 200 is the final model 202 (weighted avg.), which is synonymous with the aggregated model 114 of Fig. 1. The final model 202 represents the culmination of a federated learning process, wherein it is configured to estimate an impact value of various forms, such as environmental sustainability metrics, without directly accessing private data from local clients. This aggregated model is an embodiment of the collective intelligence derived from individual local models while maintaining each participant's data privacy.

[0181] The local client #1 210, local client #2 214, and local client #3 218, each include a respective dataset, labeled as data set #1 212, data set #2 216, and data set #3 220. These datasets contain private data that is utilized to train individual models specific to each client, referred to as model #1 204, model #2 206, and model #3 208. The relationship between each local client and its corresponding data set is depicted by the containment of the dataset within its respective local client.

[0182] The training process is indicated by arrows labeled with the word 'train', proceeding from each local client to their corresponding models. During this training process, each local client 210, 214, 216 utilizes its processor to train a local model using its local dataset 212, 216, 220, subsequently generating privatized update data that reflects the learning achieved without exposing the raw data itself. In one embodiment of this training is the addition of '+10% noise' above each of the local models, which indicates the application of techniques such as differential privacy to ensure that the updates derived from the private data do not allow for the reconstruction or inference of the original data.

[0183] The arrows extending from model #1 204, model #2 206, and model #3 208 to the final model 202 signify the aggregation process where the privatized updates from each local model are combined to improve the final model. This aggregation is performed by a central server or cloudservice provider, which updates the final model by integrating the contributions from each local model. The final model 202 is then able to provide an impact value that is reflective of the collective learning from all participating local clients.

[0184] Test data 208 is depicted in Fig. 2 and is utilized to assess the performance of each local model as well as the final model. The arrows from the test data 208 to the individual models and the final model indicate the evaluation or testing phase where the predictive capabilities of the models are measured against unseen data.

[0185] In some embodiments, each local model can be trained using various machine learning algorithms and techniques. For instance, the training component within each local client may implement algorithms such as stochastic gradient descent, support vector machines, decision trees, or neural networks. The specific algorithm chosen can depend on the nature of the impact being estimated and the characteristics of the local dataset.

[0186] Furthermore, the privatization of update data, as indicated by the '+10% noise', may involve the use of noise addition techniques like Gaussian or Laplacian noise, as well as other advanced privacy-preserving methods such as homomorphic encryption or secure multi-party computation protocols. These methods serve to obfuscate the updates to prevent any reverse engineering that could compromise the private data.

[0187] Moreover, in some alternative embodiments, the weighting of the individual models' contributions to the final model 202 can be varied based on factors such as the reliability of the data, the volume of the data contributed, or the historical accuracy of each local client's predictions. This means that not all model updates may be considered equally when updating the final model, allowing for a more nuanced and potentially more accurate aggregated model.

[0188] Additionally, as described herein, the training process for each local client may involve various preprocessing techniques to prepare the data. These can include normalization, outlier removal, feature extraction, and dimensionality reduction methods like Principal Component Analysis (PCA) ort-Distributed Stochastic Neighbor Embedding (t-SNE). Such preprocessing steps serve to enhance the training efficiency and the predictive performance of the local models.

[0189] In some embodiments of the system 200, variations in how the test data 208 is applied to the models may be used. For example, cross-validation techniques may be employed where the test data is partitioned into different subsets, with each subset used to evaluate the model and the remaining data used for training. This can provide a more robust assessment of the model's generalizability.

[0190] The central server, incorporating the final model 202, may also employ various optimization techniques to refine the aggregated model. These may include, but are not limited to, gradient descent with momentum, adaptive learning rate algorithms like AdaGrad or Adam, or even evolutionary algorithms that simulate processes of natural selection.

[0191] Fig. 3 shows a block diagram 300 illustrating gradient descent in accordance with an embodiment of the present disclosure. The process begins with the "Training" phase, where the model is exposed to data and commences the learning process. This stage is foundational, as it allows the model to identify patterns and relationships within the data. An arrow labeled "compute" extends from this training phase to the next step, indicating the action where computation of the model's updates occurs.

[0192] Proceeding from the training phase, the "Model Update (g)" 302 step is encountered. In this step, the updates to the model's weights are calculated by determining the difference between the current weights of the model, denoted as (w_{local}), and the weights at a baseline state, represented as (w_{baseline}). An arrow, annotated with "clipping," connects this step to the following one, illustrating the sequence of the procedure.

[0193] The subsequent step is the "Clipping" 304 phase, where updates exceeding a predefined norm, denoted as (C), are scaled down to meet this upper limit. The clipped update maintains the magnitude of the update vector at or below the set threshold. This action is conveyed by an arrow from the model update block to the clipping block, visually guiding one through the progression.

[0194] Following the clipping process, an arrow with the label "Adding Noise" points towards the next step. Here, noise is introduced to the clipped update to further secure the privacy of the data. The noise, characterized by a normal distribution with a mean of zero and a standard deviation influenced by the clipping threshold (C) and the noise scale (sigma), as expressed by (N(0, Csigma)), is added to the update. This step results in a "Privatized Update" 306, which represents the update that has been adjusted for privacy through the addition of noise. Item 308 illustrates the gradient descent that may be used. Here is pseudocode presented for the gradient descent shown as follows in a Table 1 :_ _

[0196] Referring to Table 1 : The algorithm delineates a differentially private stochastic gradient descent method, encapsulated as Algorithm 1. It initiates with an array of examples{xl,...,xN} and a loss function L(9), computed as the mean loss across all examples. The process is governed by parameters: a learning rate q_t, a noise scale c, a group size L, and a gradient norm bound C. The commencement involves a stochastic initialization of the parameter vector 9 0.

[0197] In each iteration of the predefined T iterations, a random subset L t is chosen with a probability correlating to the group size over the total number of examples. The gradient of the loss function with respect to the parameter vector at iteration t, V_9_t L(9_t,x_i), is calculated for each example i in L t. To preserve data privacy, the gradient g_t(x_i) is clipped by its norm, divided by the maximum of 1 and the norm-to-gradient norm bound ratio C to form a g_t.

[0198] Noise is then introduced to the gradient by adding a Gaussian random variable with zero mean and variance dependent on the square of the noise scale and the gradient norm bound. This step yields a noised gradient g_t, inversely scaled by the group size. The parameter vector is updated by descending along the noised gradient, scaled by the learning rate.

[0199] The output is the parameter vector after T iterations, the output 9_T and the overall privacy cost (a, 5) are computed using a privacy accounting method. This cost encapsulates the utility-privacy trade-off, ensuring adherence to differential privacy standards.

[0200] The algorithm features two privacy-related hyperparameters: the gradient norm bound and the noise scaler. The privacy cost (a, 5), indicative of the information leak during the process, is contingent on the values of c. A larger C implies more noise, potentially diminishing model accuracy. The privacy guarantee (a, 5) for one update is mathematically expressed as a = 2 log(1.25 / 5) / o, with specific a values for varying 5. In practical applications, the magnitude of (a, 5) is influenced by industry benchmarks, typically utilizing an a value around 10, extending up to 40 for specific scenarios.

[0201] Fig. 4 shows a flow chart diagram illustrating an embodiment of method 400 for estimating impact while preserving privacy according to certain aspects of the present disclosure. The method 400 exemplified in Fig. 4 outlines a process that integrates with the systems and apparatuses detailed in Figures 1, 2, and 3, and as described in the prior sections of the present disclosure. The method 400 provides a structured approach to achieving the dual objectives of accurate impact estimation and privacy preservation across a network of local clients and a central server.

[0202] The process begins at block 402, signifying the initiation of the method 400. This initial step may involve preparatory actions such as initializing the system components, setting up communication channels between the local clients and the central server, and ensuring that the system is in a ready state to receive and process data.

[0203] At block 404, a central server distributes an aggregate model to a plurality of local clients. This step can involve the mechanisms and technologies described in relation to Figure l's cloud service provider 102, where the cloud service provider 102 communicates data over a network 108.

[0204] At block 405, each local client replaces its respective local model with the distributed aggregate model. This aggregate model may be a seed model or may be a model that was updated during a recent round of tuning and / or training the aggregate model. The local model can either incorporate the new model in its entirety, or by weighting it and blending it with its own model.

[0205] Block 406 outlines the process where each local client trains a local model using its local dataset containing private data. This may correspond with the descriptions provided in Fig. 1, where each local client 104a-c has a local dataset 156 containing private data 160 and public data 158. The training is performed using processors 162 and training components 166 as described in the earlier figures. The training may incorporate various algorithms and techniques, such as those mentioned in the preceding text of the present disclosure, including but not limited to stochastic gradient descent, batch gradient descent, genetic optimization, etc.

[0206] In block 408, each local client generates privatized update data from the trained local model. This may be utilized to preserve privacy. It can utilize the privatization components 170 as shown in Figure 1, which may perform computations to transform the update data into privatized update data. The specific techniques for privatizing the update data, such as the addition of Gaussian or Laplacian noise or the use of homomorphic encryption, are further elucidated in Fig. 3's illustration of the gradient descent and privatization process.

[0207] Block 410 illustrates the transmission of the privatized update data from each local client to the central server. The clients make use of transmitters, like transmitter 176 in Figure 1, to send the data securely over the network 108 to the central server's receiver 116. This step is used in the federated learning process as it allows the central server to aggregate multiple updates without direct access to the sensitive private data.

[0208] At block 412, the central server aggregates the received privatized update data from all local clients. This aggregation process can leverage an optimization component 115, as described in reference to Fig. 1. The optimization component 115 may use various algorithms, potentially including differentially private stochastic gradient descent, to integrate the contributions from multiple local clients into the aggregated model 114.

[0209] Block 414 proceeds with the central server updating the aggregated model using the aggregated data, without accessing the private data. This may ensure that the privacy of the local clients' data is preserved while the aggregated model benefits from the collective learning represented by the privatized update data.

[0210] At block 415, the central server updates the updated aggregate model to the plurality of local clients which the local client may use to replace and / or update its local model again. In some embodiments, a local client may decide to cache the updated aggregate model for testing, validation, authentication, and / or to run diagnostics on the newly received model prior to using it to replace, supplement, or reject the updated aggregate model.

[0211] In block 416, the method 400 queries one of the models to receive an impact value. The query may occur at a local client where the local model is queried and there is no need to communicate the query to the central server. In some embodiments, the query may be where the central server performs a query either for internal testing and / or to provide an outside client the results of the query. Here, the query component 120, mentioned in Fig. 1 or the local query component 187, may play a role in interfacing with the aggregated model or local model, respectively to extract relevant impact estimations, which could include various environmental or social sustainability metrics.

[0212] Block 418 describes generating an impact score based on the queried impact value. The server and / or local client may use a scoring component, similar to the one described in the earlier sections, to convert impact values into interpretable scores that reflect environmental sustainability metrics or other relevant indices.

[0213] Subsequently, at block 420, the impact score is communicated to the user making the query which may mean transmitting back to the requesting local client(s) from a central server or via displaying the results to a user on a GUI. This step closes the feedback loop, providing a user with actionable insights derived from the collaborative learning process. If transmission back to the local clients is utilized, it can be facilitated through communication channels established between the central server and the local clients.

[0214] The method concludes at block 422, indicating the end of the process. However, it should be noted that in practice, the method 400 may be iterative, with the central server periodically distributing updated aggregated models back to local clients for further refinement and training, as suggested by the embodiments within the present disclosure. Additionally or alternatively, the central server may facilitate a distributed learning approach. In yet additional embodiments, themethod may avoid using a central server at all and may utilize a distributed, cooperative approach of model training and / or generation.

[0215] The embodiment described above in Fig. 4 may undergo various modifications to adapt to different scenarios and requirements. For example, the central server may incorporate machine learning models other than those explicitly mentioned, such as various deep learning neural networks, ensemble learning models, or Bayesian networks, to estimate the impact values with higher accuracy or specificity.

[0216] Additionally, the privatization techniques used by the local clients to generate privatized update data can include advanced cryptographic methods such as secure multi-party computation or zero-knowledge proofs, providing additional layers of security and privacy. The central server may also use these or other cryptographic techniques when aggregating the privatized update data to further enhance privacy.

[0217] Furthermore, the central server may apply different weighting factors to the privatized update data during the aggregation process as described herein. These weighting factors can be based on various metrics, such as the quality or quantity of the data provided by each local client, the historical accuracy of the local models, or the relevance of the data to the impact being estimated.

[0218] The process of querying the aggregated model and generating impact scores may also involve additional computational steps or algorithms. For instance, the central server might apply multi -criteria decision analysis (MCDA) techniques to synthesize impact values into a composite score that accounts for various dimensions of sustainability.

[0219] Moreover, the transmission of impact scores back to local clients can be customized based on the preferences or requirements of each client. For instance, clients may receive scores in different formats or through different mediums, such as through a web portal, as an API response, or embedded within an interactive dashboard.

[0220] In some embodiments, the central server may further analyze the impact scores to identify patterns, trends, or anomalies. This analysis could involve statistical methods, data visualization techniques, or other data analytics processes that provide additional insights into the impact data.

[0221] The iterative nature of the method 400 may be tailored to the specific learning rates, convergence criteria, or privacy thresholds set by the participating local clients or dictated by the central server's policies. The method 400 can also accommodate various operational modes, such as batch processing, real-time updates, or periodic synchronization, depending on the network's configuration and the participants' needs.

[0222] The method 400 can be modified to include additional steps, to employ alternative technologies, or to be executed in different sequences, as dictated by practical considerations and the specific objectives of the system in which it is implemented.

[0223] Various alternatives and modifications can be devised by those skilled in the art without departing from the disclosure. Accordingly, the present disclosure is intended to embrace all such alternatives, modifications and variances. Additionally, while several embodiments of the present disclosure have been shown in the drawings and / or discussed herein, it is not intended that the disclosure be limited thereto, as it is intended that the disclosure be asbroad in scope as the art will allow and that the specification be read likewise. Therefore, the above description should not be construed as limiting, but merely as exemplifications of particular embodiments. And, those skilled in the art will envision other modifications within the scope and spirit of the claims appended hereto. Other elements, steps, methods and techniques that are insubstantially different from those described above and / or in the appended claims are also intended to be within the scope of the disclosure.

[0224] The embodiments shown in the drawings are presented only to demonstrate certain examples of the disclosure. And, the drawings described are only illustrative and are non-limiting. In the drawings, for illustrative purposes, the size of some of the elements may be exaggerated and not drawn to a particular scale. Additionally, elements shown within the drawings that have the same numbers may be identical elements or may be similar elements, depending on the context.

[0225] Where the term "comprising" is used in the present description and claims, it does not exclude other elements or steps. Where an indefinite or definite article is used when referring to a singular noun, e.g., "a," "an," or "the,” this includes a plural of that noun unless something otherwise is specifically stated. Hence, the term "comprising" should not be interpreted as being restricted to the items listed thereafter; it does not exclude other elements or steps, and so the scope of the expression "a device comprising items A and B" should not be limited to devices consisting only of components A and B. This expression signifies that, with respect to the present disclosure, the only relevant components of the device are A and B.

[0226] Furthermore, the terms "first," "second," "third," and the like, whether used in the description or in the claims, are provided for distinguishing between similar elements and not necessarily for describing a sequential or chronological order. It is to be understood that the terms so used are interchangeable under appropriate circumstances (unless clearly disclosed otherwise) and that the embodiments of the disclosure described herein are capable of operation in other sequences and / or arrangements than are described or illustrated herein.

[0227] Additional embodiments or features of a system may optionally include any one or any combination of features as described, for example, but not limited to:• A plurality of local clients, each local client comprising: a database having a local dataset containing private data; a processor configured to train a local model using the local dataset and generate privatized update data; and a transmitter configured to transmit the privatized update data from the local model; and a central server comprising: a receiver configured to receive the privatized update data from each of the plurality of local clients; an aggregated model configured to estimate an impact value; and a processor configured to update the aggregated model without accessing the private data in the local dataset of any of the plurality of local clients.• Each local client further comprises a local query component configured to query the local model.• The local model is an identical copy of the aggregated model.• The local model is a copy of the aggregated model that has been updated by the local dataset.• The central server comprises a query component configured to query the aggregated model to receive the impact value while thereby preserving privacy of the private data of all of the plurality of local clients.• The query performed by the query component within the central server is configured to perform the query during the updating of the aggregated model.• The query performed by the query component within the central server performs the queries for at least one epoch of training.• The performed by the query component within the central server corresponds to a single impact value.• The central server further comprises a score component configured to generate a life-cycle score using a plurality of queries to the query component configured to receive a plurality of impact values including the impact value.• The privatized update data consists of only weight update values.• The processor of the central server is further configured to update the aggregated model by aggregating the privatized update data from the plurality of local clients without accessing the private data in the local dataset of any of the plurality of clients.• The impact value is one of a greenhouse gas emission value, an energy consumption value, a water use value, a waste generation value, and an economic impact value.• The private data corresponds to data of a product.• The impact value is a product impact value.• The product impact value corresponds to a production activity of the product.• The production activity is one of a material transport, a material extraction, a material handling, a material use, a material disposal.• The impact value is an impact score.• The impact value is one of a climate change value, a greenhouse gas emission value, a carbon-dioxide value, a carbon-di oxide equivalent value, an ozone depletion value, a CFC- 11 equivalent value, a toxicity value, a human toxicity value, a particulate matter value, an animal toxicity value, an aquatic toxicity value, a particulate matter formation value, an ionizing radiation value, a photochemical ozone formation value, an acidification value, a water acidification value, an ocean acidification value, a SO2 value, an NOx value, a NH3 value, an eutrophication value, a terrestrial eutrophication value, a freshwater eutrophication value, a marine eutrophication value, a nitrogen eutrophication value, a phosphorus eutrophication value, an ecotoxicity value, a water use value, a resource depletion value, a mineral depletion value, a fossil depletion value, a land-use value, a habitat transformation value, a habitat loss value, a biodiversity loss value, a renewable-energy consumption value, a non-renewable-energy consumption value, an energy consumption value, a waste generation value, an economic impact value, a carbon footprint value, a water footprint value, a nitrification value, a rare earth element depletion value, a volatile organic compound emissions value, a heavy metal emissions value, a pesticide residues value, a herbicide residues value, a radioactive contamination value, a biocidal product residues value, a dioxin emission value, a furan emission value, a polychlorinated biphenyls (pcb) emissions value, a noise pollution levels value, an alight pollution value, a thermal pollution from heat emissions value, a microplastic emissions value, a nanomaterial releases value, an antibiotic residues value, an endocrine disruptor emissions value, an algal bloom potential value, a soil contamination levels value, a soil salinization value, a land degradation value, a desertification risk value, a cultural heritage impact value, a visual impact value, an electromagnetic radiation value, a biodiversity impact on endangered species value, an invasive species propagation value, a life cycle sustainability index value, a carbon sequestration impact value, a greenhouse gas intensity value, a water quality index value, a sediment quality value, a turbidity levels value, a disposal impact value, a recycling rate value, a material recovery efficiency value, a supply chain vulnerability value, a labor practices impact value, a community health impact value, an economic dependency value, a regulatory compliance level value, a green certification status value, an operational energy efficiency value, an embodied energy value, an embodied carbon value, a total ownership cost value, a market dependency value, a technological obsolescence rate value, an adaptability to climate change value, a resilience to natural disasters value, and a crisis response capability value.• The central server further comprises a score component configured to generate at least one of a health score, an ecosystem score, and a natural resources score based upon the impact value.• The central server further comprises a score component configured to convert the impact value to a score using a conversion factor obtained from a database.• Each local client of the plurality of local clients further comprises a training component configured to train the local model using the local dataset to generate update data.• The training component fine tunes the local model to thereby train the local model using the local dataset to generate the update data.• Each local client of the plurality of local clients further comprises a privatization component configured to privative the update data to generate the privatized updated data by added noise to the update data.• Each local client of the plurality of local clients further comprises a privatization component configured to privatize the update data to generate the privatized updated data by clipping the update data.• Each local client of the plurality of local clients further comprises a privatization component configured to clip the update data and add noise to the clipped update data to generate the privatized update data.• The added noise is one of Gaussian noise and Laplacian noise.• The privatized update data transmitted from each local client corresponds to only differences between the local model and the aggregated model.• The central server further comprises an optimization component configured to minimize an error in the aggregated model by adjusting a plurality of model parameters of the aggregated model.• The central server is configured to receive a plurality of privatized update data including the privatized update data from the plurality of local clients, wherein the central server further comprises an optimization component configured to apply a weighting factor to each of the plurality of privatized update data received from the plurality of local clients based on a metric of the respective local client of the plurality of local clients.• The metric corresponds to a size of a company.• The metric corresponds to one of: amount of data in the local dataset, variability of data in the local dataset, industry segment of the local client, historical accuracy of update data from the local client.• The central server is configured to add noise to the privatized update data.• The central server is configured to add noise to weights of the aggregated model.• The server is configured to communicate the aggregated model to each of the local clients, wherein each of the plurality of local clients uses the communicated aggregated model as the local model.• The central server is configured to employ secure multi-party computation (SMPC) during model updates to allow for privacy -preserving computation and aggregation of the privatized update data.• The secure multi-party computation enables clients to compute an average of private inputs without revealing individual data points.• The central server utilizes homomorphic encryption (HE) for the transmission and aggregation of the privatized update data, enabling clients to decrypt the aggregated model without the aggregator accessing the actual model updates.• The homomorphic encryption employs the Paillier algorithm to facilitate the encryption process.• The central server applies a differential privacy algorithm with a set gradient norm bound (C) and noise scale (c) during the aggregation process to ensure privacy guarantees.• The differential privacy algorithm utilizes a privacy accounting method to compute an overall privacy cost using parameters a and 5.• The model updates between the local clients and the central server are transmitted by communicating only the differences between models, thereby enhancing data security during transmission.• The differences between models are computed based on the current local model and a previously aggregated model provided by the central server.• The central server employs a stochastic gradient descent (SGD) algorithm with differentially private updates to minimize an error in the aggregated model while preserving privacy.• The stochastic gradient descent algorithm includes steps for sampling, computing gradients, clipping gradients, adding noise, and descent updates.• The central server is configured to compute an aggregate function over the private data from the local clients using secure transmission methods without revealing individual private data.• The central server is configured to adjust the weighting factor of the aggregated model based on a reliability metric associated with each local client.• The central server and the local clients utilize Secure Multi-Party Computation (SMPC) protocols to prepare local model updates for transmission thereby preserving privacy by obscuring individual updates.• The SMPC protocols include operations to add cryptographic noise to the local model updates before transmission.• Each local client performs secure transformations on the local model updates as part of the SMPC protocol to thereby generate secure, shareable versions of their updates.• The server distributes the aggregated global model to the clients in a manner that each client can only decrypt the updates relevant to their local model, maintaining the confidentiality of the global model.• The privatized update data from a subset of local clients is aggregated and transmitted to the central server as a batch to thereby obscure the contribution of individual local clients.• Each local client's processor is further configured to: execute a training process on the local model using the local dataset to generate update data corresponding to the private data; and perform computations in accordance with the SMPC protocol to generate a secure, shareable version of the update data thereby defining the privatized update data, the computations including adding cryptographic noise or performing secure transformations to thereby obscure the update data to prevent inference of the private data; and wherein the central server's processor is further configured to: aggregate the privatized update data from the plurality of local clients using SMPC protocols thereby ensuring the aggregation process does not reveal any local update data from any of the plurality of local clients; update the aggregated model to form an updated version of the aggregate model thereby reflecting combined learning from the plurality of local clients without accessing the private data of any of the plurality of local clients; and distribute the new version of the aggregate model to the plurality of local clients for updating each respective local models.• Each local client further comprises a data preprocessing component configured to transform the local dataset prior to training the local model.• The data preprocessing component is configured to perform at least one of data normalization, feature extraction, dimensionality reduction, and handling of missing values.• The dimensionality reduction is performed by at least one of principal component analysis, independent component analysis, and t-distributed stochastic neighbor embedding.• The central server further comprises a validation component configured to assess accuracy of the aggregated model using a held-out test dataset.• The validation component is further configured to quantify privacy risk of the aggregated model.• The central server further comprises a monitoring component configured to evaluate fairness and bias metrics for the aggregated model.• The query component is further configured to provide an explanation of the reasoning behind an impact value prediction.• The central server further comprises a user interface configured to visually depict life cycle stages corresponding to an impact value.• The privatized update data comprises differentially private stochastic gradient updates.• The transmitter and receiver employ federated learning protocols for decentralized training.

[0228] Additional embodiments or features, which may include any one or any combination of the following, but not limited to:• A method for estimating impact while preserving privacy, the method comprising: training a local model using a local dataset containing private data at each of a plurality of local clients; generating privatized update data from the local model at each of the plurality of local clients; transmitting the privatized update data from each of the plurality of local clients to a central server; receiving the privatized update data from each of the plurality of local clients at the central server; and updating an aggregated model at the central server by aggregating the privatized update data from the plurality of local clients without accessing the private data in the local dataset of any of the plurality of local clients.• Querying the aggregated model at the central server to receive an impact value while thereby preserving privacy of the private data of all of the plurality of local clients.• Distributing the updated aggregated model to the plurality of local clients.• Replacing the local model with the updated aggregated model at each of the plurality of local clients.• Querying the aggregated model at the central server for updating the aggregated model.• Querying the local model prior to training the local model.• The local model and the aggregated model are identical.• The impact value is one of a greenhouse gas emission value, an energy consumption value, a water use value, a waste generation value, and an economic impact value.• The private data corresponds to data of a product.• The impact value is a product impact value.• The product impact value corresponds to a production activity of the product.• The production activity is one of a material transport, a material extraction, a material handling, a material use, a material disposal.• The impact value is an impact score.• generating a life-cycle score using a plurality of queries to the query component configured to receive a plurality of impact values including the impact value at the central server.• Each query of the plurality of queries corresponds to a single impact value of the plurality of impact values.• Generating at least one of a health score, an ecosystem score, and a natural resources score based upon the impact value at the central server.• Converting the impact value to a score using a conversion factor obtained from a database at the central server.• Training the local model using the local dataset to generate update data at each local client of the plurality of local clients.• Fine-tuning the local model to thereby train the local model using the local dataset to generate the update data at each local client of the plurality of local clients.• Privatizing the update data to generate the privatized updated data by adding noise to the update data at each local client of the plurality of local clients.• Privatizing the update data to generate the privatized updated data by clipping the update data at each local client of the plurality of local clients.• Clipping the update data and adding noise to the clipped update data to generate the privatized update data at each local client of the plurality of local clients.• The added noise is one of Gaussian noise and Laplacian noise.• The privatized update data transmitted from each local client corresponds to only differences between the local model and the aggregated model.• Minimizing an error in the aggregated model by adjusting a plurality of model parameters of the aggregated model at the central server.• receiving a plurality of privatized update data including the privatized update data from the plurality of local clients, and applying a weighting factor to each of the plurality of privatized update data received from the plurality of local clients based on a metric of the respective local client at the central server.• The metric corresponds to a size of a company.• The metric corresponds to one of: amount of data in the local dataset, variability of data in the local dataset, industry segment of the local client, historical accuracy of update data from the local client.• Adding noise to the privatized update data at the central server.• Adding noise to weights of the aggregated model at the central server.• The aggregated model to each of the local clients, wherein each of the plurality of local clients uses the communicated aggregated model as the local model.• Employing secure multi-party computation (SMPC) during model updates to allow for privacy -preserving computation and aggregation of the privatized update data at the central server.• The secure multi-party computation enables clients to compute an average of private inputs without revealing individual data points.• Utilizing homomorphic encryption (HE) for the transmission and aggregation of the privatized update data, enabling clients to decrypt the aggregated model without the aggregator accessing the actual model updates at the central server.• The homomorphic encryption employs the Paillier algorithm to facilitate the encryption process.• Applying a differential privacy algorithm with a set gradient norm bound (C) and noise scale (G) during the aggregation process to thereby ensure privacy guarantees at the central server.• The differential privacy algorithm utilizes a privacy accounting method to compute an overall privacy cost using parameters a and 5.• The model updates between the local clients and the central server are transmitted by communicating only the differences between models, thereby enhancing data security during transmission.• The differences between models are computed based on the current local model and a previously aggregated model provided by the central server.• Employing a stochastic gradient descent (SGD) algorithm with differentially private updates to minimize an error in the aggregated model while preserving privacy at the central server.• The stochastic gradient descent algorithm includes steps for sampling, computing gradients, clipping gradients, adding noise, and descent updates.• Computing an aggregate function over the private data from the local clients using secure transmission methods without revealing individual private data at the central server.• Adjusting the weighting factor of the aggregated model based on a reliability metric associated with each local client at the central server.• Utilizing Secure Multi-Party Computation (SMPC) protocols to prepare local model updates for transmission thereby preserving privacy by obscuring individual updates at both the central server and the local clients.• The SMPC protocols include operations to add cryptographic noise to the local model updates before transmission.• Each local client performs secure transformations on the local model updates as part of the SMPC protocol to thereby generate secure, shareable versions of their updates.• Distributing the aggregated global model to the clients in a manner that each client can only decrypt the updates relevant to their local model, maintaining the confidentiality of the global model at the central server.• The privatized update data from a subset of local clients is aggregated and transmitted to the central server as a batch to thereby obscure the contribution of individual local clients.• Executing a training process on the local model using the local dataset to generate update data corresponding to the private data, and performing computations in accordance with the SMPC protocol to generate a secure, shareable version of the update data thereby defining the privatized update data, the computations including adding cryptographic noise or performing secure transformations to thereby obscure the update data to prevent inference of the private data at each local client; and aggregating the privatized update data from the plurality of local clients using SMPC protocols thereby ensuring the aggregation process does not reveal any local update data from any of the plurality of local clients, updating the aggregated model to form an updated version of the aggregate model thereby reflecting combined learning from the plurality of local clients without accessing the private data of any of the plurality of local clients, and distributing the new version of the aggregate model to the plurality of local clients for updating each respective local models at the central server.• A system for estimating impact while preserving privacy, comprising: a plurality of local clients, each local client comprising: a database having a local dataset containing private data; a processor configured to train a local model using the local dataset and generate privatized update data; and a transmitter configured to transmit the privatized update data from the local model, wherein the plurality of local clients are configured to update the aggregated model in a distributed manner without accessing another local client’s local dataset.

[0229] The following examples pertain to further embodiments. The following examples of the present disclosure may comprise subject material such as at least one device, a method, at least one machine-readable medium for storing instructions that when executed cause a machine to perform acts based on the method, means for performing acts based on the method and / or a system for estimating impact value.• According to example 1, there is provided a system. A system for estimating impact while preserving privacy, which may comprise: a plurality of local clients, each local client comprising: a database having a local dataset containing private data; a processor configuredto train a local model using the local dataset and generate privatized update data; and a transmitter configured to transmit the privatized update data from the local model; and a central server comprising: a receiver configured to receive the privatized update data from each of the plurality of local clients; an aggregated model configured to estimate an impact value; and a processor configured to update the aggregated model without accessing the private data in the local dataset of any of the plurality of local clients.• Example 2 may comprise the system according to example 1, where each local client may further comprise a local query component configured to query the local model.• Example 3 may comprise the system according to examples 1 or 2, wherein the local model is an identical copy of the aggregated model.• Example 4 may comprise the system according to example 1 or 2, wherein the local model is a copy of the aggregated model that has been updated by the local dataset.• Example 5 may comprise the system according to example 1, wherein the central server may comprise a query component configured to query the aggregated model to receive the impact value while thereby preserving privacy of the private data of all of the plurality of local clients.• Example 6 may comprise the system according to example 5, wherein the query performed by the query component within the central server is configured to perform the query during the updating of the aggregated model.• Example 7 may comprise the system according to example 6, wherein the query performed by the query component within the central server performs the queries for at least one epoch of training.• Example 8 may comprise the system according to example 6, the query performed by the query component within the central server corresponds to a single impact value.• Example 9 may comprise the system according to example 5, wherein the central server may further comprise a score component configured to generate a life-cycle score using a plurality of queries to the query component configured to receive a plurality of impact values comprising the impact value.• Example 10 may comprise the system according to example 1 , wherein the privatized update data consists of only weight update values.• Example 11 may comprise the system according to example 1, wherein the processor of the central server is further configured to update the aggregated model by aggregating the privatized update data from the plurality of local clients without accessing the private data in the local dataset of any of the plurality of clients.• Example 12 may comprise the system according to example 1, wherein the impact value is one of a greenhouse gas emission value, an energy consumption value, a water use value, a waste generation value, and an economic impact value.• Example 13 may comprise the system according to example 1, wherein the private data corresponds to data of a product.• Example 14 may comprise the system according to example 13, wherein the impact value is a product impact value.• Example 15 may comprise the system according to example 14, wherein the product impact value corresponds to a production activity of the product.• Example 16. may comprise the system according to example 15, wherein the production activity is one of a material transport, a material extraction, a material handling, a material use, a material disposal.• Example 17 may comprise the system according to example 1, wherein the impact value is an impact score.• Example 18 may comprise the system according to example 1, wherein the impact value is one of a climate change value, a greenhouse gas emission value, a carbon-dioxide value, a carbon-dioxide equivalent value, an ozone depletion value, a CFC-11 equivalent value, a toxicity value, a human toxicity value, a particulate matter value, an animal toxicity value, an aquatic toxicity value, a particulate matter formation value, an ionizing radiation value, a photochemical ozone formation value, an acidification value, a water acidification value, an ocean acidification value, a SO2 value, an NOx value, a NH3 value, an eutrophication value, a terrestrial eutrophication value, a freshwater eutrophication value, a marine eutrophication value, a nitrogen eutrophication value, a phosphorus eutrophication value, an ecotoxicity value, a water use value, a resource depletion value, a mineral depletion value, a fossil depletion value, a land-use value, a habitat transformation value, a habitat loss value, a biodiversity loss value, a renewable-energy consumption value, a non-renewable-energy consumption value, an energy consumption value, a waste generation value, an economic impact value, a carbon footprint value, a water footprint value, a nitrification value, a rare earth element depletion value, a volatile organic compound emissions value, a heavy metal emissions value, a pesticide residues value, a herbicide residues value, a radioactive contamination value, a biocidal product residues value, a dioxin emission value, a furan emission value, a polychlorinated biphenyls (pcb) emissions value, a noise pollution levels value, an alight pollution value, a thermal pollution from heat emissions value, a microplastic emissions value, a nanomaterial releases value, an antibiotic residues value, an endocrine disruptor emissions value, an algal bloom potential value, a soil contamination levels value, a soil salinization value, a land degradation value, a desertification risk value, a cultural heritage impact value, a visual impact value, an electromagnetic radiation value, a biodiversity impact on endangered species value, an invasive species propagation value, a life cycle sustainability index value, a carbon sequestration impact value, a greenhouse gas intensity value, a water quality index value, a sediment quality value, a turbidity levels value, a disposal impact value, a recycling rate value, a material recovery efficiency value, a supply chain vulnerability value, a labor practices impact value, a community health impact value, an economic dependency value, a regulatory compliance level value, a green certification status value, an operational energy efficiency value, an embodied energy value, an embodied carbon value, a total ownership cost value, a market dependency value, a technological obsolescence rate value, an adaptability to climate change value, a resilience to natural disasters value, and a crisis response capability value.• Example 19 may comprise the system according to example 1, wherein central server may further comprise a score component configured to generate at least one of a health score, an ecosystem score, and a natural resources score based upon the impact value.• Example 20 may comprise the system according to example 1, wherein the central server may further comprise a score component configured to convert the impact value to a score using a conversion factor obtained from a database• Example 21 may comprise the system according to example 1, wherein each local client of the plurality of local clients may further comprise a training component configured to train the local model using the local dataset to generate update data.• Example 22 may comprise the system according to example 21, wherein the training component fine tunes the local model to thereby train the local model using the local dataset to generate the update data.• Example 23 may comprise the system according to example 21, wherein each local client of the plurality of local clients may further comprise a privatization component configured to privative the update data to generate the privatized updated data by added noise to the update data.• Example 24 may comprise the system according to example 21 or 23, wherein each local client of the plurality of local clients may further comprise a privatization component configured to privatize the update data to generate the privatized updated data by clipping the update data.• Example 25 may comprise the system according to example 21, wherein each local client of the plurality of local clients may further comprise a privatization component configured to clip the update data and add noise to the clipped update data to generate the privatized update data.• Example 26 may comprise the system according to example 23 or 25, wherein the added noise is one of Gaussian noise and Laplacian noise.• Example 27 may comprise the system according to example 1 , wherein the privatized update data transmitted from each local client corresponds to only differences between the local model and the aggregated model.• Example 28 may comprise the system according to example 1, wherein the central server may further comprise an optimization component configured to minimize an error in the aggregated model by adjusting a plurality of model parameters of the aggregated model.• Example 29 may comprise the system according to example 1, wherein the central server is configured to receive a plurality of privatized update data comprising the privatized update data from the plurality of local clients, wherein the central server may further comprise an optimization component configured to apply a weighting factor to each of the plurality of privatized update data received from the plurality of local clients based on a metric of the respective local client of the plurality of local clients.• Example 30 may comprise the system according to example 29, wherein the metric corresponds to a size of a company.• Example 31 may comprise the system according to example 29, wherein the metric corresponds to one of: amount of data in the local dataset, variability of data in the local dataset, industry segment of the local client, historical accuracy of update data from the local client.• Example 32 may comprise the system according to example 1, wherein the central server is configured to add noise to the privatized update data.• Example 33 may comprise the system according to example 1, wherein the central server is configured to add noise to weights of the aggregated model.• Example 34 may comprise the system according to example 1, wherein the central server is configured to communicate the aggregated model to each of the local clients, wherein each of the plurality of local clients uses the communicated aggregated model as the local model.• Example 35 may comprise the system according to example 1, wherein the central server is configured to employ secure multi-party computation (SMPC) during model updates to allow for privacy-preserving computation and aggregation of the privatized update data.• Example 36 may comprise the system according to example 35, wherein the secure multiparty computation enables clients to compute an average of private inputs without revealing individual data points.• Example 37 may comprise the system according to example 1, wherein the central server utilizes homomorphic encryption (HE) for the transmission and aggregation of the privatizedupdate data, enabling clients to decrypt the aggregated model without the aggregator accessing the actual model updates.• Example 38 may comprise the system according to example 37, wherein the homomorphic encryption employs the Paillier algorithm to facilitate the encryption process.• Example 39 may comprise the system according to example 1, wherein the central server applies a differential privacy algorithm with a set gradient norm bound (C) and noise scale (G) during the aggregation process to ensure privacy guarantees.• Example 40 may comprise the system according to example 39, wherein the differential privacy algorithm utilizes a privacy accounting method to compute an overall privacy cost using parameters a and 5.• Example 41 may comprise the system according to example 1, wherein model updates between the local clients and the central server are transmitted by communicating only the differences between models, thereby enhancing data security during transmission.• Example 42 may comprise the system according to example 41, wherein the differences between models are computed based on the current local model and a previously aggregated model provided by the central server.• Example 43 may comprise the system according to example 1, wherein the central server employs a stochastic gradient descent (SGD) algorithm with differentially private updates to minimize an error in the aggregated model while preserving privacy.• Example 44 may comprise the system according to example 43, wherein the stochastic gradient descent algorithm comprises steps for sampling, computing gradients, clipping gradients, adding noise, and descent updates.• Example 45 may comprise the system according to example 1, wherein the central server is configured to compute an aggregate function over the private data from the local clients using secure transmission methods without revealing individual private data.• Example 46 may comprise the system according to example 1, wherein the central server is configured to adjust the weighting factor of the aggregated model based on a reliability metric associated with each local client.• Example 47 may comprise the system according to example 1, wherein the central server and the local clients utilize Secure Multi-Party Computation (SMPC) protocols to prepare local model updates for transmission thereby preserving privacy by obscuring individual updates.• Example 48 may comprise the system according to example 47, wherein the SMPC protocols comprise operations to add cryptographic noise to the local model updates before transmission.• Example 49 may comprise the system according to example 1, wherein each local client performs secure transformations on the local model updates as part of the SMPC protocol to thereby generate secure, shareable versions of their updates. <PLACEHOLDER>• Example 50 may comprise the system according to example 1, wherein the central server distributes the aggregated global model to the clients in a manner that each client can only decrypt the updates relevant to their local model, maintaining the confidentiality of the global model.• Example 51 may comprise the system according to example 1 , wherein the privatized update data from a subset of local clients is aggregated and transmitted to the central server as a batch to thereby obscure the contribution of individual local clients.• Example 52 may comprise the system according to example 1, wherein each local client's processor is further configured to: execute a training process on the local model using thelocal dataset to generate update data corresponding to the private data; and perform computations in accordance with the SMPC protocol to generate a secure, shareable version of the update data thereby defining the privatized update data, the computations comprising adding cryptographic noise or performing secure transformations to thereby obscure the update data to prevent inference of the private data; and wherein the central server's processor is further configured to: aggregate the privatized update data from the plurality of local clients using SMPC protocols thereby ensuring the aggregation process does not reveal any local update data from any of the plurality of local clients; update the aggregated model to form an updated version of the aggregate model thereby reflecting combined learning from the plurality of local clients without accessing the private data of any of the plurality of local clients; and distribute the new version of the aggregate model to the plurality of local clients for updating each respective local models.• Example 53 may comprise the system according to example 1, wherein each local client may further comprise a data preprocessing component configured to transform the local dataset prior to training the local model.• Example 54 may comprise the system according to example 1, wherein the data preprocessing component is configured to perform at least one of data normalization, feature extraction, dimensionality reduction, and handling of missing values.• Example 55 may comprise the system according to example 54, wherein the dimensionality reduction is performed by at least one of principal component analysis, independent component analysis, and t-distributed stochastic neighbor embedding.• Example 56 may comprise the system according to example 1, wherein the central server may further comprise a validation component configured to assess accuracy of the aggregated model using a held-out test dataset.• Example 57 may comprise the system according to example 56, wherein the validation component is further configured to quantify privacy risk of the aggregated model.• Example 58 may comprise the system according to example 1, wherein the central server may further comprise a monitoring component configured to evaluate fairness and bias metrics for the aggregated model.• Example 59 may comprise the system according to example 1, wherein the query component is further configured to provide an explanation of the reasoning behind an impact value prediction.• Example 60 may comprise the system according to example 1, wherein the central server may further comprise a user interface configured to visually depict life cycle stages corresponding to an impact value.• Example 61 may comprise the system according to example 1 , wherein the privatized update data may comprise differentially private stochastic gradient updates.• Example 62 may comprise the system according to example 1, wherein the transmitter and receiver employ federated learning protocols for decentralized training.• According to example 63, there is provided a method. The method for estimating impact while preserving privacy, the method may further comprise: training a local model using a local dataset containing private data at each of a plurality of local clients; generating privatized update data from the local model at each of the plurality of local clients; transmitting the privatized update data from each of the plurality of local clients to a central server; receiving the privatized update data from each of the plurality of local clients at the central server; and updating an aggregated model at the central server by aggregating theprivatized update data from the plurality of local clients without accessing the private data in the local dataset of any of the plurality of local clients.• Example 64 may comprise the method according to example 63, may further comprise querying the aggregated model at the central server to receive an impact value while thereby preserving privacy of the private data of all of the plurality of local clients.• Example 65 may comprise the method according to example 63, may further comprise distributing the updated aggregated model to the plurality of local clients.• Example 66 may comprise the method according to example 65, may further comprise replacing the local model with the updated aggregated model at each of the plurality of local clients.• Example 67 may comprise the method according to example 63, may further comprise querying the aggregated model at the central server for updating the aggregated model.• Example 68 may comprise the method according to example 63, may further comprise querying the local model prior to training the local model.• Example 69 may comprise the method according to example 68, wherein the local model and the aggregated model are identical.• Example 70 may comprise the method according to example 63, wherein the impact value is one of a greenhouse gas emission value, an energy consumption value, a water use value, a waste generation value, and an economic impact value.• Example 71 may comprise the method according to example 63, wherein the private data corresponds to data of a product.• Example 72 may comprise the method according to example 71, wherein the impact value is a product impact value.• Example 73 may comprise the method according to example 72, wherein the product impact value corresponds to a production activity of the product.• Example 74 may comprise the method according to example 73, wherein the production activity is one of a material transport, a material extraction, a material handling, a material use, a material disposal.• Example 75 may comprise the method according to example 63, wherein the impact value is an impact score.• Example 76 may comprise the method according to example 63, may further comprise generating a life-cycle score using a plurality of queries to the query component configured to receive a plurality of impact values comprising the impact value at the central server.• Example 77 may comprise the method according to example 76, wherein each query of the plurality of queries corresponds to a single impact value of the plurality of impact values.• Example 78 may comprise the method according to example 63, may further comprise generating at least one of a health score, an ecosystem score, and a natural resources score based upon the impact value at the central server.• Example 79 may comprise the method according to example 63, may further comprise converting the impact value to a score using a conversion factor obtained from a database at the central server.• Example 80 may comprise the method according to example 63, may further comprise training the local model using the local dataset to generate update data at each local client of the plurality of local clients.• Example 81 may comprise the method according to example 80, may further comprise finetuning the local model to thereby train the local model using the local dataset to generate the update data at each local client of the plurality of local clients.• Example 82 may comprise the method according to example 80 or 84, may further comprise privatizing the update data to generate the privatized updated data by adding noise to the update data at each local client of the plurality of local clients.• Example 83 may comprise the method according to example 80 or 82, may further comprise privatizing the update data to generate the privatized updated data by clipping the update data at each local client of the plurality of local clients.• Example 84 may comprise the method according to example 80, may further comprise clipping the update data and adding noise to the clipped update data to generate the privatized update data at each local client of the plurality of local clients.• Example 85 may comprise the method according to example 82 or 84, wherein the added noise is one of Gaussian noise and Laplacian noise.• Example 86 may comprise the method according to example 63, wherein the privatized update data transmitted from each local client corresponds to only differences between the local model and the aggregated model.• Example 87 may comprise the method according to example 63, may further comprise minimizing an error in the aggregated model by adjusting a plurality of model parameters of the aggregated model at the central server.• Example 88 may comprise the method according to example 63, may further comprise receiving a plurality of privatized update data comprising the privatized update data from the plurality of local clients, and applying a weighting factor to each of the plurality of privatized update data received from the plurality of local clients based on a metric of the respective local client at the central server.• Example 89 may comprise the method according to example 88, wherein the metric corresponds to a size of a company.• Example 90 may comprise the method according to example 88, wherein the metric corresponds to one of: amount of data in the local dataset, variability of data in the local dataset, industry segment of the local client, historical accuracy of update data from the local client.• Example 91 may comprise the method according to example 63, may further comprise adding noise to the privatized update data at the central server.• Example 92 may comprise the method according to example 63, may further comprise adding noise to weights of the aggregated model at the central server.• Example 93 may comprise the method according to example 63, may further comprise communicating the aggregated model to each of the local clients, wherein each of the plurality of local clients uses the communicated aggregated model as the local model.• Example 94 may comprise the method according to example 63, may further comprise employing secure multi-party computation (SMPC) during model updates to allow for privacy -preserving computation and aggregation of the privatized update data at the central server.• Example 95 may comprise the method according to example 94, wherein the secure multiparty computation enables clients to compute an average of private inputs without revealing individual data points.• Example 96 may comprise the method according to example 63, may further comprise utilizing homomorphic encryption (HE) for the transmission and aggregation of the privatized update data, enabling clients to decrypt the aggregated model without the aggregator accessing the actual model updates at the central server.• Example 97 may comprise the method according to example 96, wherein the homomorphic encryption employs the Paillier algorithm to facilitate the encryption process.• Example 98 may comprise the method according to example 63, may further comprise applying a differential privacy algorithm with a set gradient norm bound (C) and noise scale (G) during the aggregation process to thereby ensure privacy guarantees at the central server.• Example 99 may comprise the method according to example 98, wherein the differential privacy algorithm utilizes a privacy accounting method to compute an overall privacy cost using parameters a and 5.• Example 100 may comprise the method according to example 63, wherein model updates between the local clients and the central server are transmitted by communicating only the differences between models, thereby enhancing data security during transmission.• Example 101 may comprise the method according to example 100, wherein the differences between models are computed based on the current local model and a previously aggregated model provided by the central server.• Example 102 may comprise the method according to example 63, may further comprise employing a stochastic gradient descent (SGD) algorithm with differentially private updates to minimize an error in the aggregated model while preserving privacy at the central server.• Example 103 may comprise the method according to example 102, wherein the stochastic gradient descent algorithm comprises steps for sampling, computing gradients, clipping gradients, adding noise, and descent updates.• Example 104 may comprise the method according to example 63, may further comprise computing an aggregate function over the private data from the local clients using secure transmission methods without revealing individual private data at the central server.• Example 105 may comprise the method according to example 63, may further comprise adjusting the weighting factor of the aggregated model based on a reliability metric associated with each local client at the central server.• Example 106 may comprise the method according to example 63, may further comprise utilizing Secure Multi-Party Computation (SMPC) protocols to prepare local model updates for transmission thereby preserving privacy by obscuring individual updates at both the central server and the local clients.• Example 107 may comprise the method according to example 106, wherein the SMPC protocols comprise operations to add cryptographic noise to the local model updates before transmission.• Example 108 may comprise the method according to example 63, wherein each local client performs secure transformations on the local model updates as part of the SMPC protocol to thereby generate secure, shareable versions of their updates.• Example 109 may comprise the method according to example 63, may further comprise distributing the aggregated global model to the clients in a manner that each client can only decrypt the updates relevant to their local model, maintaining the confidentiality of the global model at the central server.• Example 110 may comprise the method according to example 63, wherein the privatized update data from a subset of local clients is aggregated and transmitted to the central server as a batch to thereby obscure the contribution of individual local clients.• Example 111 may comprise the method according to example 63, may further comprise executing a training process on the local model using the local dataset to generate update data corresponding to the private data, and performing computations in accordance with the SMPC protocol to generate a secure, shareable version of the update data thereby definingthe privatized update data, the computations comprising adding cryptographic noise or performing secure transformations to thereby obscure the update data to prevent inference of the private data at each local client; and aggregating the privatized update data from the plurality of local clients using SMPC protocols thereby ensuring the aggregation process does not reveal any local update data from any of the plurality of local clients, updating the aggregated model to form an updated version of the aggregate model thereby reflecting combined learning from the plurality of local clients without accessing the private data of any of the plurality of local clients, and distributing the new version of the aggregate model to the plurality of local clients for updating each respective local models at the central server.• According to example 112, a system may be provided. The system for estimating impact while preserving privacy, which may comprise: a plurality of local clients, each local client comprising: a database having a local dataset containing private data; a processor configured to train a local model using the local dataset and generate privatized update data; and a transmitter configured to transmit the privatized update data from the local model, wherein the plurality of local clients are configured to update the aggregated model in a distributed manner without accessing another local client’s local dataset.

Claims

What is claimed is:

1. A system for estimating impact while preserving privacy, comprising: a plurality of local clients, each local client comprising: a database having a local dataset containing private data; a processor configured to train a local model using the local dataset and generate privatized update data; and a transmitter configured to transmit the privatized update data from the local model; and a central server comprising: a receiver configured to receive the privatized update data from each of the plurality of local clients; an aggregated model configured to estimate an impact value; and a processor configured to update the aggregated model without accessing the private data in the local dataset of any of the plurality of local clients.

2. The system according to any preceding claim, where each local client further comprises a local query component configured to query the local model.

3. The system according to any preceding claim, wherein the local model is an identical copy of the aggregated model.

4. The system according to any preceding claim, wherein the local model is a copy of the aggregated model that has been updated by the local dataset.

5. The system according to any preceding claim 1, wherein the central server comprises a query component configured to query the aggregated model to receive the impact value while thereby preserving privacy of the private data of all of the plurality of local clients.

6. The system according to any preceding claim, wherein the query performed by the query component within the central server is configured to perform the query during the updating of the aggregated model.

7. The system according to any preceding claim, wherein the query performed by the query component within the central server performs the queries for at least one epoch of training.

8. The system according to any preceding claim, the query performed by the query component within the central server corresponds to a single impact value.

9. The system according to any preceding claim, wherein the central server further comprises a score component configured to generate a life-cycle score using a plurality of queries to the query component configured to receive a plurality of impact values including the impact value.

10. The system according to any preceding claim, wherein the privatized update data consists of only weight update values.

11. The system according to any preceding claim, wherein the processor of the central server is further configured to update the aggregated model by aggregating the privatized update data fromthe plurality of local clients without accessing the private data in the local dataset of any of the plurality of clients.

12. The system according to any preceding claim, wherein the impact value is one of a greenhouse gas emission value, an energy consumption value, a water use value, a waste generation value, and an economic impact value.

13. The system according to any preceding claim, wherein the private data corresponds to data of a product.

14. The system according to any preceding claim, wherein the impact value is a product impact value.

15. The system according to any preceding claim, wherein the product impact value corresponds to a production activity of the product.

16. The system according to any preceding claim, wherein the production activity is one of a material transport, a material extraction, a material handling, a material use, a material disposal.

17. The system according to any preceding claim, wherein the impact value is an impact score.

18. A method for estimating impact while preserving privacy, the method comprising: training a local model using a local dataset containing private data at each of a plurality of local clients; generating privatized update data from the local model at each of the plurality of local clients; transmitting the privatized update data from each of the plurality of local clients to a central server; receiving the privatized update data from each of the plurality of local clients at the central server; and updating an aggregated model at the central server by aggregating the privatized update data from the plurality of local clients without accessing the private data in the local dataset of any of the plurality of local clients.

19. The method according to claim 18, further comprising querying the aggregated model at the central server to receive an impact value while thereby preserving privacy of the private data of all of the plurality of local clients.

20. The method according to claims 18 to 19, further comprising distributing the updated aggregated model to the plurality of local clients.

21. The method according to claims 18 to 20, further comprising replacing the local model with the updated aggregated model at each of the plurality of local clients.

22. The method according to claims 18 to 21, further comprising querying the aggregated model at the central server for updating the aggregated model.

23. The method according to claims 18 to 22, wherein the impact value is one of a greenhouse gas emission value, an energy consumption value, a water use value, a waste generation value, and an economic impact value.

24. The method according to claims 18 to 23, further comprising generating a life-cycle score using a plurality of queries to the query component configured to receive a plurality of impact values including the impact value at the central server.

Citation Information

Patent Citations

  • Privacy-safe building cluster energy consumption collaborative prediction method and system

    CN115409370A

  • Multi-region water demand prediction method for urban graded collaborative water supply

    CN115630745A

  • Photovoltaic power generation capacity prediction method adopting environmental perception under federated learning architecture

    CN117540422A