A zero-trust driven supply chain real-time collaborative optimization and security protection method

CN121503797BActive Publication Date: 2026-08-21CHAOHU UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511687446.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-18
Publication Date
2026-08-21
Estimated Expiration
2045-11-18

AI Technical Summary

Technical Problem

[0004]风险识别滞后:传统供应链依赖周期性审计或事后评估,无法实时感知参与方的可靠性变化(如供应商财务危机、合规违规或生产中断)

Benefits of technology

[0056] 1. The zero-trust driven supply chain real-time collaborative optimization and security protection method abandons the assumption of traditional static trust under the zero-trust architecture, constructs a credibility assessment strategy and a multi-objective collaborative optimization method, and achieves the following objectives of the supply chain system: (1) Credibility assessment method: a process of comprehensively evaluating the reliability, compliance and risk resistance of each participant in the supply chain, identifying potential risk points, and ensuring the stability and resilience of the supply chain; (2) Multi-objective collaborative optimization: coordinating the resources, information and decision-making processes of each link in the supply chain, with the goal of maximizing credibility, minimizing overall cost, maximizing efficiency and improving system flexibility, to achieve a dynamic optimization process that achieves a win-win situation for all parties.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121503797B_ABST
    Figure CN121503797B_ABST
Patent Text Reader

Abstract

The application discloses a kind of zero trust drive's supply chain real-time collaborative optimization and security protection method, and the application is related to supply chain management technical field.The zero trust drive's supply chain real-time collaborative optimization and security protection method includes: S1 can be trusted to assess calculation: based on direct trust and indirect trust comprehensive calculation final trust;S2 multi-objective collaborative optimization;S3 utilize MAPPO to solve communication strategy optimization problem;Under the architecture of zero trust, discard the assumption of traditional static trust, build the trust evaluation strategy and multi-objective collaborative optimization method, realize the reliability, compliance and risk resistance ability of each participant in supply chain are comprehensively evaluated process, potential risk point can be identified, ensure the stability and flexibility of supply chain;And realize the coordination of supply chain each link's resource, information and decision-making process, to maximize the trust, minimize overall cost, maximize efficiency, improve system flexibility as goal, realize the dynamic optimization process of multi-party win-win.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of supply chain management technology, specifically a zero-trust-driven method for real-time collaborative optimization and security protection of the supply chain. Background Technology

[0002] A supply chain refers to the entire network structure that begins with the procurement of raw materials, proceeds through production, processing, transportation, warehousing, and sales, and ends with the delivery of the final product or service to the customer. It encompasses all activities related to the flow of products or services both inside and outside the enterprise, including participants at each stage such as suppliers, manufacturers, distributors, retailers, and end consumers.

[0003] In traditional supply chain management models, trust relationships among various participants (suppliers, manufacturers, distributors, retailers, etc.) are largely based on static assumptions, establishing a one-time trust assessment through contracts, historical cooperation records, or industry certifications. For example, patent application CN115146455B discloses a computationally experimentally supported multi-objective decision-making method for complex supply chains. In this patent application, the supply chain model mainly consists of a static directed graph reflecting the supply chain structure and a function DF reflecting the supply chain evolution trend. In the directed graph of the supply chain structure, node V has its individual attribute characteristics C, its decision-making mechanism D, and constraints R. E represents the interaction methods and connections between supply chain entities, establishing a static supply chain model based on the directed graph. This model exposes the following limitations of static trust when facing complex and ever-changing business environments:

[0004] Lagging risk identification: Traditional supply chains rely on periodic audits or post-event assessments, making it impossible to detect changes in the reliability of participants in real time (such as supplier financial crises, compliance violations, or production disruptions). For example, suppliers may experience delivery delays due to sudden natural disasters, but traditional systems cannot provide early warnings.

[0005] Insufficient compliance verification: Compliance checks typically rely on manual audits or periodic reports, making it difficult to cover dynamic compliance risks across the entire supply chain (such as updates to environmental regulations and data privacy requirements).

[0006] Lack of risk resilience assessment: The traditional model does not establish a quantitative mechanism to assess the participants' ability to withstand risks such as market fluctuations and supply chain disruptions, resulting in insufficient resilience. Summary of the Invention

[0007] To address the shortcomings of existing technologies, this invention provides a zero-trust-driven real-time collaborative optimization and security protection method for supply chains, thus solving the problems of existing technologies.

[0008] To achieve the above objectives, the present invention provides the following technical solution: a zero-trust-driven real-time collaborative optimization and security protection method for supply chains, comprising the following steps:

[0009] S1 Trustworthiness Assessment Calculation: The final trust is calculated based on a combination of direct and indirect trust; where:

[0010] S1.1 Direct Trust Calculation: Based on historical performance and multiple trust dimensions of compliance certification, historical data is collected and normalized, and the time decay factor is taken into account for comprehensive calculation. The weight of each dimension is dynamically adjusted according to the market environment. The model is trained using machine learning algorithms to automatically adjust, while real-time data is also included in the evaluation.

[0011] S1.2 Indirect Trust Calculation: Trust information of neighboring nodes is aggregated using a Graph Neural Network (GCN). A spatiotemporal graph network and multi-head attention mechanism are introduced for calculation. Node feature vectors and edge weights are defined, and transaction frequency is considered during modeling. A time dimension is added to the graph convolution operation to construct a spatiotemporal graph. Attention weights are calculated using a multi-head attention mechanism, and the output is obtained. Finally, the node embeddings are updated by combining the outputs of spatiotemporal graph convolution and multi-head attention. A fully connected network maps the node embeddings to indirect trust values.

[0012] S2 Multi-Objective Collaborative Optimization: Defines a multi-dimensional state space vector, a hybrid action space, and an objective vector; calculates the Pareto solution based on the objective function; designs basic rewards, Pareto distance rewards, and dominance intensity rewards; combines them to obtain a comprehensive reward value; and adds risk perception factors to form a risk perception reward mechanism.

[0013] S3 uses MAPPO to solve the communication policy optimization problem: it performs initialization, simulates the interaction between the reinforcement learning agent and the environment to collect data, updates the network parameters from the buffer batch through the updates of the Critic network and the Actor network, and optimizes the policy set.

[0014] Preferably, the formula for calculating the final trust is:

[0015] T final =α·T direct +(1-α)T indirect ;

[0016] Among them, T direct and T indirect Let represent direct trust and indirect trust respectively, and let α be the weight of direct trust, which is an adjustable hyperparameter.

[0017] Preferably, the direct trust is calculated based on multiple trust dimensions, including historical performance, compliance certification, financial health, technical capabilities, risk response, and data credibility; and the weight ω of each trust dimension is determined. i Historical data for each dimension is collected and normalized. Trust scores for each dimension are calculated, and the time decay factor is considered. The scores for each dimension are then summarized to obtain the direct trust value T. direct ;

[0018] Weights ω of each trust dimension i The system is dynamically adjusted based on market conditions, industry trends, and supply chain characteristics; machine learning algorithms are used to train models based on historical data and automatically adjust the weights of each dimension.

[0019] Incorporate real-time data into direct trust assessments to reflect the current status of supply chain entities in a timely manner.

[0020] Preferably, the indirect trust calculation aggregates the trust information of neighboring nodes through a graph neural network (GCN), and introduces a spatiotemporal graph temporal network and a multi-head attention mechanism to calculate the indirect trust degree T of the nodes. indirect To capture the spatiotemporal dependencies and complex interactions between entities in the supply chain.

[0021] Preferably, the specific steps of the indirect trust calculation include:

[0022] Graph modeling: Define the node feature vector X i =d g , i∈I, g∈G, including direct trust score, business reputation score, and number of associated entities; where I represents the number of supply chain entities, d g Let G represent the g-th feature of the supply chain entity, where G represents the number of features.

[0023] Define the weight value E of the edge and the variable attribute feature e. ij ∈E, i, j∈I, represents the cooperative relationship between supply chain entities i and j, and the transaction frequency and cooperation amount are considered when modeling.

[0024] Introducing a spatiotemporal graph temporal network: By adding a time dimension to graph convolution operations, a spatiotemporal graph is constructed to capture the trust relationships of supply chain entities that change over time;

[0025] Introducing a multi-head attention mechanism: Using a multi-head attention mechanism to calculate attention weights;

[0026] Node embedding update: Update node embeddings by combining the outputs of spatiotemporal graph convolution and multi-head attention mechanism;

[0027] Indirect trust calculation: The output layer embeds nodes into an indirect trust value T through a fully connected network. indirect ;

[0028] Define a Graph Convolutional Network (GCN) to extract indirect trust from supply entities;

[0029] The output layer embeds nodes into an indirect trust value T via a fully connected network. indirect ∈[0,1].

[0030] Preferably, the state space definition in S2 includes:

[0031] State space: Define a multidimensional state space vector o t ∈O, containing the final trust level T final_t Supplier Service Quality (QoS) t Inventory levels, order backlog characteristics, and supplier service quality including cost and time, in the following specific forms:

[0032] ;

[0033] Among them, T final_t ∈[0,1), M is the time step length, and t∈1~M.

[0034] Preferably, the action space definition in S2 includes:

[0035] Action space: Define the hybrid action space a t ∈A, including the allocation of purchasing power b t Production plan adjustments and logistics route optimizations take the following forms:

[0036] ;

[0037] Among them, b t,i ∈(0,1), indicating whether the current purchasing right has been allocated to supplier i.

[0038] Preferably, the objective function design includes:

[0039] Define the target vector F=(f QoS f TR This includes maximizing Quality of Service (QoS) and maximizing Trustworthiness Reward (TR); specific objectives are as follows:

[0040] Maximizing Quality of Service (QoS): By weighting logistics time and price cost metrics, a unified objective is achieved, as shown in the following formula:

[0041] ;

[0042] Maximize Trustworthiness Reward (TR): Used to incentivize suppliers to increase trustworthiness; the specific formula is as follows:

[0043] .

[0044] Preferably, the reward function design includes:

[0045] Calculate the Pareto solution (GP) based on the objective function, and then calculate the reward function, which includes the basic reward r. base Pareto distance reward r dist and dominance intensity reward r domCombined with basic reward r base Pareto distance reward r dist and dominance intensity reward r dom The comprehensive reward value is obtained using the following formula:

[0046] r(o) t a t )= λ base r base + λ dist r dist + λ dom r dom ;

[0047] Where, λ base , λ dist , λ dom These represent the weights of the base reward, Pareto distance reward, and dominance reward, respectively.

[0048] By adding risk perception factors, penalizing high-risk decisions and rewarding low-risk decisions, defining risk indicators and incorporating them into the reward function, a risk perception-based reward mechanism is formed. The updated reward function is expressed as follows:

[0049] ;

[0050] Where, λ risk γ is the risk penalty coefficient. z For the weight of the z-th risk indicator, Risk z (o t a t ) represents the value of the z-th risk indicator.

[0051] Preferably, S3 specifically includes:

[0052] Initialization: Initialize network parameters, Pareto solution front set, and historical solution set;

[0053] Environmental interaction and experience collection: Simulate the interaction process between the reinforcement learning agent and the environment to collect new interaction data;

[0054] Network Update and Policy Optimization: The policy set is optimized by updating the Critic and Actor networks. Specific steps include sampling batches from the buffer, updating the Critic and Actor networks, and updating network parameters through gradient descent and gradient ascent.

[0055] This invention provides a zero-trust-driven method for real-time collaborative optimization and security protection of the supply chain. Compared with existing technologies, it has the following advantages:

[0056] 1. The zero-trust driven supply chain real-time collaborative optimization and security protection method abandons the assumption of traditional static trust under the zero-trust architecture, constructs a credibility assessment strategy and a multi-objective collaborative optimization method, and achieves the following objectives of the supply chain system: (1) Credibility assessment method: a process of comprehensively evaluating the reliability, compliance and risk resistance of each participant in the supply chain, identifying potential risk points, and ensuring the stability and resilience of the supply chain; (2) Multi-objective collaborative optimization: coordinating the resources, information and decision-making processes of each link in the supply chain, with the goal of maximizing credibility, minimizing overall cost, maximizing efficiency and improving system flexibility, to achieve a dynamic optimization process that achieves a win-win situation for all parties.

[0057] 2. This zero-trust-driven real-time collaborative optimization and security protection method for the supply chain provides a more comprehensive and accurate trust assessment method by comprehensively considering ultimate trust, direct trust, and indirect trust, thereby improving the reliability of the assessment results. Secondly, the direct trust calculation incorporates multiple trust dimensions and combines time decay factors and real-time data, making the assessment results more dynamic and real-time, reflecting the current state of supply chain entities in a timely manner. Furthermore, the indirect trust calculation aggregates trust information from neighboring nodes through a graph neural network (GCN) and introduces a spatiotemporal graph temporal network and a multi-head attention mechanism, effectively capturing the spatiotemporal dependencies and complex interactions between supply chain entities, improving the depth and breadth of the assessment. These improvements collectively enhance the accuracy and real-time performance of supply chain trust assessment, providing strong support for real-time collaborative optimization and security protection of the supply chain.

[0058] 3. This zero-trust-driven real-time collaborative optimization and security protection method for the supply chain comprehensively considers key factors in the supply chain, such as trust level, service quality, and inventory levels, by defining a multi-dimensional state space vector and a hybrid action space, making decision-making more scientific and rational. Secondly, the objective function design balances service quality and trustworthiness rewards, unifying the objective through weighted indicators and improving the overall optimization effect. Furthermore, the reward function is cleverly designed, combining basic rewards, Pareto distance rewards, and dominance strength rewards to effectively incentivize the agent to explore Pareto optimal solutions and favor solutions with strong dominance, driving the frontier expansion. The introduction of risk perception factors, penalizing high-risk decisions and rewarding low-risk decisions, enhances the robustness of decision-making. These improvements collectively enhance the comprehensiveness, scientific nature, and robustness of supply chain optimization, providing strong support for supply chain management.

[0059] 4. This zero-trust driven real-time collaborative optimization and security protection method for the supply chain utilizes MAPPO to solve the communication strategy optimization problem. By initializing network parameters, collecting experience through environmental interaction, and updating the network to optimize the strategy, it achieves dynamic adaptation and efficient decision-making in complex communication environments. Compared with existing technologies, it improves the real-time performance and accuracy of strategy optimization and enhances the adaptability and robustness of the system. Attached Figure Description

[0060] Figure 1 This is a schematic diagram of the overall steps of the present invention;

[0061] Figure 2 This is a schematic diagram of the process of S2 in this invention;

[0062] Figure 3 This is a schematic diagram of the process of S3 in this invention. Detailed Implementation

[0063] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0064] See Figures 1-3 The present invention provides the following four technical solutions:

[0065] First implementation method: A zero-trust driven real-time collaborative optimization and security protection method for supply chains, comprising the following steps:

[0066] S1 Trustworthiness Assessment Calculation: The final trust is calculated based on a combination of direct and indirect trust; where:

[0067] S1.1 Direct Trust Calculation: Based on historical performance and multiple trust dimensions of compliance certification, historical data is collected and normalized, and the time decay factor is taken into account for comprehensive calculation. The weight of each dimension is dynamically adjusted according to the market environment. The model is trained using machine learning algorithms to automatically adjust, while real-time data is also included in the evaluation.

[0068] S1.2 Indirect Trust Calculation: Trust information of neighboring nodes is aggregated using a Graph Neural Network (GCN). A spatiotemporal graph network and multi-head attention mechanism are introduced for calculation. Node feature vectors and edge weights are defined, and transaction frequency is considered during modeling. A time dimension is added to the graph convolution operation to construct a spatiotemporal graph. Attention weights are calculated using a multi-head attention mechanism, and the output is obtained. Finally, the node embeddings are updated by combining the outputs of spatiotemporal graph convolution and multi-head attention. A fully connected network maps the node embeddings to indirect trust values.

[0069] S2 Multi-Objective Collaborative Optimization: Defines a multi-dimensional state space vector, a hybrid action space, and an objective vector; calculates the Pareto solution based on the objective function; designs basic rewards, Pareto distance rewards, and dominance intensity rewards; combines them to obtain a comprehensive reward value; and adds risk perception factors to form a risk perception reward mechanism.

[0070] S3 uses MAPPO to solve the communication policy optimization problem: it performs initialization, simulates the interaction between the reinforcement learning agent and the environment to collect data, updates the network parameters from the buffer batch through the updates of the Critic network and the Actor network, and optimizes the policy set.

[0071] This invention, under a zero-trust architecture, abandons the assumption of traditional static trust and constructs a credibility assessment strategy and a multi-objective collaborative optimization method to achieve the following objectives of the supply chain system:

[0072] (1) Credibility assessment method: The process of comprehensively evaluating the reliability, compliance and risk resistance of each participant in the supply chain, identifying potential risk points, and ensuring the stability and resilience of the supply chain;

[0073] (2) Multi-objective collaborative optimization: Coordinate the resources, information and decision-making processes of each link in the supply chain to maximize credibility, minimize overall cost, maximize efficiency and improve system flexibility, and achieve a dynamic optimization process that wins the interests of all parties.

[0074] The second implementation differs from the first in that the trust assessment is based on the calculation of final trust, direct trust, and indirect trust. The formula for calculating final trust is:

[0075] T final =α·T direct +(1-α)T indirect ;

[0076] Among them, T direct and T indirect These represent direct trust and indirect trust, respectively, with α being the direct trust weight, which is an adjustable hyperparameter.

[0077] Direct Trust Calculation: Direct trust is calculated based on multiple trust dimensions, including historical performance, compliance certification, financial health, technological capabilities, risk response, and data credibility; the weight ω of each trust dimension is determined. i Historical data for each dimension is collected and normalized. Trust scores for each dimension are calculated, and the time decay factor is considered. The scores for each dimension are then summarized to obtain the direct trust value T. direct T direct The calculation formula is:

[0078] ;

[0079] Where 's' represents the trust dimension, see Table 1 for details. x s,t Let represent the normalized score of the s-th trust dimension at time t, M be the length of the time step, and t∈1~M; λ represent the time decay coefficient, and t0-t represent the difference between the current time and the historical data timestamp;

[0080] Weights ω of each trust dimension i Dynamic adjustments are made based on market conditions, industry trends, and supply chain characteristics; machine learning algorithms (such as decision trees or random forests) are used to train models based on historical data, automatically adjusting the weights of each dimension, as expressed by the formula:

[0081] ω i_new = f(ω i_old (market environment, industry trends, supply chain characteristics), where f is the weight adjustment function obtained through training;

[0082] Incorporating real-time data (such as real-time logistics information and inventory status) into direct trust assessments allows for a more timely reflection of the current status of supply chain entities. The formula is as follows:

[0083] x s,t_real = g(x s,t_historical (real-time data), where g is a data fusion function that combines historical data with real-time data to obtain an accurate normalized score;

[0084] Table 1. Recommendation factors for direct trust modeling

[0085]

[0086] Indirect Trust Calculation: Trust information of neighboring nodes is aggregated through a graph neural network (GCN), and a spatiotemporal graph temporal network and a multi-head attention mechanism are introduced to calculate the indirect trust degree T of nodes. indirect To capture the spatiotemporal dependencies and complex interactions between supply chain entities, the calculation formula is as follows:

[0087] T indirect =GNN(E,X);

[0088] Where X is the feature vector of a node in the graph neural network GCN, and E is the weight value of an edge in the graph neural network GCN.

[0089] The specific steps include:

[0090] Graph modeling: Define the node feature vector X i =d gLet i ∈ I, g ∈ G, including direct trust score, business reputation score, and number of associated entities; where I represents the number of supply chain entities (suppliers, logistics providers, etc.), and d g Let G represent the g-th feature of the supply chain entity, where G represents the number of features.

[0091] Define the weight value E of the edge and the variable attribute feature e. ij ∈E, i, j∈I, represents the cooperative relationship between supply chain entities i and j, and factors such as transaction frequency and cooperation amount are considered when modeling;

[0092] Introducing a spatiotemporal graph temporal network: By adding a time dimension to graph convolution operations, a spatiotemporal graph is constructed to capture the changing trust relationships between supply chain entities over time. The formula is expressed as:

[0093] ;

[0094] in, For attention weights, W l This is the weight matrix. Let be the embedding of node k in layer l, σ be the activation function, and b be the value of the embedding. l For bias terms;

[0095] Introducing a multi-head attention mechanism: Using a multi-head attention mechanism to calculate attention weights. :

[0096] ;

[0097] in, , These are the query and key matrices, respectively. Indicates to Perform the transpose operation, d k The dimension of the key;

[0098] Output of the multi-head attention mechanism:

[0099] ;

[0100] ;

[0101] Where s∈1~S, W O The output weight matrix is ​​used to concatenate the outputs of multiple heads and perform a linear transformation, mapping the output of the multi-head attention to the final output space.

[0102] This represents the query weight matrix for the s-th head, which incorporates the input node features. Mapped to the query space;

[0103] This represents the key weight matrix of the s-th head, which incorporates the input node features. Mapped to the key space;

[0104] The weight matrix represents the value of the s-th head, which incorporates the input node features. Mapped to value space;

[0105] Node embedding update:

[0106] By combining the outputs of spatiotemporal graph convolution and multi-head attention mechanism, update the node embeddings:

[0107] ;

[0108] Here, STGCN represents the spatiotemporal graph convolutional network operation, used to combine the current embedding of node s. Messages from its neighboring nodes To update the embedding of node s; It is a message (or feature information) passed from a neighboring node u of node s to node s at the (l+1)th layer; It is the embedding of node s in the l-th layer; N(s) is the set of messages passed from all neighboring nodes u of node s to node s; u∈N(s) means that u is a neighboring node of node s, and N(s) is the set of neighboring nodes of node s, that is, the set of nodes directly connected to node s in the graph structure.

[0109] Indirect trust level calculation:

[0110] The output layer embeds nodes into an indirect trust value T via a fully connected network. indirect :

[0111] ;

[0112] Among them, W out This is the weight matrix of the output layer. It is the embedding of node s in the Lth layer, b out Here, σ is the bias term, and σ is the activation function.

[0113] A Graph Convolutional Network (GCN) is defined to extract indirect trust from supply entities, and its formal representation is as follows:

[0114] ;

[0115] ;

[0116] ;

[0117] in, Let I(I) represent the nf-dimensional embedding of node i at layer l; let I(I) represent the set of neighborhood nodes of node i. gcn e (.) and gcn h (.) represent the edge operation and node operation functions, respectively, which are approximated using a multilayer perceptron.

[0118] The output layer embeds nodes into an indirect trust value T via a fully connected network. indirect ∈[0,1].

[0119] By comprehensively considering ultimate trust, direct trust, and indirect trust, a more comprehensive and accurate trust assessment method is provided, improving the reliability of the assessment results. Secondly, the direct trust calculation incorporates multiple trust dimensions and combines time decay factors and real-time data, making the assessment results more dynamic and real-time, reflecting the current state of supply chain entities in a timely manner. Furthermore, the indirect trust calculation aggregates trust information from neighboring nodes through a graph neural network (GCN) and introduces a spatiotemporal graph temporal network and a multi-head attention mechanism, effectively capturing the spatiotemporal dependencies and complex interactions between supply chain entities, improving the depth and breadth of the assessment. These improvements collectively enhance the accuracy and real-time performance of supply chain trust assessment, providing strong support for real-time collaborative optimization and security protection of the supply chain.

[0120] The third implementation method differs from the first in that it employs a multi-objective collaborative optimization method.

[0121] The definitions of state space and action space include:

[0122] State space: Define a multidimensional state space vector o t ∈O, containing the final trust level T final_t Supplier Service Quality (QoS) t Characteristics such as inventory levels and order backlog, and supplier service quality including cost and time, are detailed below:

[0123] ;

[0124] Among them, T final_t ∈[0,1), M is the time step length, and t∈1~M;

[0125] Action space: Define the hybrid action space a t ∈A, including the allocation of purchasing power b t Actions such as production plan adjustments and logistics route optimization take the following forms:

[0126] ;

[0127] Among them, bt,i ∈(0,1), indicating whether the current purchasing power has been allocated to supplier i;

[0128] The objective function design includes:

[0129] Define the target vector F=(f QoS f TR This includes maximizing Quality of Service (QoS) and maximizing Trustworthiness Reward (TR); specific objectives are as follows:

[0130] Maximizing Quality of Service (QoS): By weighting indicators such as logistics time and price cost, a unified objective is achieved, as shown in the following formula:

[0131] ;

[0132] Maximize Trustworthiness Reward (TR): Used to incentivize suppliers to increase trustworthiness; the specific formula is as follows:

[0133] ;

[0134] Reward function design:

[0135] Calculate the Pareto solution (GP) based on the objective function, and then calculate the reward function:

[0136] Basic reward r base Incentivize the agent to explore the Pareto optimal solution:

[0137] ;

[0138] Pareto distance reward r dist To incentivize the agent to explore Pareto optimal solutions, solutions that deviate significantly from the forefront are penalized.

[0139] ;

[0140] Where, β GP To control the hyperparameter of reward decay rate with distance, β GP The larger the β value, the more sensitive the reward is to distance; the reward for solutions far from the Pareto front decreases rapidly. GP The smaller the value, the higher the tolerance of the reward to changes in distance, encouraging broader exploration; F max and F min These represent the historical maximum and minimum values ​​of each objective function, respectively; F P This represents the Pareto solution that is closest to the current solution F;

[0141] Domination Intensity Reward r dom The preference for solutions with strong dominance drives the frontier to expand outward, as shown in the following formula:

[0142] ;

[0143] Domination bonus = Total number of Pareto front solutions / Number of historical solutions dominated by the current solution

[0144] Overall Reward: The overall reward value is obtained by combining the base reward, Pareto distance reward, and dominance reward. The specific formula is as follows:

[0145] r(o) t a t )= λ base r base + λ dist r dist + λ dom r dom ;

[0146] Where, λ base , λ dist , λ dom These represent the weights of the base reward, Pareto distance reward, and dominance reward, respectively.

[0147] By adding risk perception factors, penalizing high-risk decisions and rewarding low-risk decisions, defining risk indicators (such as supplier bankruptcy risk and logistics disruption risk) and incorporating them into the reward function, a risk perception-based reward mechanism is formed. The updated reward function is expressed as follows:

[0148] ;

[0149] Where, λ risk γ is the risk penalty coefficient. z For the weight of the z-th risk indicator, Risk z (o t a t ) represents the value of the z-th risk indicator.

[0150] By defining a multi-dimensional state space vector and a hybrid action space, key factors in the supply chain, such as trust, service quality, and inventory levels, are comprehensively considered, making decision-making more scientific and rational. Secondly, the objective function design balances service quality and trust rewards, unifying the objective through weighted indicators and improving overall optimization performance. Furthermore, the reward function is cleverly designed, combining basic rewards, Pareto distance rewards, and dominance strength rewards to effectively incentivize the agent to explore Pareto optimal solutions and favor solutions with strong dominance, driving the frontier expansion. The introduction of risk perception factors, penalizing high-risk decisions and rewarding low-risk decisions, enhances the robustness of decision-making. These improvements collectively enhance the comprehensiveness, scientific rigor, and robustness of supply chain optimization, providing strong support for supply chain management.

[0151] The fourth implementation method differs from the first in that it utilizes MAPPO to solve the communication strategy optimization problem.

[0152] Initialization: Initialize network parameters, Pareto solution front set, historical solution set, etc.; network parameters include;

[0153] Environmental interaction and experience collection: Simulate the interaction process between the reinforcement learning agent and the environment to collect new interaction data;

[0154] Network Update and Policy Optimization: The policy set is optimized by updating the Critic and Actor networks. Specific steps include sampling batches from the buffer, updating the Critic and Actor networks, and updating network parameters through gradient descent and gradient ascent.

[0155] By using MAPPO to solve the communication strategy optimization problem, the system collects experience through initializing network parameters and interacting with the environment, and updates the network to optimize the strategy. This enables dynamic adaptation and efficient decision-making in complex communication environments. Compared with existing technologies, it improves the real-time performance and accuracy of strategy optimization, and enhances the system's adaptability and robustness.

[0156] Furthermore, any content not described in detail in this specification is existing technology known to those skilled in the art.

[0157] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0158] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A zero-trust-driven real-time collaborative optimization and security protection method for supply chains, characterized in that, Includes the following steps: S1 Trustworthiness Assessment Calculation: Calculates the final trust based on a comprehensive analysis of direct and indirect trust; where: S1.1 Direct Trust Calculation: Based on multiple trust dimensions such as historical performance, compliance certification, financial health, technical capabilities, risk response, and data credibility, historical data is collected and normalized, and the time decay factor is taken into account for comprehensive calculation. The weight of each dimension is dynamically adjusted according to the market environment, and the model is trained using machine learning algorithms to automatically adjust it. Real-time data is also included in the evaluation. S1.2 Indirect Trust Calculation: Trust information of neighboring nodes is aggregated through a Graph Neural Network (GCN). A spatiotemporal graph temporal network and multi-head attention mechanism are introduced for calculation, defining a node feature vector X. i =d g , i∈I, g∈G, including direct trust score, business reputation score, and number of associated entities; where I represents the number of supply chain entities, d g Let G represent the g-th feature of a supply chain entity, where G represents the number of features; define the edge weight E, and the variable attribute feature e. ij ∈E, i, j∈I, representing the cooperative relationship between supply chain entities i and j, considering transaction frequency and cooperation amount factors during modeling; adding a time dimension to the graph convolution operation to construct a spatiotemporal graph; using a multi-head attention mechanism to calculate attention weights and obtain the output; finally, combining the spatiotemporal graph convolution and multi-head attention mechanism output to update node embeddings, and mapping node embeddings to indirect trust values ​​through a fully connected network; S2 Multi-Objective Collaborative Optimization: Defines a multi-dimensional state space vector, a hybrid action space, and an objective vector; calculates the Pareto solution based on the objective function; designs basic rewards, Pareto distance rewards, and dominance intensity rewards; combines them to obtain a comprehensive reward value; and adds risk perception factors to form a risk perception reward mechanism. S3 uses MAPPO to solve the communication policy optimization problem: it performs initialization, simulates the interaction between the reinforcement learning agent and the environment to collect data, updates the network parameters from the buffer batch through the updates of the Critic network and the Actor network, and optimizes the policy set. The state space definition in S2 includes: State space: Define a multidimensional state space vector o t ∈O, containing the final trust level T final_t Supplier Service Quality (QoS) t Inventory levels, order backlog characteristics, and supplier service quality including cost and time; The action space definition in S2 includes: Action space: Define the hybrid action space a t ∈A, including the allocation of purchasing power b t Production plan adjustments and logistics route optimization actions; The objective function design includes: Define the target vector F=(f QoS f TR The objective vector aims to maximize f. QoS and f TR f QoS The specific formula is as follows: ; f TR The specific formula is as follows: ; Among them, T final_t ∈[0,1), M is the time step length, and t∈1~M; The reward function design includes: Calculate the Pareto solution (GP) based on the objective function, and then calculate the reward function, which includes the basic reward r. base Pareto distance reward r dist and dominance intensity reward r dom Combined with basic reward r base Pareto distance reward r dist and dominance intensity reward r dom The comprehensive reward value is obtained using the following formula: r(o) t ,a t )=λ base r base +λ dist r dist +λ dom r dom ; Where, λ base , λ dist , λ dom These represent the weights of the base reward, Pareto distance reward, and dominance reward, respectively. By adding risk perception factors, penalizing high-risk decisions and rewarding low-risk decisions, defining risk indicators and incorporating them into the reward function, a risk perception-based reward mechanism is formed. The updated reward function is expressed as follows: ; Where, λ risk γ is the risk penalty coefficient. z For the weight of the z-th risk indicator, Risk z (o t a t ) represents the value of the z-th risk indicator.

2. The zero-trust driven real-time collaborative optimization and security protection method for supply chains according to claim 1, characterized in that: The formula for calculating the final trust is: T final =α·T direct +(1-a)T indirect ; Among them, T direct and T indirect Let represent direct trust and indirect trust respectively, and let α be the weight of direct trust, which is an adjustable hyperparameter.

3. The zero-trust driven real-time collaborative optimization and security protection method for supply chains according to claim 2, characterized in that: The direct trust is calculated based on multiple trust dimensions, including historical performance, compliance certification, financial health, technical capabilities, risk response, and data credibility; the weight ω of each trust dimension is determined. i Historical data for each dimension is collected and normalized. Trust scores for each dimension are calculated, and the time decay factor is considered. The scores for each dimension are then summarized to obtain the direct trust value T. direct ; Weights ω of each trust dimension i The system is dynamically adjusted based on market conditions, industry trends, and supply chain characteristics; machine learning algorithms are used to train models based on historical data and automatically adjust the weights of each dimension. Incorporating real-time data into direct trust assessments allows for timely reflection of the current status of supply chain entities.

4. The zero-trust driven real-time collaborative optimization and security protection method for supply chains according to claim 2, characterized in that: The specific steps of the indirect trust calculation include: Introducing a spatiotemporal graph temporal network: By adding a time dimension to graph convolution operations, a spatiotemporal graph is constructed to capture the trust relationships of supply chain entities that change over time; Introducing a multi-head attention mechanism: Using a multi-head attention mechanism to calculate attention weights; Node embedding update: Update node embeddings by combining the outputs of spatiotemporal graph convolution and multi-head attention mechanism; Indirect trust calculation: The output layer embeds nodes into an indirect trust value T through a fully connected network. indirect ; Define a Graph Convolutional Network (GCN) to extract indirect trust from supply entities; The output layer embeds nodes into an indirect trust value T via a fully connected network. indirect ∈[0,1].

5. The zero-trust driven real-time collaborative optimization and security protection method for supply chains according to claim 1, characterized in that: S3 specifically includes: Initialization: Initialize network parameters, Pareto solution front set, and historical solution set; Environmental interaction and experience collection: Simulate the interaction process between the reinforcement learning agent and the environment to collect new interaction data; Network Update and Policy Optimization: The policy set is optimized by updating the Critic and Actor networks. Specific steps include sampling batches from the buffer, updating the Critic and Actor networks, and updating network parameters through gradient descent and gradient ascent.

Citation Information

Patent Citations

  • A multi-objective decision-making method for complex supply chains supported by computational experiments

    CN115146455B

  • Cross-domain collaborative trust evaluation mechanism construction method and system under zero-trust system

    CN117675402A

  • KR20210037934A