Industrial chain strategy acquisition method and device, equipment and storage medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA SHENHUA ENERGY CO LTD
- Filing Date
- 2026-06-09
- Publication Date
- 2026-08-04
AI Technical Summary
[0003]本申请提供了一种产业链策略的获取方法、装置、设备及存储介质,以解决现有技术中模型无法适应动态环境、规划精准度低的问题
[0003] This application provides a method, apparatus, device, and storage medium for obtaining supply chain strategies to solve the problems of existing models being unable to adapt to dynamic environments and having low planning accuracy.
Smart Images

Figure CN122509728A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of supply chain modeling and optimization technology, specifically to a method, apparatus, equipment, and storage medium for obtaining supply chain strategies. Background Technology
[0002] In related technologies, supply chain modeling has failed to effectively address the complex environment of multiple stakeholders and highly dynamic changes within the supply chain, resulting in models that cannot accurately simulate the autonomous decision-making behavior and interaction logic among participants. Models often simplify stakeholder behavior into static rules, making planning solutions ill-suited to the dynamic needs of modern industrial systems. Summary of the Invention
[0003] This application provides a method, apparatus, device, and storage medium for obtaining supply chain strategies to solve the problems of existing models being unable to adapt to dynamic environments and having low planning accuracy.
[0004] Firstly, this application provides a method for obtaining supply chain strategies. The method includes: constructing a supply chain game model based on an intelligent agent layer, a relational layer, and a rule layer, wherein the intelligent agent layer is used to simulate the autonomous decision-making of supply chain participants, the relational layer is used to define upstream and downstream interaction logic, and the rule layer is used to integrate global constraints; obtaining an initial supply chain strategy based on the supply chain game model; adjusting the corresponding parameters in the supply chain game model in response to feedback data, and obtaining a target supply chain strategy based on the adjusted supply chain game model.
[0005] Based on the aforementioned technical means, it is possible to more accurately simulate the autonomous decision-making and complex interactions of participants in the industrial chain, and correct model deviations through real-time feedback data, thereby obtaining more adaptable industrial chain strategies that are more in line with actual operating conditions, and improving the accuracy and dynamic adaptability of industrial chain planning and optimization.
[0006] In some optional embodiments, a supply chain game model based on an agent layer, a relation layer, and a rule layer is constructed, including: constructing an agent layer, wherein the construction of the agent layer includes abstracting each supply chain entity into different agents and determining the attributes and decision rules of each agent; constructing a relation layer, wherein the construction of the relation layer includes constructing topological relationships between agents and determining information interaction rules and benefit distribution mechanisms between agents, wherein the topological relationships include at least one of material flow, capital flow, and information flow; constructing a rule layer, wherein the construction of the rule layer includes defining operational constraints of the supply chain and adjusting constraint weights through dynamic adjustment mechanisms and iterative optimization rules; and obtaining the supply chain game model based on the constructed agent layer, relation layer, and rule layer.
[0007] Based on the aforementioned technical means, the layered game theory model can provide a more refined, dynamic, and robust industrial chain analysis framework, laying a solid foundation for obtaining more effective and adaptable industrial chain strategies.
[0008] In some optional embodiments, the determination of decision rules in the agent layer includes: based on the mapping relationship between the objective function and the state space, using a preset algorithm to enable the agent to adjust its behavior according to environmental changes, wherein the preset algorithm includes reinforcement learning strategies or multi-criteria decision analysis algorithms, the behavior includes production, supply, and sales, and the environmental changes include market environment changes.
[0009] Based on the above technical means, by mapping the target to the state and combining advanced algorithms to achieve dynamic adjustment of behavior, the shortcomings of poor environmental adaptability and insufficient model robustness in the existing technology are overcome, enabling the industry chain strategy to accurately respond to the dynamic environment, thereby improving the operational efficiency and competitiveness of the entire industry chain. In some optional embodiments, the determination of information interaction rules in the relational layer includes: dividing the interaction information into a common signal layer, a business collaboration layer, and a core state layer; using intelligent routing mode and subscription and publish mechanism to transmit information at each layer of the interaction information, and controlling the perception accuracy of the intelligent agent to the external environment through view filtering rules; setting information threshold filters to fuse and compress repetitive or low-value signals to prevent overload and feedback suppression.
[0010] Based on the aforementioned technical means, the robustness and stability of the entire industry chain game model are enhanced, enabling it to more effectively acquire target industry chain strategies in complex and ever-changing market environments.
[0011] In some optional embodiments, the dynamic adjustment mechanism in the rule body layer includes: when the actual operating parameters of the agent deviate from the corresponding parameters predicted by the industry chain game model by more than a first preset threshold, correction is made through a parameter fine-tuning mechanism; when the change in the environment exceeds the corresponding second preset threshold, a decision is made through a backup strategy.
[0012] Based on the above technical means, the adaptability problem of the industrial chain game model caused by prediction deviation and drastic environmental changes in actual operation can be effectively solved.
[0013] In some optional embodiments, the iterative optimization rules in the rule body layer include: adjusting the corresponding parameters when the deviation rate between the current running data and the prediction of the industry chain game model exceeds a third threshold; and injecting random noise into the policy function or increasing the exploration rate when the industry chain game model gets stuck in a local optimum, so that the agent can jump out of the existing policy domain.
[0014] Based on the above technical means, the problem of lacking an effective automatic optimization mechanism is solved when there is a large deviation between the model prediction and the actual operating data or when the model gets stuck in a local optimum during the operation of the industrial chain game model.
[0015] In some alternative embodiments, the objective function includes: , Where, Ui,t(s t ,a i Let Φi,t(s) be the utility of the i-th agent at step t. t ,a i ,t) represents the constraint term of the i-th agent at step t, λ is the penalty factor, γ is the discount factor, and Eπheta [ ] represents the expectation under strategy πheta.
[0016] Based on the above technical means, the decision-making process was optimized, and the model's adaptability and robustness in dynamic environments were enhanced.
[0017] Secondly, this application provides a device for acquiring supply chain strategies. The device includes: a construction module for constructing a supply chain game model based on an intelligent agent layer, a relational layer, and a rule layer, wherein the intelligent agent layer is used to simulate the autonomous decision-making of supply chain participants, the relational layer is used to define upstream and downstream interaction logic, and the rule layer is used to integrate global constraints; a first acquisition module for acquiring an initial supply chain strategy based on the supply chain game model; and a second acquisition module for adjusting the corresponding parameters in the supply chain game model in response to feedback data, and acquiring a target supply chain strategy based on the adjusted supply chain game model.
[0018] Thirdly, this application provides an electronic device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to perform a method for acquiring a supply chain strategy as described in the first aspect or any corresponding embodiment.
[0019] Fourthly, this application provides a computer-readable storage medium storing computer instructions for causing a computer to execute a method for obtaining a supply chain strategy as described in the first aspect or any corresponding embodiment.
[0020] Fifthly, this application provides a computer program product, including computer instructions, which are used to cause a computer to execute a method for obtaining a supply chain strategy as described in the first aspect or any corresponding embodiment. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the specific embodiments of this application or the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0022] Figure 1 This is a schematic diagram of the first method for obtaining a supply chain strategy according to an embodiment of this application; Figure 2 A second flowchart illustrating the method for obtaining supply chain strategies according to embodiments of this application; Figure 3 A schematic diagram of the third process for obtaining the supply chain strategy according to an embodiment of this application; Figure 4 A structural block diagram of a supply chain strategy acquisition device according to an embodiment of this application; Figure 5 A schematic diagram of the hardware structure of the electronic device according to an embodiment of this application. Detailed Implementation
[0023] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0024] It is understood that before using the technical solutions disclosed in the various embodiments of this application, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this application in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0025] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.
[0026] In related technologies, supply chain modeling and optimization methods, such as integrated energy system planning optimization based on multiple factors and three levels, or wind turbine generator optimization layout methods based on three-level dynamic variation rates, have shortcomings in planning accuracy. The actual energy efficiency of equipment deviates from the model's preset values, and they fail to fully consider the dynamic changes in equipment energy efficiency during actual operation and the complex relationships between multiple factors. Furthermore, they are highly dependent on data accuracy and lack dynamic adjustment mechanisms, making it difficult to adapt to the complex and ever-changing real-world environment, resulting in insufficient model robustness.
[0027] In response, this application proposes a method for obtaining supply chain strategies, such as... Figure 1 As shown, the method includes: Step S101: Construct an industry chain game model based on the intelligent agent layer, the relation layer, and the rule layer. The intelligent agent layer is used to simulate the autonomous decision-making of industry chain participants, the relation layer is used to define the upstream and downstream interaction logic, and the rule layer is used to integrate global constraints.
[0028] Among them, the supply chain game model refers to a mathematical or computational model used to simulate and analyze the interactions, decisions, and outcomes among multiple participants in a supply chain. This model typically contains multiple layers to capture the complex dynamics and structure of the supply chain.
[0029] The intelligent agent layer refers to an abstract layer in the supply chain game model, used to abstract various entities in the supply chain, such as enterprises, departments, and individuals, as intelligent agents with autonomous decision-making capabilities. This layer simulates the behavior of intelligent agents making decisions based on their own goals and environmental information.
[0030] The relational layer, in the context of supply chain game theory, is an abstract level that defines the various interactive relationships and logics between agents. These relationships can include material flows, capital flows, and information flows, and specify how agents exchange information and distribute benefits.
[0031] The rule layer, in the context of a supply chain game theory model, is an abstract layer used to integrate and manage the global constraints and rules governing the operation of the supply chain. This layer ensures that the decisions and interactions of agents comply with relevant laws, regulations, industry standards, and resource limitations.
[0032] Industry chain participants refer to various entities that play specific roles and conduct economic activities in the industry chain, such as raw material suppliers, component manufacturers, product assemblers, distributors, retailers, and end consumers.
[0033] Upstream and downstream interaction logic refers to the rules and processes by which participants at different stages of the industrial chain exchange resources, information, products, or services. Upstream typically refers to the stage that provides raw materials or semi-finished products, while downstream refers to the stage that receives and further processes or sells them.
[0034] Global constraints refer to the restrictions imposed on the entire industry chain system. These constraints may originate from external environments, such as policies and regulations, market capacity, or internal resources, such as total production capacity and total budget. All intelligent agents must act within these constraints.
[0035] Specifically, this method first constructs a supply chain game model based on three layers: agent layer, relation layer, and rule layer. When constructing the agent layer, each entity in the supply chain, such as suppliers, manufacturers, and distributors, can be abstracted as an independent agent. The autonomous decision-making behavior of these agents can be simulated using a pre-set decision tree or a simple rule-based expert system. For example, a production agent can be programmed to automatically send a purchase request to its upstream supplier agent when its inventory falls below a preset threshold; when market demand increases, the production agent adjusts its output according to a pre-set production plan. This approach gives the agents a degree of autonomy, but their decision-making logic remains relatively fixed.
[0036] When constructing the relational layer, the upstream and downstream interaction logic between intelligent agents can be implemented by defining fixed data interfaces and communication protocols. For example, after an upstream intelligent agent completes product delivery, it transmits information such as product quantity and quality to a downstream intelligent agent via a standardized data packet. The transmission paths and rules of material flow, capital flow, and information flow can be explicitly defined, forming an interactive network. For example, the capital flow can be set so that after product delivery, the downstream intelligent agent settles accounts with its upstream intelligent agent according to a preset payment cycle.
[0037] When constructing the rule layer, global constraints of the industry chain, such as total production capacity limits, market access standards, and environmental regulations, can be integrated into the model as static parameters. These constraints remain unchanged during model operation, and all agent decisions and interactions must be conducted within these pre-defined static constraints. For example, the total output of the entire industry chain cannot exceed the set market capacity limit, or the emissions of any agent cannot exceed the national environmental standards.
[0038] Step S102: Obtain the initial industrial chain strategy based on the industrial chain game model.
[0039] The initial supply chain strategy refers to a preliminary, unoptimized supply chain operation plan or decision set obtained through model operation or calculation when the supply chain game model is first constructed or under specific initial conditions.
[0040] Specifically, this is typically achieved by running a one-time simulation of the model. For example, given initial market conditions and agent states, the model runs for a pre-set simulation period, recording the decision sequences of each agent and the overall operational results of the industry chain, such as total output, profit distribution, and resource consumption. The collection of these operational results constitutes the initial industry chain strategy, providing a foundation for subsequent analysis and optimization.
[0041] Step S103: In response to the feedback data, adjust the corresponding parameters in the industrial chain game model, and obtain the target industrial chain strategy based on the adjusted industrial chain game model.
[0042] Feedback data refers to data collected from the external environment or internal systems during the actual operation or simulation of the supply chain game model, used to evaluate model performance or guide model adjustments. This data reflects the discrepancies between actual conditions and model predictions.
[0043] The target supply chain strategy refers to a final set of supply chain operation plans or decision sets that are obtained after the supply chain game model has been adjusted and optimized based on feedback data, and are aimed at achieving a specific optimization goal.
[0044] Specifically, feedback data can come from various indicators monitored during actual operation, such as actual market demand, product price fluctuations, and the deviation between the agent's actual output and the model's predictions. When feedback data indicates a discrepancy between the model's predictions and the actual situation, the relevant parameters in the model can be corrected through manual intervention or a simple scaling mechanism. For example, if actual market demand consistently exceeds the model's predictions, the weight or value of the market demand parameter in the model can be increased manually or through a simple algorithm.
[0045] After adjusting the model parameters based on feedback data, the entire industry chain game model is run again. By resimulating or recalculating, a new industry chain operation plan that better reflects the current situation can be obtained. This new plan is the target industry chain strategy, which reflects the optimization result of the model after absorbing the latest feedback information, aiming to improve the adaptability and effectiveness of the strategy.
[0046] It is understandable that this application addresses the shortcomings of existing industry chain modeling methods in terms of accuracy, lack of dynamic adjustment mechanisms, and low robustness in environments with complex multi-agent characteristics by constructing a multi-level industry chain game model and introducing a feedback adjustment mechanism. This method can more accurately simulate the autonomous decision-making and complex interactions of industry chain participants, and correct model biases through real-time feedback data, thereby obtaining more adaptable industry chain strategies that better reflect actual operating conditions, and improving the accuracy and dynamic adaptability of industry chain planning and optimization.
[0047] In some of the solutions mentioned above in this application, a supply chain game model is proposed to simulate the autonomous decision-making, upstream and downstream interaction logic, and global constraints of supply chain participants. However, in the process of its implementation, the existing methods have problems such as insufficient model accuracy and lack of dynamic description ability of complex relationships. Specifically, it is difficult to accurately abstract entity behavior, define interaction rules, and adjust constraint weights in real time, which leads to the model being out of touch with the actual operating environment and unable to effectively cope with market changes and data deviations.
[0048] In response, this application further proposes a method for obtaining supply chain strategies, such as... Figure 2 As shown, the method includes: Step S201: Construct a supply chain game model based on the intelligent agent layer, relation layer, and rule layer. Specifically, this includes: Step S2011: Construct the intelligent agent layer. The construction of the intelligent agent layer includes abstracting each industry chain entity into different intelligent agents and determining the attributes and decision rules of each intelligent agent.
[0049] Specifically, when constructing the intelligent agent layer, various entities in the industry chain are abstracted into different intelligent agents. This helps to decompose complex real-world systems into manageable and analyzable units. Each intelligent agent represents an independent decision-making entity capable of simulating its behavior within the industry chain. Besides abstracting enterprises or organizations in the industry chain, such as suppliers, manufacturers, distributors, retailers, and consumers, into intelligent agents, specific functional departments, such as purchasing departments and sales departments, or key resources, such as specific production lines and logistics centers, can also be abstracted into intelligent agents to simulate their behavior and interactions with greater granularity. Simultaneously, the attributes and decision-making rules of each intelligent agent are determined. Attributes refer to the inherent characteristics of the intelligent agent, such as production capacity, inventory levels, cost structure, market share, and risk preference. Decision-making rules refer to the logic by which the intelligent agent makes choices in a specific environment, such as adjusting output based on market demand, optimizing procurement strategies based on cost, or adjusting pricing based on competitor behavior. These rules can be pre-defined heuristic rules, strategies based on optimization algorithms, or behavioral patterns learned from machine learning models. In addition to reinforcement learning strategies or multi-criteria decision analysis algorithms, decision rules can also be determined through expert systems, rule-based inference engines, or bounded rationality models based on behavioral economics, in order to simulate the decision-making process of intelligent agents in different situations.
[0050] Step S2012: Construct a relational layer. The construction of the relational layer includes constructing topological relationships between intelligent agents and determining information interaction rules and benefit distribution mechanisms between intelligent agents. The topological relationships include at least one of material flow, capital flow, and information flow.
[0051] Specifically, when constructing the relational layer, topological relationships between intelligent agents are established. These topological relationships describe the connection methods and structures between intelligent agents within the industry chain. Material flow refers to the physical movement of raw materials, semi-finished products, and finished products between intelligent agents; capital flow refers to financial transactions such as payments, settlements, and financing; and information flow refers to the transmission of data such as orders, demand forecasts, inventory data, and market intelligence. These flows collectively constitute the operational framework of the industry chain, clarifying the dependencies and interaction paths between intelligent agents. In addition to material, capital, and information flows, topological relationships can also include technology flows, such as technology licensing and R&D cooperation; service flows, such as after-sales service and consulting services; or human resource flows, such as talent mobility and skills sharing, to more comprehensively reflect the complex connections within the industry chain. Simultaneously, the rules for information interaction and the mechanism for profit distribution between intelligent agents are determined.
[0052] Information exchange rules define how agents exchange data and instructions, including the frequency, format, content, and communication protocols of the interactions. Profit-sharing mechanisms, on the other hand, stipulate how the overall profits of the industry chain are distributed among the agents. These rules and mechanisms are crucial for coordinating agent behavior and achieving overall industry chain optimization.
[0053] In one example, the profit-sharing mechanism can be implemented as follows: To transform complex business interest relationships into computable reward signals for agents, a comprehensive reward function R for agent i at time t is defined. i,t Its expression is: , where π i,t For direct profit to individuals; For the allocation of collaborative benefits; P i,t α represents the penalty for breach of contract and risk; α and β are the benefit balance coefficients for dynamic adjustment of the rule body.
[0054] The specific quantification of positive interest linkages includes using a Shapley value-based allocation algorithm to quantify the industry value of individuals. The rule body calculates the global excess profit generated within the complete industry chain set N. When the addition of an agent i improves the overall operational efficiency of the industry chain and reduces inventory costs, the system calculates its marginal contribution and, through a profit-sharing contract, such as proportionally increasing its settlement price or providing a collaborative bonus, ensures that the agent's decisions not only consider itself but also spontaneously move towards the globally optimal direction.
[0055] The specific quantification of negative interest associations includes introducing a moral hazard quantification mechanism from the principal-agent model. If an upstream agent defaults, such as with delayed delivery or substandard quality, causing downstream agents to suffer losses due to work stoppages or high emergency procurement costs, the system will quantify this consequential loss P. i,t The defaulting party's expected profits are directly deducted through automatically executed virtual smart contracts.
[0056] The comprehensive return R obtained from the above quantitative calculation i,t This directly serves as the feedback reward signal for the reinforcement learning neural network. The closer the interest correlation, i.e., the larger α and β are, the easier it is for the agent to converge to the globally co-optimal policy during iterative training, rather than getting trapped in local optima.
[0057] Step S2013: Construct the rule body layer. The construction of the rule body layer includes defining the operational constraints of the industry chain and adjusting the constraint weights through dynamic adjustment mechanisms and iterative optimization of rules.
[0058] Specifically, when constructing the rule layer, operational constraints of the industrial chain are defined. Operational constraints refer to the limitations that the industrial chain must adhere to during its operation, such as production capacity limits, inventory capacity limits, transportation time limits, regulatory and policy restrictions, environmental emission standards, and market access conditions. These constraints ensure the legal, compliant, and sustainable operation of the industrial chain and serve as the boundary conditions for agent decisions in the game theory model. Besides physical resource constraints and regulations, operational constraints can also include market demand fluctuation ranges, supply chain disruption risk thresholds, social responsibility standards, or internal strategic goals. Simultaneously, constraint weights are adjusted through dynamic adjustment mechanisms and iterative optimization rules. Dynamic adjustment mechanisms refer to methods that can adjust constraint parameters or their importance in the model in real-time or near real-time when the actual operating state of the industrial chain or the external environment changes. Iterative optimization rules are strategies that continuously evaluate the effectiveness of strategies and correct model parameters during model operation to gradually approach the optimal solution. These mechanisms and rules enable the model to adapt to uncertainty and dynamism, maintaining its effectiveness and accuracy. In addition to parameter fine-tuning mechanisms, backup strategies, injecting random noise, or increasing the exploration rate, dynamic adjustment mechanisms can also include adaptive control based on real-time data streams, early warning and adjustment based on predictive models, and iterative optimization rules can adopt metaheuristic algorithms such as genetic algorithms, particle swarm optimization, and simulated annealing, or optimization methods based on gradient descent.
[0059] Step S2014: Based on the constructed intelligent agent layer, relation layer, and rule layer, the industry chain game model is obtained.
[0060] Specifically, based on the constructed agent layer, relation layer, and rule layer, a supply chain game model is obtained. This game model can be a multi-agent system, an agent-based simulation model, a mixed integer programming model, or a hybrid model combining machine learning and optimization algorithms. Its core lies in its ability to simulate and predict the behavior and outcomes of the supply chain under different strategies.
[0061] Step S202: Obtain the initial supply chain strategy based on the supply chain game model. See details in [link to relevant documentation]. Figure 1Step S102 in the embodiment will not be described again here.
[0062] Step S203: In response to the feedback data, adjust the relevant parameters in the supply chain game model, and obtain the target supply chain strategy based on the adjusted supply chain game model. See details... Figure 1 Step S103 in the embodiment will not be described again here.
[0063] It is understood that, through the above-described technical solution in this embodiment, this application effectively solves the problems of insufficient model accuracy and lack of dynamic adjustment capability in the prior art by constructing a layered industrial chain game model. Specifically, the construction of the agent layer, by abstracting the entities in the industrial chain into agents with clear attributes and decision-making rules, enables the model to simulate the autonomous behavior and heterogeneity of individual participants more realistically and precisely, thereby improving the abstraction accuracy of complex entity behavior. The construction of the relationship layer, by clarifying the topological relationships between agents, as well as information interaction rules and benefit distribution mechanisms, clearly depicts the interaction patterns and benefit-driven factors among the participants in the industrial chain, enhancing the model's ability to dynamically describe complex upstream and downstream relationships. The construction of the rule layer, by defining the operational constraints of the industrial chain and introducing dynamic adjustment mechanisms and iterative optimization rules, enables the model to adaptively correct constraints and optimization strategies based on actual operating data and environmental changes, thereby overcoming the problems of the model being out of touch with the actual environment and unable to effectively cope with market changes and data biases. Overall, the layered game theory model provides a more refined, dynamic, and robust framework for analyzing the industrial chain, laying a solid foundation for obtaining more effective and adaptable industrial chain strategies. In some of the solutions mentioned above in this application, decision-making rules are proposed to construct the agent layer. However, in this process, there is a lack of an effective mechanism to enable the agent to dynamically adjust its behavior according to environmental changes. This results in poor model robustness in complex and ever-changing industrial chain environments, making it difficult to adapt to actual market changes. Consequently, the model predictions deviate too much from the actual operation, and it is unable to accurately respond to dynamic environments.
[0064] In response, this application further proposes the determination of decision rules in the agent layer, including: based on the mapping relationship between the objective function and the state space, a preset algorithm is adopted to enable the agent to adjust its behavior according to environmental changes. The preset algorithm includes reinforcement learning strategies or multi-criteria decision analysis algorithms, the behavior includes production, supply and sales, and the environmental changes include market environment changes.
[0065] Specifically, based on the mapping relationship between the objective function and the state space, the aim is to associate the agent's optimization objective with the current environmental state. The objective function defines the utility or benefit the agent expects to achieve by taking a specific action in a given state, such as profit maximization, cost minimization, or market share increase. The state space encompasses all environmental information perceptible to the agent, such as market prices, inventory levels, demand fluctuations, and competitor behavior. By establishing this mapping relationship, the agent can assess the impact of different actions on achieving its objective based on the real-time state of its environment, thus providing a quantitative basis for decision-making. For example, a utility function can be defined, taking current inventory levels, market demand, and production costs as inputs and outputting the expected benefit of the current production decision; or, within a reinforcement learning framework, the objective function can be transformed into a reward function, and the state space can be defined as all possible scenarios of the agent's interaction with the environment, optimizing decisions by learning the value of state-action pairs.
[0066] To enable the agent to make effective decisions, this application employs a pre-defined algorithm. This pre-defined algorithm may include reinforcement learning strategies or multi-criteria decision analysis algorithms. When using a reinforcement learning strategy, the agent learns how to optimize its behavior from experience through continuous interaction with the environment. For example, a Q-learning algorithm can be used, where the agent updates its Q-value table through trial and error and receiving reward signals, thereby learning which behavior in different states will yield the maximum cumulative reward; alternatively, a policy gradient algorithm can be used, where the agent directly learns a policy function that maps the current state to the probability distribution of the optimal behavior. When using a multi-criteria decision analysis algorithm, the agent can weigh and choose among multiple conflicting decision criteria. For example, the Analytic Hierarchy Process (AHP) can be used to decompose complex decision problems into multiple levels, determine the weights of each criterion by comparing judgment matrices, and then rank the candidate behaviors to determine the optimal behavior.
[0067] Even if the agent adjusts its behavior according to environmental changes, it no longer follows fixed rules but can adjust its production, supply, and sales behaviors in real time based on perceived environmental changes. For example, when market demand suddenly increases, the agent can immediately adjust its production plan to increase output; when raw material prices rise, the agent can seek alternative suppliers or adjust its procurement strategy. This dynamic adjustment capability ensures the flexibility and responsiveness of the supply chain strategy. "Production" behavior refers to the agent's decisions in the production stage, such as formulating production plans, allocating production resources, and selecting production processes. "Supply" behavior refers to the agent's decisions in the supply chain stage, such as purchasing raw materials, selecting and managing suppliers, formulating inventory strategies, and optimizing logistics and distribution. "Sales" behavior refers to the agent's decisions in the sales stage, such as product pricing, selecting sales channels, formulating marketing strategies, and maintaining customer relationships. These behaviors cover the core aspects of supply chain operation, ensuring the comprehensiveness and practicality of the agent's decisions. Furthermore, environmental changes include changes in the market environment, which are key external factors affecting supply chain operation. This includes, but is not limited to, changes in consumer preferences, competitors' strategic adjustments, changes in macroeconomic policies, market shocks brought about by technological innovation, and seasonal demand fluctuations. By incorporating these market environment changes into the perception scope of intelligent agents and using them as the basis for their decision-making adjustments, supply chain strategies can be made more closely aligned with actual market demands, thereby enhancing their market competitiveness.
[0068] It is understood that, through the above-described technical solutions in this embodiment, this application effectively solves the problem of insufficient adaptability of intelligent agents under dynamic environmental changes, and improves the robustness and accuracy of the industrial chain game model. Specifically, based on the mapping relationship between the objective function and the state space, the intelligent agent can closely associate its optimization objective with the real-time environmental state, ensuring that the decision-making process not only relies on preset rules but can also be adjusted according to the dynamic changes in the current market environment, thereby avoiding the defect of static models failing in complex and ever-changing environments. By adopting preset algorithms such as reinforcement learning strategies or multi-criteria decision analysis algorithms, the intelligent agent is endowed with the ability to self-learn and make multi-dimensional decisions, enabling it to learn from historical data and optimize its core behaviors such as production, supply, and sales, greatly enhancing the flexibility and adaptability of decision-making. Overall, by mapping the objective to the state and combining advanced algorithms to achieve dynamic adjustment of behavior, the shortcomings of poor environmental adaptability and insufficient model robustness in the prior art are overcome, enabling the industrial chain strategy to accurately respond to the dynamic environment, thereby improving the operational efficiency and competitiveness of the entire industrial chain. In some of the embodiments described above in this application, information interaction rules are proposed to define information interaction between intelligent agents. However, in the process of implementation, information transmission may lack effective hierarchical management, intelligent routing mechanisms, and filtering of low-value signals, resulting in information overload, feedback suppression, and insufficient perception accuracy, which affects the dynamic adaptability and robustness of the model.
[0069] In response, this application further proposes that the determination of information interaction rules in the relational layer includes: dividing the interaction information into a common signal layer, a business collaboration layer, and a core state layer; using intelligent routing mode and subscription and publication mechanism to transmit information at each layer of the interaction information, and controlling the perception accuracy of the intelligent agent to the external environment through view filtering rules; setting information threshold filters to fuse and compress repetitive or low-value signals to prevent overload and feedback suppression.
[0070] Specifically, the interactive information is divided into a common signal layer, a business collaboration layer, and a core state layer to achieve refined information management. The common signal layer carries general, broadcast information such as market dynamics and policies / regulations, for reference by all or most agents. The business collaboration layer focuses on processing interactive data generated between agents to complete specific business processes, such as order processing, inventory allocation, and production collaboration. The core state layer is used to transmit key state variables that have a decisive impact on agent decision-making, such as key resource inventory, core equipment operating status, and key performance indicators. The layered processing method can be defined based on the universality, timeliness, importance of the information content, or the role and permissions of the receiving agent, or it can be automatically categorized through preset information type tags or metadata.
[0071] This system employs intelligent routing and a publish-subscribe mechanism to optimize information transmission efficiency and accuracy at each layer of the interactive information exchange. The intelligent routing model dynamically selects the optimal transmission path based on factors such as network topology, the strength of business relationships between agents, and information priority. For example, it can be implemented using routing algorithms based on proxies, content awareness, or Quality of Service (QoS). The publish-subscribe mechanism allows agents to actively subscribe to information topics or events of interest based on their own needs. When relevant information is published, the system automatically pushes it to subscribers, thus avoiding unnecessary network-wide broadcasts and reducing redundant transmission. This mechanism can be implemented through message queue systems, event buses, or distributed publish / subscribe frameworks.
[0072] View filtering rules control the precision of an agent's perception of the external environment, aiming to prevent the agent from receiving too much irrelevant information and thus overburdening its decision-making. View filtering rules can dynamically adjust the granularity of information received by the agent based on its current task, role permissions, computing power limitations, or preset focus. For example, an agent focused on production planning might only need to receive information on raw material inventory and order changes, without needing to pay attention to marketing data. These rules can be implemented through configuration parameters, rule-based engine strategies, or lightweight machine learning models, ensuring that the agent can focus on the information most critical to its decision-making.
[0073] An information threshold filter is implemented to fuse and compress repetitive or low-value signals, preventing overload and feedback suppression, thereby improving system stability and robustness. The information threshold filter identifies repetitive or low-value signals with minimal impact on decision-making based on preset criteria such as information change rate, importance score, novelty, or business rules. Once identified, these signals are processed using fusion and compression techniques, such as aggregating continuous small changes, deduplicating identical information, or extracting features from large amounts of similar data. This effectively reduces the amount of information transmitted and the processing burden, preventing system overload due to information overload and avoiding decision oscillations or suppression caused by invalid feedback.
[0074] In one example, message delivery can be based on an intent-driven intelligent routing pattern: The agent dynamically subscribes to messages from upstream supply and downstream demand nodes, avoiding computational overload caused by irrelevant and redundant data. The system generates an independent observation view vector for each agent. , of which M i This is the permission mask for the agent. The mask determines the agent's permissions to the external environment S. env The accuracy of perception.
[0075] For example, downstream power plant agents can only observe the supplier's price range, but cannot know their remaining inventory. This technical reduction of information forces the agents to engage in game theory under incomplete information, improving the model's realism.
[0076] An information threshold filter is set up so that when the interaction frequency exceeds the set decision step length period in a short period of time, repetitive or low-value signals will be fused and compressed to ensure that the input dimension of the agent's decision neural network remains stable and to prevent model oscillation.
[0077] It is understood that, through the above-described technical solutions in this embodiment, this application effectively solves the problems of low efficiency, overload risk, and insufficient perception accuracy in information interaction, significantly improving the dynamic adaptability and stability of the industry chain game model. Specifically, hierarchical management of interactive information allows different types of information to be processed and transmitted in a targeted manner, avoiding information mixing and disordered transmission, thereby improving the efficiency and accuracy of information transmission. The introduction of intelligent routing mode and subscription-publishing mechanism ensures that information can be accurately delivered to the truly needed intelligent agent via the optimal path, reducing network load and redundant communication, enabling the intelligent agent to obtain key information and respond in a timely manner. At the same time, the application of view filtering rules enables the intelligent agent to dynamically adjust the perception granularity of the external environment according to its own needs and task focus, effectively avoiding the interference of information overload on the decision-making process, ensuring that the intelligent agent can focus on core decisions, and improving the accuracy of decision-making. Furthermore, the setting of the information threshold filter, by fusing and compressing repetitive or low-value signals, further reduces the information processing burden of the system, effectively preventing system overload and feedback inhibition, thereby enhancing the robustness and stability of the entire industry chain game model, enabling it to more effectively acquire target industry chain strategies in a complex and ever-changing market environment. In some of the solutions mentioned above in this application, a rule body layer is proposed to integrate global constraints and adjust constraint weights through dynamic adjustment mechanisms and iterative optimization rules. However, in this process, when the actual operating parameters of the agent deviate significantly from the model prediction or when the environment changes drastically, there is a lack of effective threshold triggering mechanisms and targeted adjustment strategies, which makes it impossible for the model to correct deviations or adapt to sudden changes in a timely manner, thereby affecting the accuracy of decision-making and the robustness of the system.
[0078] In this regard, this application further proposes that the dynamic adjustment mechanism in the rule body layer includes: when the actual operating parameters of the agent are detected to deviate from the corresponding parameters predicted by the industry chain game model by more than a first preset threshold, correction is made through a parameter fine-tuning mechanism; when the change in the environment exceeds the corresponding second preset threshold, a decision is made through a backup strategy.
[0079] Specifically, the dynamic adjustment mechanism aims to ensure that the supply chain game model can promptly and effectively self-correct and adapt to actual operational deviations and changes in the external environment, thereby maintaining the accuracy of decision-making and the stability of the system. Specifically, when the deviation between the agent's actual operating parameters and the corresponding parameters predicted by the supply chain game model exceeds a first preset threshold, the system will activate the parameter fine-tuning mechanism. The agent's actual operating parameters may include its output, input costs, inventory levels, market share, etc., while the corresponding parameters predicted by the supply chain game model are the expected values deduced by the model based on the current state and strategy. The first preset threshold is a quantitative standard for judging whether the deviation is "significant." The system can continuously collect the agent's actual operating data and compare it with the model's real-time output prediction data, calculating the absolute or relative deviation between the two, and then comparing it with the preset first threshold. When the deviation exceeds this threshold, the parameter fine-tuning mechanism is triggered. The parameter fine-tuning mechanism refers to a method of making small, gradual adjustments to some parameters in the supply chain game model when a deviation between the model prediction and actual operation is detected. Parameters may include weights in the agent's decision rules, interaction coefficients in the relation layer, and constraint strengths in the rule layer. For example, gradient descent or heuristic search algorithms can be used to iteratively adjust key parameters affecting the deviation in small steps based on the magnitude and direction of the deviation; or a set of adjustable parameters and their adjustment step sizes can be preset, and when fine-tuning is triggered, the corresponding parameters can be adjusted incrementally or decrementally according to predefined rules or lookup tables.
[0080] Furthermore, when the magnitude of environmental change exceeds a corresponding second preset threshold, the system will make decisions using backup strategies. Environmental changes can include market demand fluctuations, drastic changes in raw material prices, policy and regulatory adjustments, technological breakthroughs, etc. The second preset threshold is a quantitative standard for measuring the "severity" of environmental change. The system can continuously monitor key environmental indicators, such as market indices, macroeconomic data, and policy releases, and calculate the rate or magnitude of change of these indicators within a certain time window, comparing it with the preset second preset threshold. When the magnitude of change exceeds this threshold, the backup strategy is activated. Backup strategies refer to a set of alternative decision-making schemes prepared in advance when drastic changes in the external environment render the original industrial chain strategy inapplicable or potentially have negative impacts. These strategies are usually designed for specific extreme situations or significant risks. For example, multiple backup strategies for different environmental change scenarios can be pre-designed and stored. When the triggering conditions are met, the system selects and executes the most suitable backup strategy based on the specific type of current environmental change; or a rule-based expert system can be used, where the system automatically derives and executes the corresponding backup decision based on a predefined rule base when environmental changes exceed the threshold.
[0081] In one example, the dynamic adjustment mechanism may include: The agent collects its own attribute features and external environment features in real time and encodes them into a high-dimensional state feature vector S. t .
[0082] The agent internally integrates a policy function π(a|s;θ) based on a deep neural network. Using weight parameters θ obtained through training on historical game data, the agent calculates the policy function π(a|s;θ) for the current state S. t The probability distribution of taking different business actions, such as increasing production, decreasing production, raising prices, and lowering prices.
[0083] The decision-making logic is centered on balancing short-term and long-term gains, aiming to maximize the expected cumulative reward. , where r t Given the current step's reward, γ is the discount factor, and the final optimal action instruction a is output. t .
[0084] To ensure the model accurately maps to the actual industry ecosystem, the input dimensions are defined using the following example: For a production-type intelligent agent, such as a chemical plant, its attributes include real-time inventory rate (reflecting warehousing pressure), marginal production cost (determining the profit floor), equipment health (affecting production capacity reliability), and historical order fulfillment rate (affecting reputation score in the game).
[0085] External environment, taking the energy industry chain environment as an example, includes environmental parameters such as spot market commodity prices (fluctuation guidance strategy), meteorological factors (such as low temperatures leading to a surge in electricity consumption or heavy rain affecting logistics), policy restriction indicators (such as carbon emission quotas and power curtailment ratios), and publicly quoted signals from competitors.
[0086] When the external environment undergoes a sudden change, such as a surge in market prices, the agent updates its state vector s. t+1 This triggers the recalculation of the strategy function, thereby shifting the decision-making action from a conservative supply dynamic to an active capacity expansion, achieving real-time alignment between decision-making behavior and the industrial environment.
[0087] It is understood that, through the above-described technical solutions in this embodiment, this application can effectively solve the adaptability problem of the industrial chain game model caused by prediction deviations and drastic environmental changes in actual operation. In some of the schemes mentioned above in this application, a rule body layer is proposed to integrate global constraints and dynamic adjustment mechanisms. However, in this process, when the model prediction bias rate is too large or the model gets stuck in a local optimum, there is a lack of an effective automatic optimization mechanism to ensure the model's continuous accuracy and global optimization, resulting in poor model robustness and difficulty in adapting to complex and ever-changing real-world environments.
[0088] In response, this application further proposes iterative optimization rules in the rule body layer, which include: adjusting the corresponding parameters when the deviation rate between the current running data and the prediction of the industrial chain game model exceeds the third threshold; and injecting random noise into the policy function or increasing the exploration rate when the industrial chain game model gets stuck in a local optimum, so that the agent can jump out of the existing policy domain.
[0089] Specifically, this mechanism aims to ensure that the supply chain game model remains synchronized with actual operations. Current operational data refers to various indicators collected in real time during the actual operation of the supply chain, such as actual output, actual inventory, actual market prices, and actual logistics costs. The corresponding parameters predicted by the supply chain game model are estimates of these indicators for the future or present based on the model's internal logic and current state. The deviation rate is an indicator that measures the degree of difference between the predicted and actual values, and can be calculated in various ways, such as relative error, mean squared error, or mean absolute error. The third threshold is a preset tolerance level; when the deviation rate exceeds this threshold, it indicates a significant deviation between the model's predictions and the actual situation, requiring intervention. Adjusting the corresponding parameters means correcting the parameters within the supply chain game model that affect the prediction results based on the magnitude and direction of the deviation. For example, a gradient descent-based optimization algorithm can be used to iteratively update the weight coefficients, probability distribution parameters, or thresholds in the decision rules based on the gradient direction of the prediction error. Another approach is to pre-set a series of adaptive adjustment rules. When the deviation rate of a specific type exceeds a threshold, a predefined parameter correction action is triggered. For example, if the market demand forecast is consistently lower than the actual demand, the growth factor in the market demand forecast model can be slightly increased.
[0090] Furthermore, a local optimum refers to a situation where, within the current policy space, the agent cannot find a better policy, but a global optimum may exist in other unexplored policy spaces. This mechanism aims to help the model escape suboptimal states and explore a broader policy space to seek a global optimum. Injecting random noise into the policy function refers to introducing a certain degree of random perturbation when the agent makes decisions based on its policy function. For example, in reinforcement learning-based agents, Gaussian or uniform noise can be added to its output action values, such as Q-values, or policy probabilities, so that when the agent performs what it considers the "optimal" action, there is a certain probability that it will choose other non-optimal actions. This randomness helps the agent try new behaviors, thereby discovering potentially better policies. Increasing the exploration rate refers to increasing the agent's tendency to conduct random exploration during the decision-making process. For example, in optimization algorithms, such as genetic algorithms, the mutation rate can be increased; in particle swarm optimization, the random component of particle movement can be adjusted. To determine whether a model is trapped in a local optimum, you can monitor model performance metrics, such as the objective function value, whether it has not improved significantly over a long period of time in continuous iterations, or whether the magnitude of policy updates tends to zero.
[0091] It is understood that, through the above-mentioned technical solution in this embodiment, this application solves the problem of lacking an effective automatic optimization mechanism when the model prediction deviates significantly from the actual operating data or the model gets trapped in a local optimum during the operation of the industrial chain game model. In some of the schemes mentioned above in this application, an objective function is proposed to enable the agent to adjust its behavior based on environmental changes. However, in this process, the objective function fails to effectively integrate constraints and handle dynamic uncertainties, resulting in insufficient decision optimization, poor model robustness, and difficulty in adapting to complex and ever-changing environments.
[0092] In response, this application further proposes an objective function, including: , among which, U i,t (s t ,a i,t Let Φ be the utility of the i-th agent at step t. i,t (s t ,a i,t Let ) be the constraint term for the i-th agent at step t, λ be the penalty factor, γ be the discount factor, and E be the value of E. πheta [ ] represents the expectation under strategy πheta.
[0093] Specifically, U i,t (s t ,a i,tLet $\frac{i}{t}$ represent the utility of the i-th agent at step $t$. It is a quantitative indicator measuring the gain or satisfaction gained by the agent from taking a specific action $a_i,t$ in a specific state $st$, reflecting the agent's goal orientation in the decision-making process. Utility can be calculated based on economic indicators; for example, for a production agent, its utility may be related to production volume, selling price, and production cost. Furthermore, utility can also be evaluated based on non-economic indicators; for example, for agents in a supply chain, its utility may include on-time delivery rate, customer satisfaction, etc.
[0094] Φ i,t (s t ,a i,t Let represent the constraint term of the i-th agent at step t. This represents various restrictions that the agent must adhere to during the decision-making process. These restrictions can be resource limitations, policies and regulations, technical standards, or market rules. Constraint terms can be expressed as inequalities or equations, such as upper limits on production capacity, lower limits on raw material inventory, and liquidity requirements. When the agent's behavior violates these constraints, the constraint term will have a non-zero value. Alternatively, constraints can also be logical conditions, such as certain investments only possible under specific market conditions, or shipments only possible after meeting specific quality standards.
[0095] λ is the penalty factor, a positive weighting coefficient used to quantify the negative impact or cost of violating constraints. It "penalizes" behaviors that violate constraints by increasing the value of the objective function, thereby guiding the agent to make compliant decisions. The penalty factor can be a fixed constant, pre-set according to the importance of the constraint.
[0096] γ is the discount factor, a value between 0 and 1, used to measure the value of future returns relative to current returns. It reflects the agent's consideration of future uncertainty, causing the agent to favor immediate returns or discount future returns when making decisions. The discount factor can also be adjusted according to the dynamics of the environment or the agent's risk appetite. For example, in a highly uncertain market environment, the discount factor can be set lower to reduce reliance on long-term returns.
[0097] E πheta [ Let be the expectation under policy πheta, which represents the weighted average of all possible outcomes of the objective function given policy πheta. It is used to handle randomness and uncertainty in the environment, ensuring that the agent's decisions are based on long-term average performance rather than a single random event.
[0098] It is understood that, through the above-described technical solution in this embodiment, this application optimizes the decision-making process by defining a specific mathematical objective function, thereby enhancing the model's adaptability and robustness in dynamic environments. This enables the objective function to more comprehensively handle constraints and uncertainties, thus achieving efficient decision optimization in the supply chain game model and solving the problems of insufficient decision optimization and poor model robustness in existing technologies. In one example, the following provides a more detailed explanation of the above technical solution through a more specific example: A complex electronic component manufacturing supply chain involves multiple participants, including upstream raw material suppliers, midstream component manufacturers, downstream product assembly plants, and final sales channels. This supply chain faces challenges such as volatile market demand, difficulty in accurately matching production plans, inventory backlogs or shortages, and low efficiency in information collaboration across different stages. Existing technologies often struggle to effectively address such complex, multi-participant, and highly dynamic environments due to a lack of innovative models, insufficient planning precision, and a lack of dynamic adjustment mechanisms.
[0099] This solution first constructs a supply chain game model based on the intelligent agent layer, relation layer, and rule layer. For example... Figure 3 As shown.
[0100] In constructing the agent layer, each independent entity in the supply chain, such as raw material suppliers, component manufacturers, product assembly plants, and sales channels, is abstracted as a different agent, i.e., the agent layer. Each agent is assigned specific attributes, such as production capacity, inventory level, cost structure, and delivery cycle. Simultaneously, decision rules are determined for each agent. For example, the decision rule for the component manufacturer agent is based on the mapping relationship between the objective function and the state space, and is adjusted using a reinforcement learning strategy.
[0101] Next, the relationship layer is constructed. This layer defines the topological relationships between agents, including at least material flow, capital flow, and information flow. For example, raw materials flow from suppliers to manufacturers (material flow), manufacturers pay suppliers (capital flow), and manufacturers send production plans to assembly plants (information flow). Simultaneously, the information interaction rules and benefit distribution mechanisms between agents are determined. Regarding information interaction rules, interactive information is divided into a common signal layer (e.g., industry macroeconomic data), a business collaboration layer (e.g., order, inventory, and delivery status), and a core status layer (e.g., real-time data from internal production lines). The system uses intelligent routing and a publish-subscribe mechanism to transmit information between these layers, ensuring that information reaches the relevant agents efficiently and accurately. For example, when a product assembly plant publishes a new order request, this information is transmitted to the component manufacturer via intelligent routing, and the manufacturer agent subscribes to this request information. View filtering rules control the accuracy of the agent's perception of the external environment; for example, the manufacturer agent may only receive aggregated order requests, rather than detailed customer information for each sales channel. Furthermore, an information threshold filter is implemented to fuse and compress repetitive or low-value signals. For example, multiple low-inventory warning signals can be integrated into a single emergency replenishment request to prevent information overload and feedback suppression, thereby improving the efficiency and quality of information transmission and addressing the inaccuracy of information transmission in existing models. The profit-sharing mechanism, based on the contributions and risks of each agent, sets a reasonable profit-sharing or cost-sharing scheme.
[0102] Subsequently, a rule layer is constructed. This layer defines the operational constraints of the supply chain, such as the maximum production capacity of component manufacturers, minimum safety stock levels, and delivery time windows for product assembly plants. The constraint weights are adjusted through a dynamic adjustment mechanism and iterative optimization rules. The dynamic adjustment mechanism includes: when the actual output of the component manufacturer agent deviates from the corresponding parameters predicted by the supply chain game model by more than a first preset threshold (e.g., actual output is 5% lower than the predicted value), the effective production capacity parameters of the manufacturer agent are corrected through a parameter fine-tuning mechanism. When market demand drops sharply, and the corresponding change exceeds the corresponding second preset threshold, the system triggers a backup strategy, such as initiating a preset production reduction or promotion plan to cope with sudden environmental changes. The iterative optimization rules include: when the deviation rate between current operating data, such as actual inventory costs, and the predictions of the supply chain game model exceeds a third threshold, the corresponding parameters, such as the inventory holding cost coefficient, are adjusted. When the supply chain game model gets stuck in a local optimum for a certain production plan, such as when there is a persistent slight oversupply of a certain component, the system will inject random noise into the policy function or increase the exploration rate, prompting the component manufacturer agent to try to jump out of the existing policy domain and explore better production and inventory management strategies, so as to avoid being stuck in a suboptimal solution for a long time and improve the robustness and adaptability of the model.
[0103] Based on the constructed intelligent agent layer, relation layer, and rule layer, a game theory model for the industrial chain is obtained.
[0104] Based on this supply chain game model, the system obtains an initial supply chain strategy. This strategy is a comprehensive production, procurement, inventory, and sales plan aimed at optimizing the operational efficiency and effectiveness of the entire supply chain.
[0105] In actual operation, the system continuously responds to feedback data. For example, when feedback data such as actual sales data, raw material price fluctuations, and production line malfunctions are input, the system adjusts the corresponding parameters in the supply chain game model based on this data. For instance, if actual sales volume consistently falls below expectations, the model adjusts the market demand forecast parameters; if raw material prices rise, the model adjusts the procurement cost parameters. Subsequently, based on the adjusted supply chain game model, the system re-acquires the target supply chain strategy. This dynamic adjustment and iterative optimization process enables the supply chain strategy to adapt to environmental changes in real time, improving the accuracy of planning and effectively solving the problem of existing technologies being highly dependent on data accuracy and lacking dynamic adjustment mechanisms. In this way, this solution provides an innovative, accurate, and robust strategy acquisition method for complex electronic component manufacturing supply chains.
[0106] This embodiment also provides a supply chain strategy acquisition device, which is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0107] This embodiment provides a device for acquiring supply chain strategies, such as... Figure 4 As shown, it includes: Module 401 is used to build a supply chain game model based on the intelligent agent layer, the relation layer, and the rule layer. The intelligent agent layer is used to simulate the autonomous decision-making of supply chain participants, the relation layer is used to define the upstream and downstream interaction logic, and the rule layer is used to integrate global constraints. The first acquisition module 402 is used to acquire the initial industrial chain strategy based on the industrial chain game model; The second acquisition module 403 is used to adjust the corresponding parameters in the industrial chain game model in response to feedback data, and to acquire the target industrial chain strategy based on the adjusted industrial chain game model.
[0108] In some alternative implementations, the construction module 401 includes: The first unit is used to construct the intelligent agent layer. The construction of the intelligent agent layer includes abstracting the various industry chain entities into different intelligent agents and determining the attributes and decision rules of each intelligent agent. The second unit is used to construct the relational layer, wherein the construction of the relational layer includes constructing the topological relationship between intelligent agents and determining the information interaction rules and benefit distribution mechanism between intelligent agents, wherein the topological relationship includes at least one of the material flow, capital flow, and information flow; The third unit is used to construct the rule body layer. The construction of the rule body layer includes defining the operational constraints of the industry chain and adjusting the constraint weights through dynamic adjustment mechanisms and iterative optimization of rules. The fourth unit is used to obtain the industry chain game model based on the constructed intelligent agent layer, relation layer, and rule layer.
[0109] The supply chain strategy acquisition device provided in this application embodiment can execute the supply chain strategy acquisition method provided in any embodiment of this application, and has the corresponding functional modules and beneficial effects of the method execution. Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.
[0110] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0111] The following is a detailed reference. Figure 5 The diagram illustrates a structural schematic suitable for implementing the electronic device described in the embodiments of this application. The electronic device may include a processor (e.g., a central processing unit, graphics processor, etc.) 501, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 502 or a program loaded from memory 508 into random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of the electronic device. The processor 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0112] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; memory devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.
[0113] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a memory 508, or installed from a ROM 502. When the computer program is executed by the processor 501, it performs the functions defined in the supply chain strategy acquisition method of embodiments of this application.
[0114] Figure 5 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0115] This application also provides a computer-readable storage medium. The methods described in this application can be implemented in hardware or firmware, or implemented as recordable on a storage medium, or implemented as computer code downloaded over a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the method for acquiring the supply chain strategy shown in the above embodiments is implemented.
[0116] A portion of this application can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to this application through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.
[0117] Although embodiments of this application have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of this application, and all such modifications and variations fall within the scope defined by the appended claims.
Claims
1. A method for acquiring an industry chain strategy, characterized in that, The method includes: A supply chain game model is constructed based on an intelligent agent layer, a relational layer, and a rule layer. The intelligent agent layer is used to simulate the autonomous decision-making of supply chain participants, the relational layer is used to define the upstream and downstream interaction logic, and the rule layer is used to integrate global constraints. Based on the aforementioned supply chain game model, obtain the initial supply chain strategy; In response to feedback data, the corresponding parameters in the supply chain game model are adjusted, and the target supply chain strategy is obtained based on the adjusted supply chain game model.
2. The method according to claim 1, characterized in that, The construction of the industry chain game model based on the intelligent agent layer, relation layer, and rule layer includes: Constructing an intelligent agent layer, wherein the construction of the intelligent agent layer includes abstracting each industry chain entity into different intelligent agents, and determining the attributes and decision rules of each intelligent agent; Constructing a relational layer, wherein constructing the relational layer includes constructing topological relationships between the intelligent agents and determining information interaction rules and benefit distribution mechanisms between the intelligent agents, wherein the topological relationships include at least one of material flow, capital flow, and information flow; Constructing a rule body layer, wherein the construction of the rule body layer includes defining the operational constraints of the industry chain and adjusting the constraint weights through a dynamic adjustment mechanism and iterative optimization of the rules; Based on the constructed intelligent agent layer, relation layer, and rule layer, the industrial chain game model is obtained.
3. The method according to claim 2, characterized in that, The determination of decision rules in the intelligent agent layer includes: Based on the mapping relationship between the objective function and the state space, a preset algorithm is used to enable the agent to adjust its behavior according to environmental changes. The preset algorithm includes reinforcement learning strategies or multi-criteria decision analysis algorithms, the behavior includes production, supply, and sales, and the environmental changes include market environment changes.
4. The method according to claim 2, characterized in that, The determination of information interaction rules in the relational layer includes: The interactive information is divided into a public signal layer, a business collaboration layer, and a core status layer. The system employs an intelligent routing mode and a publish-subscribe mechanism to transmit information at each layer of the interactive information, and controls the accuracy of the agent's perception of the external environment through view filtering rules. Set an information threshold filter to fuse and compress repetitive or low-value signals to prevent overload and feedback suppression.
5. The method according to claim 2, characterized in that, The dynamic adjustment mechanism in the rule body layer includes: When the actual operating parameters of the intelligent agent deviate from the corresponding parameters predicted by the industrial chain game model by more than a first preset threshold, the parameter fine-tuning mechanism is used to correct the deviation. When the change in the environment exceeds the corresponding second preset threshold, a decision is made using a backup strategy.
6. The method according to claim 2, characterized in that, The iterative optimization rules in the rule body layer include: When the deviation rate between the current operating data and the prediction of the industrial chain game model exceeds the third threshold, the corresponding parameters are adjusted. When the industry chain game model gets stuck in a local optimum, random noise is injected into the policy function or the exploration rate is increased to make the agent jump out of the existing policy domain.
7. The method according to claim 3, characterized in that, The objective function includes: , where Ui,t(s t ,a i Let Φi,t(s) be the utility of the i-th agent at step t. t ,a i ,t) represents the constraint term of the i-th agent at step t, λ is the penalty factor, γ is the discount factor, and Eπheta [ ] represents the expectation under strategy πheta.
8. A device for acquiring supply chain strategies, characterized in that, The device includes: The construction module is used to build a supply chain game model based on the intelligent agent layer, the relation layer, and the rule layer. The intelligent agent layer is used to simulate the autonomous decision-making of supply chain participants, the relation layer is used to define the upstream and downstream interaction logic, and the rule layer is used to integrate global constraints. The first acquisition module is used to acquire the initial industrial chain strategy based on the industrial chain game model. The second acquisition module is used to adjust the corresponding parameters in the industrial chain game model in response to feedback data, and to acquire the target industrial chain strategy based on the adjusted industrial chain game model.
9. An electronic device, characterized in that, include: A memory and a processor are interconnected, the memory storing computer instructions, and the processor executing the computer instructions to perform the method for acquiring the supply chain strategy according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a computer to execute the method for obtaining the supply chain strategy according to any one of claims 1 to 7.