A model building method and system based on event-driven

Through an event-driven model construction method, the timing prediction model is used to predict the future state of supply chain events, and the reinforcement learning model is updated based on these states, which solves the problem of inaccurate production planning in the existing technology and achieves more accurate and reliable production planning output.

CN119886466BActive Publication Date: 2025-05-23INSPUR SMART SUPPLY CHAIN TECH (SHANDONG) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510362104.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-05-23
Estimated Expiration
2045-03-26

AI Technical Summary

Technical Problem

When optimizing and managing supply chain production plans, the existing technology ignores the impact of emergencies in each link in the supply chain on production plans, resulting in inaccurate production plans.

Method used

The event-driven model construction method is adopted to predict the future state of each event through the timing prediction model, and the applicability of the reinforcement learning model is judged based on the current state and future state of the event. If the reinforcement learning model is not applicable, update the model to ensure it outputs an accurate production plan in its future state.

Benefits of technology

By updating the reinforcement learning model in real time, it can accurately reflect the event status of each link in the supply chain, thereby outputting more accurate production plans and improving the accuracy and reliability of production plans.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119886466B_ABST
    Figure CN119886466B_ABST
Patent Text Reader

Abstract

The present application relates to the field of model building technology, and in particular to a model building method and system based on event-driven, the method comprising: inputting multiple historical states and current states of any event into a time series prediction model to obtain the future state of each event; judging whether the reinforcement learning model is applicable; if applicable, inputting order information and the production status of each production line into the reinforcement learning model to obtain the production plan for the next production cycle; if not applicable, updating the reinforcement learning model and then outputting the production plan. Through the technical solution of the present application, the reinforcement learning model can be updated in time according to the event status of each link in the supply chain, ensuring that the reinforcement learning model can output an accurate production plan.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of model building, and in particular to an event-driven model building method and system. Background Art

[0002] The optimization and management of production plans of enterprises in the supply chain are particularly important. A good production plan is not only related to the efficiency of enterprise production, but also directly affects the quality and delivery time of products. However, a complete supply chain includes multiple links such as procurement, production, warehousing and transportation. Different links affect each other, and emergencies may occur in each link. For example, there may be emergencies such as insufficient inventory or inventory backlog in the warehousing link, and there may be emergencies such as insufficient transportation capacity or traffic congestion in the transportation link. These emergencies will affect the production plan of the enterprise and lead to inaccurate production plans.

[0003] At present, a patent application document with application publication number CN118780648A discloses a production plan evaluation method and system for the aviation manufacturing industry, wherein the method includes: formulating the overall goal of the production plan of the aviation manufacturing industry, and constructing an evaluation index reflecting the execution effect of the production plan; designing a specific implementation plan for each production line according to the characteristics and overall goals of the production line; simulating the implementation plan of each production line, and processing the data of each production line by optimizing the data grouping algorithm, and calculating the actual index value of the evaluation index of each production line; establishing a dynamic monitoring mechanism, and analyzing the data of each production line by using a reinforcement learning algorithm; comparing the target value in the evaluation index with the actual index value, evaluating the overall effect of the production plan, and adjusting the production plan of the aviation manufacturing industry in combination with the analysis results of the reinforcement learning algorithm, wherein the evaluation index of each production line includes production efficiency index, quality control index and cost control index.

[0004] The above method monitors production efficiency indicators, quality control indicators and cost control indicators, compares the target values ​​with the actual indicator values ​​to evaluate the overall effect of the production plan, and uses the reinforcement learning algorithm to adjust the production plan. However, it ignores the impact of emergencies in other links in the supply chain on the production plan, resulting in inaccurate production plans. Summary of the invention

[0005] In order to solve the technical problem of inaccurate production plans, the present application provides an event-driven model building method and system, which can timely update the reinforcement learning model according to the event status of each link in the supply chain, ensuring that the reinforcement learning model can output accurate production plans.

[0006] In a first aspect, the present application provides an event-driven model construction method, the construction method comprising: inputting multiple historical states and current states of any event into a time series prediction model to obtain the future state of each event, the event comprising a transportation event, an inventory event, and a supply event; judging whether a reinforcement learning model is applicable, comprising: clustering the interaction samples in the experience pool according to the states of other events other than the target event to obtain multiple clustering clusters; determining the target cluster according to the current states of other events, and drawing a reward value change curve in the target cluster with the state of the target event as the horizontal coordinate; dividing the change curve into multiple stable segments; in response to the current state and the future state of the target event being in the same stable segment, the target event has a marking value of 0, otherwise, the difference between the average reward value of the stable segment where the future state and the current state are located is used as the marking value of the target event; in response to the sum of the marking values ​​of each event being less than 0, the reinforcement learning model is not applicable, otherwise, the reinforcement learning model is applicable; if applicable, the order information and the production status of each production line are input into the reinforcement learning model to obtain a production plan for the next production cycle; if not applicable, the reinforcement learning model is updated and then the production plan is output.

[0007] The time series prediction model is used to predict the future state of each event in the next production cycle. According to the current state and future state of each event, it is judged whether the reinforcement learning model can output the target action with better control effect in the future state of each event. If the reinforcement learning model can output the target action with better control effect, it means that the reinforcement learning model is still applicable. The order information and the production status of each production line are directly input into the reinforcement learning model to obtain the production plan for the next production cycle. If the reinforcement learning model cannot output the target action with better control effect, it means that the reinforcement learning model is not applicable. At this time, the reinforcement learning model needs to be updated, and the updated reinforcement learning model is used to output the production plan.

[0008] Preferably, the status of the transportation event is the transportation duration, the status of the inventory event is the inventory quantity, and the status of the supply event is the supply quantity of raw materials or upstream enterprises.

[0009] Based on the status of transportation events, it can be determined whether there are emergencies such as insufficient capacity or traffic congestion in the transportation link; based on the status of inventory events, it can be determined whether there are emergencies such as insufficient inventory or inventory backlog in the warehousing link; based on the status of supply events, it can be determined whether there are emergencies such as insufficient supply in the procurement link, and the event status of each link in the supply chain can be accurately quantified.

[0010] Preferably, the training method of the time series prediction model includes: collecting a state sequence of transportation events, the state sequence including multiple historical states of the transportation events, and taking the states of the transportation events in the adjacent production cycles after the state sequence as state labels; after inputting the state sequence into the time series prediction model, calculating the mean square error loss function between the output result and the state label, and updating the time series prediction model using the gradient descent method until the mean square error loss function is less than the preset loss, or the number of updates is greater than the preset number, the training is completed.

[0011] Preferably, determining the target cluster based on the current state of other events includes: calculating the average state of other events of each interaction sample in any clustering cluster, calculating the sum of the Euclidean distances between the current state of other events and the average state, and taking the cluster corresponding to the minimum sum of the Euclidean distances as the target cluster.

[0012] The states of other events of each interaction sample in the target cluster are the same as the current state. Subsequently, a change curve of the reward value is drawn in the target cluster. This change curve can reflect the impact of the state of the target event on the reward value when the state of other events is the current state.

[0013] Preferably, drawing a reward value change curve in the target cluster includes: screening out interaction samples under any state of the target event in the target cluster, and in response to the number of interaction samples being greater than a quantity threshold, taking the average true reward value of the interaction samples as the mean reward value of the state; drawing an initial curve based on the mean reward value of each state; and interpolating the initial curve to obtain a reward value change curve.

[0014] If a state in the target cluster can filter out interaction samples above the threshold, it means that the accurate mean reward value in this state can be obtained. An initial curve is drawn based on the obtained mean reward value. The initial curve includes multiple discrete points. The initial curve is further interpolated to obtain an accurate reward value change curve.

[0015] Preferably, dividing the change curve into a plurality of stable segments comprises: segmenting the change curve according to ordered sample clustering to obtain a plurality of stable segments.

[0016] Preferably, the updating of the reinforcement learning model includes: clustering the interaction samples in the experience pool according to the state of each event to obtain multiple state clusters, calculating the sum of the Euclidean distances of the average state and future state of all events in any state cluster, and taking the state cluster corresponding to the minimum sum of the Euclidean distances as the training cluster; training the reinforcement learning model based on the interaction samples in the training cluster.

[0017] The states of each event are selected from the experience pool as interactive samples corresponding to the future states to train the reinforcement learning model, so as to update the reinforcement learning model and enable the updated reinforcement learning model to output target actions with better effects in the future states of each event.

[0018] Preferably, the reinforcement learning model includes a decision sub-model and a reward value sub-model, the interaction samples include the state of each event, the environment state, the actual reward value, the target action and the environment state after executing the target action, and the environment state is the order information and the production status of each production line; the training of the reinforcement learning model based on the interaction samples in the training cluster includes: inputting the environment state in any interaction sample into the reinforcement learning model to obtain a predicted reward value; inputting the environment state after executing the target action in the interaction sample into the reinforcement learning model to obtain a future reward value; calculating the reward value loss and the decision loss, updating the reward value sub-model according to the reward value loss, and updating the decision sub-model according to the decision loss.

[0019] Preferably, the reward value loss for:

[0020] , To predict the reward value, For interactive samples The real reward value, is the future reward value, is the discount factor; the decision loss for:

[0021] , is the predicted reward value.

[0022] In a second aspect of the present application, there is also provided an event-driven model building system, comprising a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, an event-driven model building method according to the first aspect of the present application is implemented.

[0023] The technical solution of this application has the following beneficial technical effects:

[0024] The time series prediction model is used to predict the future state of each event in the next production cycle. According to the current state and future state of each event, it is judged whether the reinforcement learning model can output the target action with better control effect in the future state of each event. If the reinforcement learning model can output the target action with better control effect, it means that the reinforcement learning model is still applicable. The order information and the production status of each production line are directly input into the reinforcement learning model to obtain the production plan for the next production cycle. If the reinforcement learning model cannot output the target action with better control effect, it means that the reinforcement learning model is not applicable. At this time, the reinforcement learning model needs to be updated, and the updated reinforcement learning model is used to output the production plan.

[0025] Furthermore, in order to accurately judge whether the reinforcement learning model can output the target action with better control effect in the future state of each event, the interaction samples in the experience pool are clustered according to the states of other events other than the target event, and multiple clusters are obtained. The states of other events of each interaction sample in a cluster are the same; the cluster in which the state of other events is equal to the current state is taken as the target cluster, and a reward value change curve is drawn in the target cluster. The change curve can characterize the impact of the state change of the target event on the reward value when other events are in the current state, and then accurately quantify the change of the reward value when the target event changes from the current state to the future state; the reward value changes corresponding to all events are comprehensively considered. When the reward value becomes smaller, it means that the reinforcement learning model cannot output the target action with better control effect in the future state, and it is accurately judged whether the reinforcement learning model is applicable. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 It is a flowchart of an event-driven model building method according to an embodiment of the present application.

[0027] Figure 2 It is a structural diagram of a reinforcement learning model according to an embodiment of the present application.

[0028] Figure 3 It is a structural block diagram of an event-driven model building system according to an embodiment of the present application. DETAILED DESCRIPTION

[0029] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.

[0030] According to the first aspect of the present application, the present application provides an event-driven model building method for building a reinforcement learning model, so that the reinforcement learning model can accurately output the production plan of each production line in the next production cycle.

[0031] Figure 1 FIG. 1 is a flow chart of a method for constructing a model based on event-driven operation according to an embodiment of the present application. Figure 1 As shown, the event-driven model building method includes steps S101 to S104, which are described in detail below.

[0032] S101, input the current state of any event and the historical state of the historical production cycle into a time series prediction model to obtain the future state of each event, wherein the events include transportation events, inventory events and supply events.

[0033] In one embodiment, the transportation event is used to characterize the status of the transportation link in the supply chain. The status of the transportation event is the transportation time, which means the time it takes to transport the product to the consignee after the order is produced. According to the status of the transportation event, it can be determined whether there are emergencies such as insufficient transportation capacity or traffic congestion in the transportation link; the inventory event is used to characterize the status of the warehousing link in the supply chain. The status of the inventory event is the inventory quantity. According to the status of the inventory event, it can be determined whether there are emergencies such as insufficient inventory or inventory backlog in the warehousing link; the supply event is used to characterize the status of the procurement link in the supply chain. The status of the supply event is the supply quantity of raw materials or upstream enterprises. According to the status of the supply event, it can be determined whether there are emergencies such as insufficient supply in the procurement link.

[0034] Among them, one event corresponds to one time series prediction model. Taking the transportation event as an example, the input of the time series prediction model is the current state of the transportation event in the current production cycle and the historical state of the transportation event in multiple historical production cycles, and the output is the future state of the transportation event in the next production cycle. Among them, the number of historical production cycles is set to 4.

[0035] In order to ensure that the time series prediction model can output accurate future states, it is necessary to train the time series prediction model. The training method of the time series prediction model includes: collecting a state sequence of a transportation event, wherein the state sequence includes multiple historical states of the transportation event, and using the state of the transportation event in the adjacent production cycle after the state sequence as a state label; after inputting the state sequence into the time series prediction model, calculating the mean square error loss function between the output result and the state label, and using the gradient descent method to update the time series prediction model until the mean square error loss function is less than the preset loss, or the number of updates is greater than the preset number, the training is completed.

[0036] Among them, the production cycle is 1 day or a week, which is set according to the production law of the enterprise; the time series prediction model is a recurrent neural network, and specifically an LSTM model or a GUR model can be used.

[0037] The preset loss is 0.01, and the preset number of times is 50. The trained time series prediction model can accurately predict the future state of transportation events in the next production cycle.

[0038] In this way, the time series prediction model of each event is obtained respectively, and the future status of each event in the next production cycle is accurately predicted.

[0039] S102, determining whether the reinforcement learning model is applicable.

[0040] In one embodiment, the input of the reinforcement learning model is order information and the production status of each production line, and the output is the production plan for the next production cycle, the order information includes delivery time and delivery quantity, the production status of the production line is product quality and production efficiency, and the production plan is the output of each production line. The reinforcement learning model can adopt the DQN model.

[0041] It can be understood that the order information and the production status of each production line are defined as the environmental state, and reinforcement learning can output the target action with the maximum reward value under the environmental state. The target action is the production plan for the next production cycle. The larger the reward value, the better the effect of the target action.

[0042] In the embodiment of the present application, the reward value is the weighted sum of the delivery time sub-reward, the delivery quantum reward and the production line load sub-reward. for:

[0043] , N is the number of production lines, is the completion time of production line N, is the transportation time, T is the delivery time in the order information, and the delivery time sub-reward The larger the value of, the earlier the order information is completed than the delivery time in the order information.

[0044] The delivered quantum reward for:

[0045] , is the completion volume of production line n, Deliver quantum rewards for the delivery quantity in the order information The maximum value is 0, when the quantum reward is delivered When it reaches the maximum value, it means that the completion volume of each production line can meet the delivery volume of the order information.

[0046] The production line load sub-reward for:

[0047] , is the variance of the completion time of all production lines. When the variance of the completion time of all production lines is equal to 0, the production line load sub-reward When the maximum value is reached, it means that all production lines have completed their respective completion quantities at the same time, and there is no need to wait for other production lines to complete before shipping.

[0048] In summary, the reward value Satisfies the relationship:

[0049] ,in, , and They are the weights of the delivery time sub-reward, delivery quantum reward and production line load sub-reward respectively.

[0050] During the training process of the reinforcement learning model, an experience pool is constructed, which includes multiple interaction samples. The interaction samples include the status of each event, the environment status (i.e., order information and the production status of each production line), the real reward value, the target action, and the environment status after executing the target action. For example, the historical production cycle The status of transport events, inventory events, and supply events are , and , environmental status Execute the target action The delivery time sub-reward, delivery quantum reward and production line load sub-reward are , and , and the state of the environment after executing the target action ,in accordance with , and Calculate historical production cycles Execute the target action The real reward value ; then you can get an interactive sample .

[0051] In one embodiment, after obtaining the future state of each event, it is determined whether the reinforcement learning model is applicable based on the future state of each event. When the reinforcement learning model is not applicable, the reinforcement learning model is updated in a timely manner so that the updated reinforcement learning model is applicable to the future state of each event, thereby ensuring that the reinforcement learning model outputs an accurate production plan.

[0052] Specifically, judging whether the reinforcement learning model is applicable includes: clustering the interaction samples in the experience pool according to the states of other events other than the target event to obtain multiple clusters; determining the target cluster according to the current states of other events, and drawing a reward value change curve in the target cluster with the state of the target event as the horizontal coordinate; dividing the change curve into multiple stable segments; in response to the current state and future state of the target event being in the same stable segment, the marking value of the target event is 0, otherwise, the difference between the average reward value of the stable segment where the future state and the current state are located is used as the marking value of the target event; in response to the sum of the marking values ​​of each event being less than 0, the reinforcement learning model is not applicable, otherwise, the reinforcement learning model is applicable.

[0053] The target event is any one of the events. When the target event is a transportation event, the interaction samples in the experience pool are clustered according to the status of the inventory event and the supply event to obtain multiple clusters. A cluster includes multiple interaction samples, and the status of the inventory event and the supply event in each interaction sample of a cluster is the same. In other words, the status of other events of each interaction sample in a cluster is the same.

[0054] In one embodiment, determining the target cluster based on the current state of other events includes: calculating the average state of other events of each interaction sample in any clustering cluster, calculating the sum of the Euclidean distances between the current state of other events and the average state, and taking the cluster corresponding to the minimum sum of the Euclidean distances as the target cluster.

[0055] Among them, the states of other events of each interaction sample in the target cluster are the same as the current state, and then a reward value change curve is drawn in the target cluster, and the change curve can reflect the impact of the state of the target event on the reward value when the state of other events is the current state. Specifically, drawing a reward value change curve in the target cluster includes: screening out interaction samples in any state of the target event in the target cluster, and in response to the number of interaction samples being greater than a quantity threshold, taking the average true reward value of the interaction samples as the mean reward value of the state; drawing an initial curve based on the mean reward value of each state; and interpolating the initial curve to obtain a reward value change curve.

[0056] Among them, the quantity threshold is 2, that is, if in the target cluster, a state can screen out more than 2 interaction samples, it means that the accurate mean reward value in this state can be obtained. The initial curve is drawn based on the obtained mean reward value. The initial curve includes multiple discrete points. The initial curve is further interpolated to obtain an accurate reward value change curve.

[0057] In one embodiment, dividing the change curve into a plurality of stable segments includes: segmenting the change curve according to ordered sample clustering to obtain a plurality of stable segments. Ordered sample clustering is a well-known technique for those skilled in the art and will not be described in detail herein.

[0058] The mean reward value within a stable segment remains basically unchanged, and there is a large difference in the mean reward value between adjacent stable segments. In response to the current state and future state of the target event being in the same stable segment, it means that after the target event changes from the current state to the future state, the mean reward value remains basically unchanged, that is, the state change of the target event does not cause the reward value to decrease, and the marking value of the target event is 0; conversely, the current state and future state of the target event are not in the same stable segment, indicating that the state change of the target event causes the reward value to change, and the difference between the average reward values ​​of the stable segments in which the future state and the current state are located is used as the marking value of the target event.

[0059] It can be understood that when the difference between the average reward value of the stable segment between the future state and the current state is greater than or equal to 0, it means that the state change of the target event causes the reward value to increase. At this time, the reinforcement learning model can output a target action with better effect; when the difference between the average reward value of the stable segment between the future state and the current state is less than 0, it means that the state change of the target event causes the reward value to decrease. At this time, the effect of the target action output by the reinforcement learning model is poor.

[0060] The sum of the tag values ​​of each event is calculated. If the sum of the tag values ​​of each event is less than 0, it means that the effect of the target action output by the reinforcement learning model is poor. At this time, the reinforcement learning model is not applicable. Otherwise, the reinforcement learning model is applicable.

[0061] S103, if applicable, input the order information and the production status of each production line into the reinforcement learning model to obtain the production plan for the next production cycle.

[0062] In one embodiment, if the reinforcement learning model is applicable, the currently collected order information and the production status of each production line are directly input into the reinforcement learning model to obtain a production plan for the next production cycle.

[0063] S104: If not applicable, update the reinforcement learning model and then output the production plan.

[0064] In one embodiment, if the reinforcement learning model is not applicable, the reinforcement learning model needs to be updated, and the updated reinforcement learning model can output accurate target actions in the future state of each event. Specifically, the updating of the reinforcement learning model includes: clustering the interaction samples in the experience pool according to the state of each event to obtain multiple state clusters, calculating the sum of the Euclidean distances of the average state and future state of all events in any state cluster, and taking the state cluster corresponding to the minimum sum of the Euclidean distances as the training cluster; training the reinforcement learning model according to the interaction samples in the training cluster.

[0065] In one embodiment, Figure 2 It is a structural diagram of a reinforcement learning model according to an embodiment of the present application; the reinforcement learning model adopts a DQN neural network, including a decision sub-model and a reward value sub-model, the input of the decision sub-model is the environment state, and the output is the target action; the input of the reward value sub-model is the environment state and the target action, and the output is the predicted reward value for executing the target action under the environment state. Among them, training the reinforcement learning model according to the interaction samples in the training cluster includes: inputting the environment state in any interaction sample into the reinforcement learning model to obtain a predicted reward value; inputting the environment state after executing the target action in the interaction sample into the reinforcement learning model to obtain a future reward value, wherein the environment state is the order information and the production status of each production line; calculating the reward value loss and the decision loss, updating the reward value sub-model according to the reward value loss, and updating the decision sub-model according to the decision loss; the reward value loss for:

[0066] , To predict the reward value, For interactive samples The real reward value, is the future reward value, is the discount factor; the decision loss for:

[0067] , is the predicted reward value.

[0068] The discount coefficient is 0.5.

[0069] In this way, the state of each event is selected from the experience pool as the interactive samples corresponding to the future state to train the reinforcement learning model, so as to update the reinforcement learning model, so that the updated reinforcement learning model can output target actions with better effects in the future state of each event, thereby realizing the construction of an event-driven reinforcement learning model.

[0070] According to the second aspect of the present application, the present application also provides an event-driven model building system. Figure 3is a structural block diagram of an event-driven model building system according to an embodiment of the present application. Figure 3 As shown, the system 50 includes a processor and a memory, the memory stores computer program instructions, and when the computer program instructions are executed by the processor, an event-driven model building method according to the first aspect of the present application is implemented. The system also includes other components well known to those skilled in the art, such as a communication bus and a communication interface, whose settings and functions are known in the art, so they are not described here.

[0071] It should be pointed out that, for ordinary technicians in this field, several modifications and improvements can be made without departing from the concept of the present application, which all fall within the scope of protection of the present application.

Claims

1. A model building method based on event-driven, characterized in that: The construction method comprises: Inputting multiple historical states and current states of any event into a time series prediction model to obtain the future state of each event, wherein the event includes a transportation event, an inventory event, and a supply event; Determining whether the reinforcement learning model is applicable includes: clustering the interaction samples in the experience pool according to the states of other events other than the target event to obtain multiple clusters; determining the target cluster according to the current states of other events, and drawing a reward value change curve in the target cluster with the state of the target event as the horizontal coordinate; dividing the change curve into multiple stable segments; in response to the current state and the future state of the target event being in the same stable segment, the marking value of the target event is 0, otherwise, the difference between the average reward value of the stable segment where the future state and the current state are located is used as the marking value of the target event; in response to the sum of the marking values ​​of each event being less than 0, the reinforcement learning model is not applicable, otherwise, the reinforcement learning model is applicable; If applicable, the order information and the production status of each production line are input into the reinforcement learning model to obtain the production plan for the next production cycle; if not applicable, the reinforcement learning model is updated and then the production plan is output.

2. The event-driven model building method according to claim 1, characterized in that: The status of the transportation event is the transportation duration, the status of the inventory event is the inventory quantity, and the status of the supply event is the supply quantity of raw materials or upstream enterprises.

3. The event-driven model building method according to claim 1, characterized in that: The training method of the time series prediction model includes: Collecting a state sequence of a transport event, wherein the state sequence includes a plurality of historical states of the transport event, and using the state of the transport event in an adjacent production cycle after the state sequence as a state label; After the state sequence is input into the time series prediction model, the mean square error loss function between the output result and the state label is calculated, and the time series prediction model is updated using the gradient descent method until the mean square error loss function is less than the preset loss, or the number of updates is greater than the preset number, the training is completed.

4. The event-driven model building method according to claim 1, characterized in that: Determining the target cluster according to the current state of other events includes: Calculate the average state of other events of each interaction sample in any cluster, calculate the sum of the Euclidean distances between the current state of other events and the average state, and take the cluster corresponding to the minimum sum of the Euclidean distances as the target cluster.

5. The event-driven model building method according to claim 1, characterized in that: Drawing a reward value change curve in the target cluster includes: Filter out interaction samples in any state of the target event in the target cluster, and in response to the number of interaction samples being greater than a quantity threshold, use an average true reward value of the interaction samples as a mean reward value of the state; Draw the initial curve based on the mean reward value of each state; The initial curve is interpolated to obtain a reward value variation curve.

6. The event-driven model building method according to claim 1, characterized in that: Dividing the change curve into multiple stable segments includes: The variation curve is segmented according to ordered sample clustering to obtain multiple stable segments.

7. The event-driven model building method according to claim 1, characterized in that: The updating reinforcement learning model comprises: Cluster the interaction samples in the experience pool according to the state of each event to obtain multiple state clusters. Calculate the sum of the Euclidean distances between the average state and the future state of all events in any state cluster, and take the state cluster corresponding to the minimum sum of the Euclidean distances as the training cluster. Train the reinforcement learning model based on the interaction samples in the training cluster.

8. The event-driven model building method according to claim 7, characterized in that: The reinforcement learning model includes a decision sub-model and a reward value sub-model, and the interaction samples include the state of each event, the environment state, the real reward value, the target action, and the environment state after executing the target action. The environment state is the order information and the production state of each production line; The training of the reinforcement learning model according to the interaction samples in the training cluster includes: Input the environment state in any interaction sample into the reinforcement learning model to obtain the predicted reward value; The state of the environment after executing the target action in the interaction sample is input into the reinforcement learning model to obtain the future reward value; Calculate the reward value loss and decision loss, update the reward value sub-model according to the reward value loss, and update the decision sub-model according to the decision loss.

9. The event-driven model building method according to claim 8, characterized in that: The reward value loss for: , To predict the reward value, For interactive samples The real reward value, is the future reward value, is the discount factor; The decision loss for: , is the predicted reward value.

10. A model building system based on event-driven, characterized in that: The method comprises a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, an event-driven model building method according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Production plan evaluation method and system for aeronautical manufacturing industry

    CN118780648A

  • Off-line reinforcement learning method and device for recommending safe treatment scheme

    CN116959737A

  • Power distribution network source network load storage low-carbon optimization scheduling increment reinforcement learning method and system

    CN119623567A