Method and electronic device for state valuation of a work environment

CN122817718APending Publication Date: 2026-09-25BSH HOME APPLIANCES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610983662.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-02
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]传统方法通常基于数学或物理模型进行估值,或依赖人工经验判断,但此类方式准确性较差,难以应对动态、复杂且多变的作业环境

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122817718A_ABST
    Figure CN122817718A_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a method for state evaluation of a working environment, comprising: obtaining state information related to a specified state of the working environment; representing the state information of the specified state as heterogeneous feature units, the heterogeneous feature units comprising at least: a task type feature unit, a time slot feature unit, and a resource feature unit; and generating, by means of a trained machine learning model, a state value corresponding to the specified state based on the heterogeneous feature units. By encoding state information into independent feature units according to different semantic categories, the machine learning model can more effectively analyze and fuse multi-source heterogeneous information, improving the accuracy and robustness of state value estimation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent scheduling, and in particular to a method for estimating the state of a working environment, a method for training a machine learning model, an electronic device, a computer program product, and a computer-readable storage medium. Background Technology

[0002] With the improvement of automation and the rapid development of digital supply chain technology, in high-end manufacturing, intelligent logistics and other operating environments, it is often necessary to estimate the value of the operating environment to help optimize future scheduling plans, predict revenue or assess profit levels.

[0003] Traditional methods typically rely on mathematical or physical models for estimation or depend on human experience. However, these approaches are often inaccurate and struggle to handle dynamic, complex, and ever-changing work environments. While some machine learning-based estimation schemes have emerged in recent years, most simply input the complex state information from the work environment into the model in a flattened format, failing to effectively distinguish between different semantic categories and resulting in low model parsing efficiency. Furthermore, faced with flattened information input, machine learning models struggle to effectively uncover the interactions between different semantic dimensions, leading to state estimation results that cannot objectively and accurately reflect the true state of the work environment and its long-term profit potential.

[0004] Therefore, existing methods for estimating the state of the working environment still have significant limitations. Summary of the Invention

[0005] The purpose of embodiments of this application is to provide a method for estimating the state of a working environment, a method for training a machine learning model, an electronic device, a computer program product, and a computer-readable storage medium, to at least solve some of the problems in the prior art.

[0006] According to a first aspect of this application, a method for estimating the state of a working environment is provided, comprising: Obtain status information related to a specified state of the operating environment; The state information of the specified state is represented as a heterogeneous feature unit, which includes at least: a task type feature unit, a time slot feature unit, and a resource feature unit; Using a trained machine learning model, a state value corresponding to the specified state is generated based on the heterogeneous feature units.

[0007] This application specifically includes the following technical concept: by encoding state information into independent feature units according to different semantic categories such as task type, time slot, and resources, the model can more effectively parse and fuse multi-source heterogeneous information. After the same set of demand data is explicitly modeled from different dimensions, the model can more comprehensively grasp the demand distribution characteristics of the working environment, improve its ability to perceive the global state, effectively avoid short-sighted local optimal decisions, and enhance the long-term global performance of the scheduling strategy.

[0008] For example, time slot feature units and task type feature units are complementary representations of the same set of demand data in the work environment across different dimensions (especially the vertical dimension): the former reflects the resource competition situation within each time slot from the time dimension, while the latter reflects the long-term demand trends of each type of task from the task type dimension. This complementary representation enables the model to simultaneously perceive the resource competition situation in the time dimension and the demand evolution in the task type dimension, thereby more accurately identifying time bottlenecks and resource conflicts.

[0009] According to an optional embodiment of this application, the state value is used to characterize the level of long-term benefits that the working environment can obtain in the future, starting from the specified state.

[0010] This clarifies the business meaning of state value, gives it a specific economic interpretation, and provides a unified and quantifiable evaluation benchmark for subsequent scheduling decisions.

[0011] According to an optional embodiment of this application, each time slot feature unit is used to characterize the workload distribution of each task type within the same time interval.

[0012] Therefore, by using time slot feature units to characterize the workload distribution of task type dimension, the model can better perceive the resource competition situation within the corresponding time period.

[0013] According to an optional embodiment of this application, each task type feature unit is used to characterize the workload distribution of the same task type in different time intervals.

[0014] Therefore, by characterizing the workload distribution over time using task type feature units, the model can perceive the demand change trend of each task type on the time axis.

[0015] According to an optional embodiment of this application, the method includes: dividing a future time window into multiple time slots; determining the time slot to which each task should belong based on the delivery deadline of each task; injecting the workload of each task into the corresponding time slot according to the task type corresponding to each task to form a demand matrix; generating the task type feature unit based on the rows of the demand matrix, and / or generating the time slot feature unit based on the columns of the demand matrix.

[0016] As a result, the demand information is explicitly organized into structured data, which facilitates the extraction of time slot feature units and task type feature units.

[0017] According to an optional embodiment of this application, the method further includes: superimposing position codes on the time slot feature units, the position codes being used to characterize the sequential order of each time slot in a future time window; and / or embedding superimposed types on the heterogeneous feature units to distinguish different categories of feature units.

[0018] Therefore, position encoding and type embedding enhance the model's ability to perceive spatiotemporal dimensions and semantic categories, thereby improving the accuracy of state value estimation.

[0019] According to an optional embodiment of this application, the resource feature unit includes a work equipment feature unit and a global feature unit, wherein: each work equipment feature unit corresponds to a work equipment and includes the earliest available time of the work equipment, the task type of the current work, and the currently assembled work accessories; the global feature unit includes the availability of shared auxiliary resources, the earliest available time, and global time information.

[0020] This allows the model to focus on local resource states while also being aware of global constraints.

[0021] According to an optional embodiment of this application, the method further includes: obtaining candidate scheduling actions in a determined state of the working environment, each candidate scheduling action representing assigning a task type to a focus working device in the working environment or representing keeping the focus working device idle; simulating the execution of each candidate scheduling action using a simulator of the working environment to determine the subsequent state and immediate reward of the working environment after the execution of each candidate scheduling action; wherein the specified state of the working environment includes at least the subsequent state.

[0022] Therefore, by incorporating state estimation into scheduling decisions, the agent does not need to directly learn the mapping from state to action, but only needs to learn state value estimation, significantly simplifying the learning objective and making model training more stable. At the same time, expanding the action space does not affect model performance or generalization ability, enabling the solution to flexibly adapt to real-world scheduling scenarios with large-scale action spaces.

[0023] According to an optional embodiment of this application, the method further includes: using the machine learning model, generating a state value corresponding to the subsequent state based on the subsequent state of the working environment; using a rule decision module, for each candidate scheduling action, determining the long-term benefit level corresponding to the candidate scheduling action based on the state value of the subsequent state and the immediate reward, and selecting the candidate scheduling action with the highest long-term benefit level as the target scheduling action.

[0024] Therefore, by dividing the agent into a machine learning model and a rule-based decision-making module, the final decision incorporates both the single-step reward calculated by the rule and the future reward estimate predicted by the model, effectively improving the interpretability and reliability of the decision. Simultaneously, the learning task of the machine learning model is limited to state value estimation, reducing the difficulty of model training.

[0025] According to an optional embodiment of this application, the machine learning model includes a bipartite graph attention network, a Transformer network, and a task head. Specifically: the bipartite graph attention network is used to perform local information exchange between task type feature units and work equipment feature units, and output updated task type feature units and work equipment feature units; the Transformer network is used to globally fuse the updated task type feature units and work equipment feature units with time slot feature units and global feature units to output a fused global state representation; the task head is used to output the state value of the successor state corresponding to the candidate scheduling action based on the fused global state representation.

[0026] Therefore, by using a bipartite graph attention network, targeted local information exchange can be carried out between compatible task types and operating devices, providing more accurate structured feature inputs for the subsequent global fusion of Transformers and improving the accuracy of state value estimation.

[0027] According to an optional embodiment of this application, the method further includes: obtaining compatibility relationships between all task types and operating devices related to the operating environment; and constructing a graph structure of the bipartite graph attention network based on the compatibility relationships, wherein edges are established only between compatible task type nodes and operating device nodes.

[0028] Therefore, the graph structure only contains physically meaningful information transmission paths, avoiding interference from invalid information and improving the computational efficiency of the network.

[0029] According to an optional embodiment of this application, the bipartite graph attention network employs a multi-head attention mechanism.

[0030] Therefore, different attention points can learn the dependency relationship between the work equipment and the task type from different perspectives, which enhances the model's ability to perceive multi-dimensional information.

[0031] According to a second aspect of this application, a method for training a machine learning model is provided, the machine learning model being used to perform the method described in the first aspect of this application, comprising: acquiring experience sample data, the experience sample data including at least: a defined state of a work environment, a scheduling action corresponding to the defined state, a subsequent state of the work environment after executing the scheduling action, and an immediate reward; representing the defined state and the subsequent state as heterogeneous feature units respectively; using the machine learning model, generating a state value of the defined state based on the heterogeneous feature units corresponding to the defined state, and generating a state value of the subsequent state based on the heterogeneous feature units corresponding to the subsequent state; calculating a training error based on the immediate reward, the state value of the defined state, and the state value of the subsequent state; and updating the internal parameters of the machine learning model based on the training error.

[0032] Therefore, the model only needs to use state estimation as the learning objective, and the agent's scheduling decisions are generated based on this. This significantly reduces the number of parameters and training complexity of the machine learning model, while improving convergence speed and generalization ability. Simultaneously, expanding the action space does not affect the model structure, enhancing the scalability of the scheduling scheme. Furthermore, this training method utilizes a reinforcement learning framework, eliminating the need for manually labeled data; it learns solely from experience samples generated through interaction with the work environment, making it particularly suitable for real-world work scenarios with complex constraints, frequent resource conflicts, and difficulties in obtaining the true value of optimal decisions.

[0033] According to a third aspect of this application, an electronic device is provided, the electronic device including a memory and a processor, the memory storing computer program instructions, which, when executed by the processor, enable the processor to perform the method according to the first and / or second aspects of this application.

[0034] According to a fourth aspect of this application, a computer program product is provided, comprising computer program instructions, wherein, when executed by a processor, the computer program instructions enable the processor to perform the method according to the first and / or second aspects of this application.

[0035] According to a fifth aspect of this application, a computer-readable storage medium is provided that stores computer program instructions, which, when executed by a processor, cause the processor to perform the method described according to the first and / or second aspects of this application. Attached Figure Description

[0036] The principles, features, and advantages of this application will be better understood below with reference to the accompanying drawings. The drawings include: Figure 1A flowchart is shown of a method for state estimation of a working environment according to an exemplary embodiment of this application; Figure 2 A schematic diagram illustrating the application of a method for state estimation of a work environment according to an exemplary embodiment of this application in a work environment task scheduling framework is shown. Figure 3 A schematic diagram illustrating the principle of representing the state information of the working environment as heterogeneous feature units according to an exemplary embodiment of this application is shown. Figure 4 A schematic diagram illustrating the principle of the generation process of a time slot feature unit according to an exemplary embodiment of this application is shown; Figure 5 A schematic diagram of the network structure of a machine learning model used according to an exemplary embodiment of this application is shown; and Figure 6 A flowchart illustrating a method for training a machine learning model according to an exemplary embodiment of this application is shown; and Figure 7 A block diagram of an electronic device according to an exemplary embodiment of this application is shown. Detailed Implementation

[0037] To make the technical problems to be solved, the technical solutions, and the beneficial technical effects of this application clearer, the application will be further described in detail below with reference to the accompanying drawings and several exemplary embodiments. It should be understood that the specific embodiments described herein are only for explaining this application and are not intended to limit the scope of protection of this application.

[0038] Before proceeding with the specific description, it should be noted that the terms "first," "second," etc., used in this document are only for distinguishing descriptions and should not be construed as indicating or implying relative importance, nor should they be construed as implicitly indicating the number of technical features indicated.

[0039] Figure 1 A flowchart illustrating a method for state estimation of a working environment according to an exemplary embodiment of this application is shown. In this embodiment, the method includes steps 101 to 103.

[0040] In step 101, status information related to a specified state of the working environment is obtained.

[0041] The work environment refers to systems or places that require task scheduling and resource allocation, such as industrial manufacturing environments (covering production, processing, assembly, etc., such as injection molding workshops, machining workshops, and electronic assembly lines), logistics and warehousing environments (such as automated storage and retrieval systems, AGV handling systems, and sorting centers), as well as food processing / distribution sites, pharmaceutical production workshops, etc.

[0042] In one embodiment, the operating environment is an industrial automation environment, meaning that each operational step related to production, processing, and assembly follows controllable and predictable operating parameters, and the state of the operating environment can be simulated within a predetermined accuracy range using a simulator. For example, in an injection molding workshop, the processing time, mold change time, and crane occupancy time of each piece of equipment are all known parameters, and the state of the operating environment (such as equipment idle time, occupancy status of work parts, and crane availability time) can be accurately calculated or predicted, thereby providing a reliable simulation basis for scheduling decisions.

[0043] The specified state of a working environment refers to, for example, the specific state of the working environment at a certain moment (e.g., a decision time step). This state includes at least the resource state, task state, and global state. For example, for a scheduled task, the specified state of the working environment can include the determined state before the execution of the scheduling action, the subsequent state after the execution of the scheduling action, and other intermediate states.

[0044] Task status includes, for example, task type (e.g., product type), delivery deadline, estimated operation time on each work machine, compatibility constraints with work machines and other shared auxiliary resources, outsourcing flag, and task switching type flag. A task refers to a work unit that needs to be executed by work machines in the work environment; it can be understood as a customer order, picking order, or production order. Each task corresponds to a task type (e.g., product type, part type, or process type) and includes a corresponding delivery deadline (i.e., the latest time the task must be completed) and estimated operation time on each work machine. Taking an injection molding workshop as an example, a task is a production order, specifying that the part type to be processed is A, the quantity is 300 pieces, the delivery deadline is 9:00 AM on July 3rd, and related process requirements (e.g., only compatible with machine M1 or machine M2, requiring mold m2). Tasks are usually added to the task queue (or order pool) in the order of their arrival time, or they can be sorted according to preset priority rules (e.g., delivery deadline or urgency level) and wait to be assigned to work machines for execution.

[0045] Resource status includes at least the relevant status of the work equipment, such as the earliest available time, the currently performing task, the currently assembled work parts, and efficiency information. Work equipment refers to physical resources that perform tasks in the work environment, such as injection molding machines, CNC machine tools, machining centers, robotic arms, autonomous mobile robots, conveyor belts, testing equipment, and packaging equipment. Typically, work equipment can only perform one task at a time, corresponding to one task type (e.g., product or part type). It can only begin performing the next task after completing the previous one. Its idle / busy status and earliest available time can be obtained in real time.

[0046] In addition, resource status may include information on shared auxiliary resources. Shared auxiliary resources refer to auxiliary resources that, while not directly used to actively perform operational tasks, may be needed in conjunction with the operational equipment when performing tasks. These resources are typically limited in number and may have conflicting uses. The status of shared auxiliary resources may further include the availability, assembly status, and earliest available time of shared operational parts, as well as the occupancy status and earliest available time of shared swapping resources (such as cranes).

[0047] Shared work accessories refer to process equipment that needs to be used in sync with the work equipment during operation, including but not limited to replaceable parts such as molds, fixtures, and cutting tools. These auxiliary resources are continuously occupied during task execution, and their availability and assembly status directly affect whether the task can be carried out normally.

[0048] Shared replacement resources refer to auxiliary equipment temporarily needed during task switching or the preparation phase before a task begins. Examples include cranes, preheating devices, and cleaning equipment used for changing work parts. These auxiliary resources are generally only used during the parts replacement operation and are released after the replacement is completed, without participating in the continuous processing of the task.

[0049] Taking an injection molding workshop as an example, shared operating parts include molds mounted on machines, while shared changeover resources include the single, non-seizing overhead crane required for mold replacement. The crane can only serve one piece of equipment at a time; other equipment must wait for it to be released before it can be used. The availability of shared auxiliary resources directly affects the estimated start time of the task, especially when changing operating parts.

[0050] The global state includes information such as the position of the current decision time step in the global time and the position of the time slot window, providing a complete context for scheduling decisions.

[0051] In one embodiment, the original form of the state information specifying the state can be, for example, structured data, or a matrix, graph, or other form of data representation.

[0052] In one embodiment, the state information of the specified state can be the real state (e.g., collected from the actual working environment), the simulated state derived by the simulator, or the state data generated by a mixture of the two.

[0053] In step 102, the state information of the specified state is represented as a heterogeneous feature unit, which includes at least: a task type feature unit, a time slot feature unit, and a resource feature unit.

[0054] In this context, a feature unit refers to encoding scattered information from a specified state of the operational environment into a unified vector representation. Its specific form can be a feature vector, an embedded representation, or other forms of numerical encoding. Multiple feature units arranged sequentially can form a token sequence.

[0055] "Heterogeneous" refers to the fact that the types of feature units are heterogeneous, meaning that different categories of entities (such as task types and operating equipment) correspond to different types of feature units, which are at least semantically distinct. Each feature unit might correspond to, for example, task type, time slot, operating equipment, and global state. Compared to directly inputting the raw state information of the operating environment into the machine learning model, the representation of heterogeneous feature units provides the model with a more compact and interpretable state representation, helping the model better understand the relationships and differences between different semantic categories of information.

[0056] In one embodiment, each task type feature unit corresponds to a task type, representing the workload distribution (i.e., time dimension distribution) of the same task type across different time intervals. Workload can be expressed as the expected work duration (e.g., in minutes), reflecting the time requirements of all tasks of that task type in different time periods. Furthermore, in some cases, workload can also be expressed as the number of products to be processed.

[0057] In one embodiment, each time slot feature unit corresponds to a future time interval, used to characterize the workload distribution (i.e., type dimension distribution) of each task type within the same time interval.

[0058] In one embodiment, task type feature units and time slot feature units represent the same set of demand information from different dimensions. Specifically, time slot feature units and task type feature units are, for example, perpendicular to each other and complementary in terms of workload distribution: the former is organized according to the time slot dimension, and the latter is organized according to the task type dimension. They describe the same set of demand data from different perspectives. This complementary representation enables the model to observe complete demand information from two vertical dimensions, simultaneously perceiving the resource competition situation in the time dimension and the demand trend in the task type dimension. This allows for more accurate identification of time bottlenecks and resource conflicts, effectively avoiding short-term local optimal decisions, and improving the long-term global performance of decisions.

[0059] In one embodiment, the resource feature unit may further include a job equipment feature unit and a global feature unit. Each job equipment feature unit corresponds to a job equipment, and is used to characterize, for example, the earliest available time of the job equipment, the task type of the current job, and the currently assembled job accessories. The global feature unit is used to characterize global information in the job environment, such as the availability status (availability, earliest available time) of shared auxiliary resources, the current time window position, global timing information, etc.

[0060] In one embodiment, there are usually multiple heterogeneous feature units of the above types, and therefore it can also be called a heterogeneous feature unit sequence.

[0061] In one embodiment, the original state information can be directly mapped to a fixed-dimensional feature vector through a trained neural network encoder (such as a multilayer perceptron, graph neural network, or Transformer network). Alternatively, the integrated discrete entity features (such as task type or operating equipment) can be encoded into a dense vector representation through a learnable embedding layer. Furthermore, structured state data (such as a demand matrix) can be extracted by rows or columns and mapped to corresponding feature units through rule-based mathematical transformations.

[0062] In step 103, a state value corresponding to the specified state is generated based on the heterogeneous feature units using a trained machine learning model.

[0063] In one embodiment, the machine learning model is a trained artificial neural network, such as a deep neural network. This artificial neural network may include multilayer perceptrons (MLPs), convolutional neural networks (CNNs), recurrent neural networks (RNNs), and / or graph neural networks (GNNs), etc. The following will combine... Figure 5 Describe a specific implementation method of a network structure.

[0064] In one embodiment, the machine learning model is a reinforcement learning model, whose training process is optimized based on reward signals obtained from interaction with the environment.

[0065] The specific meaning of the state value of a work environment can vary depending on the type and requirements of the work environment. In one embodiment, state value is used to characterize the level of long-term profitability that the work environment can obtain starting from a specified state. Here, "long-term profitability" does not refer only to the single-step profitability at the current decision time step, but rather to the overall profitability in the long term. For example, long-term profitability reflects, for instance, the comprehensive measure of the cascading effects that choosing a particular scheduling action may have on the current work situation and the future overall work situation.

[0066] In industrial manufacturing scenarios, state value is primarily used to characterize the overall manufacturing profit level. In other operational environments, state value can also be used to characterize scheduling performance indicators such as operational efficiency, resource utilization, and task delay risk in the future operational environment starting from a given state.

[0067] In one embodiment, state value can be used to evaluate the merits of candidate scheduling actions, for example, to provide a basis for agent action selection during scheduling decision-making. In another embodiment, state value can also be used to evaluate the overall operational status of the work environment, serving as a reference indicator for production plan adjustments or resource optimization.

[0068] In one embodiment, the state estimation method can be executed continuously and iteratively. For example, a corresponding state value estimation is performed once at each decision time step. Alternatively, the estimation results from each decision time step are continuously used to guide the selection of scheduling actions, and the execution result of the previous decision time step will affect the working environment state of the next decision time step, thus forming a closed-loop iterative process of decision-making and state evolution.

[0069] Figure 2 A schematic diagram illustrating the application of a method for state estimation of a job environment according to an exemplary embodiment of this application in a job environment task scheduling framework is shown.

[0070] exist Figure 2 The diagram illustrates the basic structure of a task scheduling framework for a work environment. This framework, as a whole, is used at each decision time step to select a target scheduling action 24 from candidate scheduling actions 21 and issue it to the work environment for execution. In this framework, state estimation serves to provide a basis or evaluation standard for the selection and decision-making of scheduling actions.

[0071] Specifically, the framework includes, for example, a simulator 33, a preprocessing module 34, and an agent 32. The agent 32, for example, is built based on a reinforcement learning framework and further includes a machine learning model 321 and a rule-based decision-making module 322.

[0072] like Figure 2 As shown, the operating environment is first determined. Candidate scheduling actions 21 are listed below. Each candidate scheduling action 21 represents assigning a task type to the focus job device or keeping the focus job device idle. Multiple candidate actions typically exist within a single decision time step. Action Space The size is determined by the number of potentially executable task types in the job environment, and is independent of the number of specific task instances in the task queue. Candidate scheduling actions also include idle actions, where the focused job device remains temporarily idle. The original action space is, for example, [example of action space]. .

[0073] For example, if there are four task types available for work in the work environment: P1, P2, P3, and P4, and the focus work device is M1 (i.e., in the current state, work device M1 is expected to be idle the earliest), then candidate scheduling actions may include: assigning task type P1 to M1, assigning task type P2 to M1, assigning task type P3 to M1, assigning task type P4 to M1, and including an idle action (i.e., not assigning any task type, keeping M1 idle). Candidate scheduling actions may also be a subset of the above set (e.g., only considering task types with job requirements in the current task queue), and do not necessarily have to include all task types.

[0074] In one embodiment, once a candidate scheduling action is selected, it can be mapped to the next pending task instance of the same task type in the task queue for actual execution. For example, candidate scheduling actions... To assign task type P1 to job device M1, the next pending task job1 of task type P1 is retrieved from the task queue and assigned to job device M1 for execution. The subsequent simulation process of simulator 33 can be carried out based on this.

[0075] A focal job is the earliest available job selected from all available jobs at the current decision time step. Specifically, it is determined by comparing the earliest available times of each job (i.e., the estimated completion time of the last task scheduled on that job). If multiple jobs have the same earliest available time, one is selected as the focal job according to a pre-defined priority rule (e.g., job number order, work efficiency, historical workload, etc.). The focal job is the target of the current scheduling decision, and all allocation relationships in candidate scheduling actions are directed towards this job.

[0076] After obtaining candidate scheduling actions 21, the simulator 33 is used to simulate the execution of each candidate scheduling action in order to determine the subsequent state of the job environment after the execution of each candidate scheduling action. and instant rewards In this embodiment, the successor state This serves as the designated state of the operating environment, and subsequent state value estimation is performed on it.

[0077] In one embodiment, the working environment is an injection molding workshop, the focal operation equipment is one of multiple injection molding machines present in the working environment, and the task type is the type of product to be injection molded corresponding to the production order. The working environment also includes an overhead crane for performing mold lifting operations when the injection molding machine needs to change molds. The number of overhead cranes is limited (e.g., one unit), and when multiple injection molding machines simultaneously require mold changing, they must queue up and wait for the overhead crane.

[0078] In one embodiment, when simulating the execution of candidate scheduling actions using simulator 33 and / or making decisions using agent 32, the product switching type corresponding to two production orders processed consecutively on the same injection molding machine is considered. This product switching type determines whether a mold change is required for the injection molding machine and whether a crane needs to be used. Accordingly, mold changes and crane waiting incur additional time and economic costs, which will affect the determination of the subsequent state of the operating environment and may also affect the calculation of profits or costs corresponding to the candidate scheduling actions.

[0079] For example, when simulator 33 simulates the execution of candidate scheduling actions, if the production order involved in the candidate scheduling action requires different molds from the production order currently being executed on the same injection molding machine, the estimated start time of the candidate scheduling action must wait until the crane becomes available before the mold change operation can be performed. This delay further affects the estimated completion time of each task, the earliest available time of the injection molding machine, the occupancy status of the injection mold, and the available time of the crane, thus affecting the determination of subsequent states.

[0080] For example, when agent 32 makes decisions, the immediate reward it considers also depends on the switching costs corresponding to the product switching type. For instance, the switching cost is lowest if neither the mold nor the material is changed, followed by the switching cost of changing only the material, then the switching cost of changing the mold is higher, and the switching cost of changing both the mold and the material is the highest. This reflects the impact of different switching types on economic benefits during the decision-making process.

[0081] In one embodiment, the immediate reward reflects the immediate, one-step economic benefit of the task itself during the execution of the scheduled action (i.e., the entire period from start to finish of the task on the focus device). For example, the immediate reward can be determined based on the task's gross profit, inventory holding costs, and task switching costs.

[0082] For example, the instant reward is calculated using the following formula: in: The gross profit of the task k to be processed corresponding to the scheduling action is the revenue obtained by the task through internal production minus the direct production cost. Let k be the time when task k is completed ahead of schedule. Cost of holding inventory per unit time This refers to inventory holding costs. Inventory holding costs represent the expenses incurred due to storage and capital tied up while waiting for delivery after a task is completed ahead of schedule. A higher value indicates that the earlier the task is completed, the longer the inventory is held, and the higher the warehousing costs. By introducing this cost term into the training objective, the model can be guided to complete tasks closer to the delivery deadline, rather than completing them too early, thereby reducing inventory backlog and optimizing capital tied up and warehousing efficiency. Let be the sequence-dependent switching cost from the previous task j to the current task k. This cost depends on the switching type; for example, the cost is zero when no switching is required; the cost is low when only materials are changed (e.g., cleaning costs); the cost is high when only work accessories are changed (e.g., molds); and the cost is highest when both materials and work accessories are changed. By introducing switching costs into the training objective, the model can be guided to reduce unnecessary switching in scheduling decisions, thereby reducing job transition overhead and improving overall job efficiency.

[0083] Subsequently, the successor states corresponding to each candidate scheduling action are... Provided to preprocessing module 34 to represent it as heterogeneous feature unit 60, and with corresponding immediate rewards. Together, they are provided to agent 32. After decision-making by agent 32, they ultimately originate from the original action space. Select an action (e.g.) As a target scheduling action .

[0084] Machine learning model 321 successor state The heterogeneous feature unit is used as input, and the subsequent state is output. Corresponding state value This value reflects the state from the successor state. Starting point: the cumulative revenue level that the operating environment can achieve in the future.

[0085] In one embodiment, machine learning model 321 targets each candidate scheduling action. Generate its successor state one by one. Corresponding state value The output dimension is fixed at 1 and is independent of the size of the action space.

[0086] After completing the above processing for all candidate scheduling actions one by one, the rule decision module 322 summarizes the state values ​​corresponding to all actions and makes a final selection. For example, for each candidate scheduling action, the rule decision module 322, based on the subsequent state... State value With instant rewards Calculate the level of long-term returns And select the action with the highest long-term return as the target scheduling action. .

[0087] For example, the rule decision module 322 calculates the long-term revenue level corresponding to each candidate scheduling action according to the following formula. : in: This is the immediate reward corresponding to the execution of the candidate scheduling action, used to measure the single-step benefit brought by the action; This is the discount factor, and its value range is usually [range missing]. This is used to balance the importance of immediate rewards with future benefits. For example, The possible value is 1; For the successor state The corresponding state value is used to measure the cumulative reward level starting from that subsequent state.

[0088] For example, the rule decision module 322 selects target scheduling actions according to the following rules. : in, For the action space of candidate scheduling actions, Schedule actions for the target.

[0089] Within the aforementioned framework, by introducing a deterministic simulator 33, the calculation of action-to-state transitions can be separated from the decision-making calculations of agent 32, reducing the learning burden on machine learning model 321 and improving the stability and interpretability of decisions. Thus, agent 32 does not need to implicitly learn the environmental state transition patterns through extensive trial and error, but can focus solely on making subsequent decisions based on state information. Figure 3 A schematic diagram illustrating the principle of representing the state information of the working environment as heterogeneous feature units according to an exemplary embodiment of this application is shown.

[0090] like Figure 3 As shown, the specified state of the job environment (e.g., the subsequent state after the execution of a candidate scheduled action). These features are provided to preprocessing module 34 and represented as heterogeneous feature units 61, 62, 63, and 64. These heterogeneous feature units 61, 62, 63, and 64 are then provided to machine learning model 321 to generate the subsequent state. Corresponding state value .

[0091] The preprocessing module 34 can be implemented, for example, as a rule-based and / or machine learning-based encoder. For each entity type in the state information, the preprocessing module 34 extracts the corresponding raw feature information and then maps it to a fixed dimension, for example, through a learnable linear projection layer, to obtain the corresponding feature units. Different entity types can use independent or shared projection parameters. This process transforms the unstructured raw state information into a structured sequence of fixed-dimensional vectors, providing a standardized input form for subsequent machine learning models.

[0092] Heterogeneous feature units 61, 62, 63, and 64 include at least: task type feature unit 61, time slot feature unit 62, and resource feature unit. The resource feature unit may further include operating equipment feature unit 63 and global feature unit 64.

[0093] In one embodiment, there are typically multiple feature units of each type, and therefore it can also be understood as a sequence of feature units. For example, a task-type feature unit sequence can be represented as follows: Each task type feature unit (such as Each time slot corresponds to a task type (e.g., product types A, B, C). The sequence of time slot feature units can be represented as... Each time slot feature unit (e.g.) Each slot corresponds to a future time slot (e.g., slot1, slot2, slot3). The sequence of characteristic units of the operating equipment can be represented as... Each feature unit corresponds to a specific work device. A single global feature unit, denoted as G, is typically used to represent global state information in the work environment, such as the availability of shared auxiliary resources, the current time window position, and global timing information. By encoding different categories of entities into independent sequences of feature units, the structured semantics of the state information can be preserved, providing a clear and standardized input representation for subsequent machine learning models.

[0094] For example, assuming there are three task types A, B, and C in the task queue, and the future time window is divided into three time slots, slot1, slot2, and slot3, the corresponding task type feature unit sequence can be represented as: in, This indicates that the workload of task type A in time slots 1, 2, and 3 is 200 minutes, 60 minutes, and 100 minutes, respectively. and And so on.

[0095] By extracting the raw state information of the work environment into task type feature units, machine learning models can obtain semantic information about the future demand distribution related to task types. This allows them to anticipate the resource demand intensity of various types of tasks at different time periods, which helps to identify time bottlenecks, rationally arrange production sequences, and improve adaptability to dynamic demand changes.

[0096] For example, assuming there are three task types A, B, and C in the task queue, and the future time window is divided into three time slots, slot1, slot2, and slot3, the corresponding time slot feature unit sequence can be represented as: in, This indicates that within slot 1 of the time slot, the workloads for task types A, B, and C are 200 minutes, 150 minutes, and 50 minutes, respectively. and And so on.

[0097] For example, for the working device M1, its corresponding working device feature unit 63 can be represented as follows: , This is the earliest available time for the work equipment M1. This indicates the type of task the device is currently performing. This indicates the currently assembled work accessory (e.g., a mold). The same logic applies to other work equipment in the work environment, such as M2, M3, etc.

[0098] For example, the global feature unit 64 can be represented as .in Indicates the availability of shared auxiliary resources (e.g., whether cranes are available). Indicates the earliest available time of the shared auxiliary resource. This indicates global time information.

[0099] Through the equipment feature unit 63 and the global feature unit 64, the machine learning model 321 can perceive the occupancy status of each piece of equipment, the assembly status of equipment parts, the competition and occupancy of shared auxiliary resources, and the position of the current decision time step within the time window. This information provides crucial basis for avoiding resource conflicts and rationally arranging task order in scheduling decisions.

[0100] In one embodiment, before inputting the heterogeneous feature units 61, 62, 63, and 64 into model 321, type embeddings can be superimposed on them to enable machine learning model 321 to distinguish between different categories of feature units (e.g., task type feature units, time slot feature units, work equipment feature units, and global feature units). The type embeddings can use fixed orthogonal vectors and are not updated during training.

[0101] For example, task type feature units can be overlaid. , is the superposition of time slot feature units , for superimposed feature units of operating equipment , is a superposition of global feature units The types are orthogonal to each other to ensure that Model 321 can clearly distinguish inputs with different semantics.

[0102] In one embodiment, before inputting the heterogeneous feature units 61, 62, 63, and 64 into the machine learning model 321, positional encoding may be superimposed on the time slot feature units 62. This positional encoding may be used, for example, to characterize the chronological order of each time slot in a future time window.

[0103] For example, linear position encoding can be used, assigning a learnable linear position vector to each time slot and adding it element-wise to the feature units of each time slot. This allows model 321 to perceive the sequence of time slots, thereby more accurately understanding the distribution of workload on the time axis. Furthermore, in some cases, sine / cosine position encoding or relative position encoding can also be used.

[0104] It should be noted that Figure 3 The number and sequence length of the heterogeneous feature units shown are merely illustrative; in actual applications, the number of various feature units may be more or less.

[0105] Figure 4 A schematic diagram illustrating the principle of the generation process of a time slot feature unit according to an exemplary embodiment of this application is shown.

[0106] like Figure 4 As shown, a demand matrix 70 is first constructed based on the task information in the subsequent states. The rows of the demand matrix 70 correspond to task types P1, P2, P3, and P4 (e.g., product types), and the columns correspond to time slots T1, T2, T3, and T4. The demand matrix 70 is used to structure the workload of each task type (P1, P2, P3, and P4) in different time slots T1, T2, T3, and T4.

[0107] To construct the demand matrix 70, the future time window first needs to be divided into multiple time slots T1, T2, T3, and T4. These future time windows can be dynamically updated using a sliding window mechanism, with window lengths determined in days, weeks, or months. Each time, only tasks within that time window from the current moment are released from the task queue / task pool and injected into the demand matrix 70; tasks exceeding the time window are not included. As time progresses, expired time slots are removed, and corresponding future time slots are added. Figure 4 As shown in the example, the future time window, which is based on days, is divided into 4 time slots, each time slot corresponding to a 6-hour time interval, such as 0-6:00, 6-12:00, 12-18:00, and 18-24:00.

[0108] During the injection process, the time slot to which each task should belong is determined based on its delivery deadline. Then, for example, the workload can be calculated according to the task type corresponding to each task and injected into the corresponding task type of the corresponding time slot, thereby constructing a complete demand matrix 70. For different tasks of the same task type within the same time slot, their workloads are aggregated and accumulated, no longer distinguishing between specific task instances, thus integrating discrete task information into the total workload of each task type within each time slot.

[0109] In one embodiment, assume a task of type P1 has a delivery deadline of 9:00 AM and an estimated job duration of 1 hour (i.e., a workload of 60 minutes). Since 9:00 AM falls within the 6:00-12:00 time slot T2, and the task needs to be completed within the delivery deadline, the 60-minute workload of this task is injected into time slot T2. After all tasks have been injected, the workloads of all tasks of type P1 within time slot T2 are aggregated, resulting in a total workload of 200 minutes, indicating that task type P1 within time slot T2 requires a total of 200 minutes of work time.

[0110] In one embodiment, tasks can be directly injected into the time slots where their delivery deadlines are located, meaning each task enters only one time slot. In this case, the workload within each time slot can be characterized, for example, by the number of tasks falling into it.

[0111] In another embodiment, when injecting tasks into time slots, the estimated operation time of the task or the corresponding number of products to be processed can also be considered. Therefore, a task may fall into exactly one time slot or it may span multiple adjacent time slots. For example, if the delivery deadline for a task is 7:00 AM and the estimated operation time is 180 minutes, then this task may need to continuously occupy the operation equipment resources in two adjacent time slots T1 and T2. For such tasks that span time slots, the workload can be split according to the estimated operation time in each time slot T1 and T2, and injected into the corresponding time slots respectively.

[0112] Based on the constructed demand matrix 70, the original feature representation of the time slot feature unit can be generated by taking each column of the matrix. This means that within time slot T1, the workloads of task types P1, P2, P3, and P4 are 200, 0, 100, and 30 (in minutes, for example). If necessary, the extracted original feature representations can be further transformed through a learnable linear projection layer to map them to a fixed dimension, thereby unifying the vector dimensions of feature units from different sources.

[0113] Accordingly, the original feature representation of the task type feature unit can be generated by extracting each row of the demand matrix 70. This means that for task type P1, the workload in time slots T1, T2, T3, and T4 is 200, 220, 100, and 150 respectively (in minutes, for example).

[0114] In the above manner, the task type feature unit and the time slot feature unit encode the demand information from the time dimension and the type dimension, respectively, and together provide the machine learning model with comprehensive spatiotemporal demand distribution information.

[0115] In one embodiment, the granularity of the time slots corresponding to the time slot feature units can be dynamically adjusted. For example, the length or number of time slots can be dynamically adjusted according to the demand distribution characteristics of the work environment. For instance, when future demand distribution is relatively concentrated, a finer time slot granularity (such as one time slot per hour) can be used to improve the accuracy of demand description. In addition, the number of time slots can be adaptively adjusted according to the density of task arrivals, increasing the number of time slots during high-load periods to finely characterize resource competition, and reducing the number of time slots during low-load periods to compress the state space. This dynamic adjustment mechanism enables the method of this application to flexibly adapt to the scheduling requirements of different work environments, achieving a balance between computational accuracy and state complexity.

[0116] Figure 5 A schematic diagram of the network structure of a machine learning model used according to an exemplary embodiment of this application is shown. Figure 5 In the illustrated embodiment, Figures 1 to 4 The machine learning model used in the method shown is, for example, a trained artificial neural network, specifically including a bipartite graph attention network 71, a Transformer network 72, and a task head 73.

[0117] The input to the bipartite graph attention network 71 is the task type feature unit 61 and the work equipment feature unit 63. The network 71 is used to perform local information exchange between these two types of feature units 61 and 63 to output updated task type feature units and work equipment feature units 61' and 63'.

[0118] In one embodiment, the graph structure of the bipartite graph attention network 71 is constructed as follows: each task type and each operating device are treated as two types of nodes, the compatibility relationships between each task type and each operating device in the operating environment are obtained, and edges are established based on these compatibility relationships. For example, edges are only built between compatible task type nodes and operating device nodes, thereby forming the graph structure of the bipartite graph attention network.

[0119] like Figure 5As illustrated, the graph contains four task type nodes P1, P2, P3, and P4, and three job device nodes M1, M2, and M3. Edges are established only between compatible task type-device pairs, rather than full connections. For example, task type node P1 is only connected to job device nodes M1 and M2, indicating that task type P1 can only be performed on job devices M1 and M2, and not on M3. This graph structure design based on compatibility constraints ensures that attention computation is performed only between physically feasible node pairs, avoiding the transmission of invalid information and improving the computational efficiency of the network.

[0120] In the bipartite graph attention network 71, a multi-head attention mechanism is used, where each node aggregates the features of its neighboring nodes and updates them using attention weights. These attention weights are dynamically calculated based on the feature similarity and compatibility between nodes, enabling the network to automatically learn the importance of different task types and work equipment nodes. Through this layer, task type nodes can perceive the load status and processing capacity of compatible work equipment; work equipment nodes can also perceive the urgency and workload of the tasks to be processed.

[0121] For example, the bipartite graph attention network 71 uses the following node update formula: in, For nodes The original feature vector; For nodes Updated feature vector; It is a learnable linear transformation weight matrix used to transform the input features; The attention weights represent the distances from node j to node j. The degree of importance or strength of the relationship; For nodes The set of neighboring nodes (i.e., nodes with the node) (Nodes with edges) It is a non-linear activation function (such as LeakyReLU or ELU).

[0122] Through this residual connection, nodes retain their own characteristics while absorbing relevant information from their compatible neighboring nodes, thus achieving effective exchange of local information.

[0123] The aforementioned attention mechanism enables the model to assign different levels of importance to different neighboring nodes, thereby selectively focusing on neighboring nodes that are more important to the current node during the information aggregation process.

[0124] In one embodiment, the bipartite graph attention network 71 may employ a multi-head attention mechanism, concatenating or averaging the outputs of multiple attention heads as the final node update features. Through training, different attention heads can learn the importance relationships between nodes from different feature subspaces or relational perspectives. For example, one head may focus on the idle time of the operating equipment, another head may focus on the cost of changing task types, and yet another head may focus on operational efficiency, etc.

[0125] Thus, the machine learning model can autonomously identify which neighboring nodes are more critical to the current node's value estimation (e.g., idle machines are more important than busy machines, and low-cost matches are more important than high-cost matches), thereby dynamically adjusting the allocation of attention weights to achieve selective information aggregation. This mechanism enables the network to adaptively learn the importance relationships between nodes from the data without relying on manual annotation.

[0126] In one embodiment, after the bipartite graph attention network 71 completes the modeling of the relationship between task type and operating equipment, it forms a splicing sequence 65 with the updated task type feature unit 61' and operating equipment feature unit 63', together with the time slot feature unit 62 and the global feature unit 64, and inputs it into the Transformer network 72 for global context fusion.

[0127] Through the self-attention and multi-head attention mechanisms of the Transformer network 72, the feature units in the concatenated sequence 65 can interact with each other to obtain the fused global state representation 66. This global state representation 66 is then output to the task head 73, which, for example, is a feedforward neural network (such as a multilayer perceptron), to map the high-dimensional vector form of the global state representation 66 to the final task output, i.e., the state value of the successor state. .

[0128] In one embodiment, a learnable CLS feature unit can also be added at the beginning of the splicing sequence 65. This CLS feature unit interacts with all other feature units in the sequence through a self-attention mechanism, gradually aggregating the global state information of the entire working environment. Its corresponding output vector can be used as the global state representation 66 and output to the task head 73 for state value generation.

[0129] It should be noted that, in combination Figure 5 The number of nodes, connections, network layers, and task header structure described are all exemplary and can be adjusted according to the specific working environment in actual applications.

[0130] Figure 6A flowchart illustrating a method for training a machine learning model according to an exemplary embodiment of this application is shown. The method includes steps 601 to 605.

[0131] In step 601, experience sample data is obtained, which includes at least: a defined state of the work environment, a scheduling action corresponding to the defined state, the subsequent state of the work environment after the scheduling action is executed, and the immediate reward obtained in the work environment after the scheduling action is executed.

[0132] The determined state and the scheduling action, for example, correspond to the same decision time step t. That is, the scheduling action is one of the candidate decisions generated in the determined state, representing either assigning a task type to the focus job equipment or keeping the focus job equipment idle. Accordingly, the subsequent state can be understood as the state of the work environment after the scheduling action is executed.

[0133] Immediate reward can be understood as a feedback signal that the work environment returns immediately after a scheduling action is executed, used to evaluate the single-step benefit brought about by the execution of that scheduling action. Immediate reward is usually determined based on factors such as task gross profit, inventory holding costs, and task switching costs, reflecting the instantaneous impact of the scheduling action on the economic indicators of the work environment.

[0134] In one embodiment, empirical sample data can be represented as a quadruple. ,in, The working environment is in a deterministic state at decision time step t. This is a scheduling action generated under this defined state. The immediate reward returned by the job environment after this scheduling action is executed. This refers to the subsequent state of the working environment.

[0135] In one embodiment, experience sample data can be obtained either through simulation using a simulator of the work environment or through direct interaction with the real work environment.

[0136] In simulation mode, the simulator can quickly calculate the subsequent state and immediate reward based on deterministic states and scheduled actions according to deterministic rules, thereby efficiently generating a large number of samples for training.

[0137] In real-interaction mode, scheduling actions can also be sent to the actual working environment for execution, and real-time rewards can be calculated by collecting the actual subsequent status.

[0138] In step 602, the determined state and the subsequent state are represented as heterogeneous feature units. Heterogeneous feature units include, for example, at least task type feature units, time slot feature units, and resource feature units. The specific method for generating heterogeneous feature units can be performed in a similar manner as described above, and will not be repeated here.

[0139] In step 603, a machine learning model is used to determine the state. The corresponding heterogeneous feature units generate a deterministic state. State value Based on successor state The corresponding heterogeneous feature units generate successor states. State value .

[0140] Reflecting from a definite state Starting point: the cumulative revenue level that the operating environment can achieve in the future. Reflecting from the successor state Starting point: the cumulative revenue level that the operating environment can achieve in the future. Ideally, the difference between the two is precisely the execution scheduling action. The single-step revenue brought That is, satisfying In the context of scheduling decisions, the state value of a given state can also be understood as the reflection of the long-term return level corresponding to a certain scheduling action being selected. It constitutes an assessment measure of the economic value that a scheduling action can create.

[0141] In step 604, the training error is calculated based on the immediate reward, the state value of the determined state, and the state value of the subsequent state.

[0142] In one embodiment, the timing difference error can be calculated. And based on this time-series difference error Constructing the loss function According to the loss function Update the parameters of the machine learning model. For example, time-series difference error. Defined as: in, This is the discount factor, and its value range is usually [range missing]. This is used to balance the importance of immediate rewards and future rewards. For example, The possible value is 1.

[0143] This timing difference error This reflects the state value estimation of a given state by a machine learning model. The target value integrates environmental feedback and subsequent state value estimation. The deviation between them. The value estimate of the determined state is iteratively updated by utilizing the value estimate of the successor state. This allows for single-step learning of the model.

[0144] For example, loss function Mean square error can be used: By minimizing this loss The output of the machine learning model gradually approximates the true expected long-term return. With each update of the model parameters, the value estimate of the current state is adjusted one step towards a more accurate target value.

[0145] In step 605, the internal parameters of the machine learning model are updated based on the training error. Using temporal difference learning, the model can update its parameters online and incrementally. Through iterative training with a large number of empirical samples, the model gradually converges, thereby accurately estimating the long-term return level under any state of the working environment.

[0146] In one embodiment, a semi-gradient temporal differencing approach is used to optimize the intrinsic parameters of the machine learning model. Specifically, when calculating the gradient of the loss function, the target value can be... As a fixed optimization objective, the state value is determined solely through gradient descent updates. The corresponding model parameters. By repeatedly executing the above steps, the machine learning model gradually learns the mapping from state to value estimation, thereby providing accurate state value assessment for scheduling decisions.

[0147] In one embodiment, training of the machine learning model is stopped when a preset convergence condition is met. The preset convergence condition may include, for example, the change in the loss function value over a series of iterations being lower than a preset threshold, or reaching a preset maximum number of training steps, or the absolute value of the temporal difference error remaining within a small range for an extended period.

[0148] In one embodiment, the training sample data should cover different months and combinations of different task types (e.g., demand structures of different product combinations) to improve the model's generalization ability under different operating environments. By introducing diverse demand distributions during the training phase, the machine learning model can learn more robust state-value mapping relationships, avoiding overfitting to specific months or specific product combinations. For example, actual data from multiple production months (such as August, October, and November) can be collected, each month having different distributions of order types, demand quantities, and delivery deadlines, and these can be mixed together as a training set.

[0149] Figure 7A block diagram of an electronic device according to an exemplary embodiment of this application is shown.

[0150] Electronic device 10 includes processor 11 and memory 12. Memory 12 (such as a hard disk, RAM, flash memory, or other computer-readable storage medium) stores computer program instructions. When processor 11 (e.g., a central processing unit / CPU, microprocessor, digital signal processor / DSP, or other general-purpose processor) executes these instructions, it can implement methods for estimating the state of the operating environment and / or methods for training machine learning models. The specific implementation details of these methods have been described in detail above and will not be repeated here.

[0151] The electronic device 10 can be deployed on a local terminal device or on a server (e.g., a cloud platform). The machine learning model used in the aforementioned method can be stored in the local memory of the electronic device 10 or in other locations, such as by calling a model service deployed in the cloud through an application programming interface.

[0152] The electronic device 10 may also include input / output interfaces to visualize relevant status estimates, scheduling decision results, and overall work plans when necessary. For example, such visualizations may be output in the form of charts, tables, or text.

[0153] It should be understood that the methods of the various embodiments of this application can be implemented by computer programs / software. This software can be loaded into the processor's memory and, when run, is used to execute the methods according to the various embodiments of this application.

[0154] It should be understood that the same or similar parts between the various embodiments in this specification can be referred to each other, and each embodiment focuses on describing the differences from other embodiments. In particular, for the apparatus embodiments, since their control logic basically corresponds to that of the method embodiments, the description is relatively brief, and relevant parts can be referred to the description of the method embodiments.

[0155] According to another embodiment of this application, a computer program product including computer program instructions is provided, the computer program instructions being configured to perform the methods according to various embodiments of this application when the computer program product is run on a computer or stored on a computer-readable storage medium (such as a CD-ROM). The machine-readable storage medium is, for example, an optical storage medium or a solid-state medium supplied together with or as part of other hardware.

[0156] Although specific embodiments have been described above, these embodiments are not intended to limit the scope of this application, even when only a single embodiment is described with respect to a particular feature. The feature examples provided in this application are intended to be illustrative and not limiting, unless otherwise stated. In practice, multiple features may be combined with each other as needed and where technically feasible. In particular, features from different embodiments may also be combined with each other. Various substitutions, modifications, and alterations are conceived without departing from the spirit and scope of this application.

Claims

1. A method for estimating the state of a work environment, comprising: Obtain status information related to a specified state of the operating environment; The state information of the specified state is represented as a heterogeneous feature unit, which includes at least: a task type feature unit, a time slot feature unit, and a resource feature unit; as well as Using a trained machine learning model, a state value corresponding to the specified state is generated based on the heterogeneous feature units.

2. The method according to claim 1, wherein, The state value is used to characterize the level of long-term benefits that the working environment can obtain from the specified state.

3. The method according to claim 1 or 2, wherein, Each time slot feature unit is used to characterize the workload distribution of each task type within the same time interval.

4. The method according to any one of claims 1 to 3, wherein, Each task type feature unit is used to characterize the workload distribution of the same task type in different time intervals.

5. The method according to any one of claims 1 to 4, wherein, The method includes: Divide the future time window into multiple time slots; Based on the delivery deadline of each task, determine the time slot to which each task should belong; Based on the task type, the workload of each task is injected into the corresponding time slot, forming a demand matrix; and The task type feature unit is generated based on the rows of the demand matrix, and / or the time slot feature unit is generated based on the columns of the demand matrix.

6. The method according to any one of claims 1 to 5, wherein, The method further includes: The time slot feature units are superimposed with position codes, which are used to characterize the chronological order of each time slot in a future time window; and / or The heterogeneous feature unit superposition type is embedded to distinguish different categories of feature units.

7. The method according to any one of claims 1 to 6, wherein, The resource feature unit includes an operating equipment feature unit and a global feature unit, wherein: Each work equipment feature unit corresponds to one work equipment, including the earliest available time of the work equipment, the current task type, and the currently assembled work accessories; The global feature unit includes the availability of shared auxiliary resources, the earliest available time, and global time information.

8. The method according to any one of claims 1 to 7, wherein, The method further includes: Obtain candidate scheduling actions under a defined state of the work environment. Each candidate scheduling action represents assigning a task type to the focus work equipment in the work environment or keeping the focus work equipment idle. Using a simulator of the work environment, simulate the execution of each candidate scheduling action to determine the subsequent state and immediate reward of the work environment after the execution of each candidate scheduling action. The specified state of the operating environment includes at least the subsequent state.

9. The method according to claim 8, wherein, The method further includes: Using the machine learning model, the state value corresponding to the subsequent state is generated based on the subsequent state of the working environment; With the help of the rule-based decision-making module, for each candidate scheduling action, the long-term benefit level corresponding to the candidate scheduling action is determined based on the state value of the successor state and the immediate reward, and the candidate scheduling action with the highest long-term benefit level is selected as the target scheduling action.

10. The method according to claim 8 or 9, wherein, The operating environment is an injection molding workshop, the focal operation equipment is one of the multiple injection molding machines present in the operating environment, and the task type is the type of product to be injection molded corresponding to the production order.

11. The method according to claim 10, wherein, When simulating the execution of the candidate scheduling action using the simulator, the product switching type corresponding to two production orders processed continuously on the same injection molding machine is considered. The product switching type determines whether it is necessary to change the mold for the injection molding machine and whether it is necessary to occupy the crane.

12. The method according to any one of claims 1 to 11, wherein, The machine learning model includes a bipartite graph attention network, a Transformer network, and a task head, specifically: The bipartite graph attention network is used to perform local information exchange between task type feature units and work equipment feature units, and output updated task type feature units and work equipment feature units. The Transformer network is used to globally fuse the updated task type feature unit and job device feature unit with the time slot feature unit and global feature unit to output the fused global state representation. The task header is used to output the state value of the successor state corresponding to the candidate scheduling action based on the fused global state representation.

13. The method of claim 12, wherein, The method further includes: Obtain the compatibility relationships between all task types and operating equipment related to the operating environment; and The graph structure of the bipartite graph attention network is constructed based on the compatibility relationship, wherein edges are established only between compatible task type nodes and job device nodes.

14. The method according to claim 12 or 13, wherein, The bipartite graph attention network employs a multi-head attention mechanism.

15. A method for training a machine learning model, the machine learning model being used to perform the method according to any one of claims 1 to 14, comprising: Acquire experience sample data, which includes at least: a defined state of the work environment, a scheduling action corresponding to the defined state, the subsequent state of the work environment after the scheduling action is executed, and an immediate reward. The determined state and the subsequent state are respectively represented as heterogeneous feature units; Using the machine learning model, the state value of the determined state is generated based on the heterogeneous feature unit corresponding to the determined state, and the state value of the subsequent state is generated based on the heterogeneous feature unit corresponding to the subsequent state. The training error is calculated based on the immediate reward, the state value of the determined state, and the state value of the subsequent state. The internal parameters of the machine learning model are updated based on the training error.

16. An electronic device comprising a memory and a processor, the memory storing computer program instructions which, when executed by the processor, enable the processor to perform the method according to any one of claims 1 to 15.

17. A computer program product comprising computer program instructions, wherein, When executed by a processor, the computer program instructions enable the processor to perform the method according to any one of claims 1 to 15.

18. A computer-readable storage medium storing computer program instructions that, when executed by a processor, cause the processor to perform the method according to any one of claims 1 to 15.