A process scheduling optimization method and system for quality feedback delay in intelligent manufacturing
Patent Information
- Application Number
- CN202610841159.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-11
- Publication Date
- 2026-09-18
AI Technical Summary
该方法在不改变外在质量反馈定义与评测口径的前提下,为延迟反馈条件下的调度策略优化提供结构敏感且可控衰减的辅助学习驱动支撑,系统性解决了质量反馈延迟条件下的长期价值识别困难与跨层信用分配难题
[0020](1) Solving the problem of delayed quality feedback and improving long-term value recognition: In scenarios where quality feedback signals need to undergo multiple layers of processing and cross-layer transmission before they can be concentrated, traditional scheduling methods cannot obtain effective immediate value feedback for a large number of intermediate decisions. The strategies tend to be biased towards short-sighted behavior and ignore decision paths that require long-term cross-layer advancement to realize quality benefits. This invention obtains structurally sensitive and semantically consistent state representations in the representation space through macroscopic structural feature construction and comparative representation learning. In the latent space, it generates intrinsic auxiliary learning signals oriented towards structural novelty through online clustering pseudo-counting. This provides a continuously available and controllably decaying auxiliary learning driver for policy optimization under delayed feedback conditions. This enables the agent to effectively distinguish the long-term potential differences of different intermediate paths in the long prefix interval before the milestone quality feedback is triggered. It avoids misjudging early but insufficient procedural rewards as effective guidance and staying on suboptimal paths, which significantly improves the training efficiency and convergence stability in delayed feedback environments.
Smart Images

Figure CN122776745A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent manufacturing and reinforcement learning, and specifically relates to a process scheduling optimization method and system for quality feedback delay in intelligent manufacturing. Background Technology
[0002] In the "production-monitoring-testing" environment of the entire intelligent manufacturing process, product quality feedback signals are not immediately available after each scheduling decision. Because tasks need to be processed and transferred layer by layer along the workshop, production line, and workstation levels, the final quality inspection results of the product often only become apparent after all processes are completed. Furthermore, the impact of intermediate scheduling decisions on the final quality requires a complete cross-layer transmission cycle to be evaluated. This delay in quality feedback caused by the milestone triggering mechanism means that many intermediate scheduling decisions cannot obtain effective value feedback for a considerable period, constituting the core challenge of intelligent manufacturing scheduling in the time dimension. For example, in a typical multi-layer manufacturing process, the quality inspection of a batch of products can only provide a comprehensive judgment after traversing all processing levels and completing the final test. Meanwhile, the critical impact of resource allocation and task assignment decisions in preceding processes on the quality results can only be indirectly evaluated through backtracking at the milestone trigger moment.
[0003] Traditional quality prediction and scheduling optimization methods typically rely on readily measurable process indicators for decision-making. Statistical process control (SPC) methods detect and alert on quality anomalies by monitoring real-time deviations in key process parameters. However, in scenarios with significant feedback lag, the causal relationship between parameter deviations in intermediate processes and final product quality defects is difficult to accurately identify at the decision-making stage, easily leading to false alarms or missed alarms. Supervised learning-based methods, such as neural network quality prediction models, can learn the mapping relationship between process parameters and quality results from historical data. However, in online scheduling scenarios, feedback delays result in highly sparse and unevenly distributed supervision signals for training samples. This leads to slow model updates and a tendency to overfit to recently labeled samples, making it difficult to effectively determine the long-term quality impact of intermediate decision paths.
[0004] From a reinforcement learning perspective, scheduling methods based on value functions and policy gradients model task allocation and resource scheduling as a sequential decision-making process, learning the optimal policy through continuous interaction with the environment. However, when external quality feedback signals are highly sparse and concentrated on a few milestone nodes, value estimation based on temporal differences struggles to form effective gradients over a large number of unrewarded time steps. Degradation of the advantage function estimation and increased policy gradient variance lead to a significant decrease in training efficiency. To alleviate the reward sparsity problem, exploration methods based on intrinsic motivation compensate for the lack of external rewards by constructing intrinsic driving signals related to state novelty. Representative methods include curiosity mechanisms based on prediction errors, access frequency methods based on pseudo-counting, and information gain maximization methods based on information theory. These methods have achieved some success in general benchmark tasks, but their intrinsic driving signals are mainly based on novelty measures at the instantaneous observation level, lacking explicit encoding of the evolutionary trend of multi-layer structures. In cross-layer coupling scenarios, they are prone to generating noisy driving signals unrelated to milestone progress and cannot reasonably attribute the driving signals to policy updates at each layer. Furthermore, existing methods typically assume that the external reward mapping remains constant throughout the scheduling process. However, in multi-layer manufacturing scenarios with delayed quality feedback, the feedback properties before and after the milestone triggering moment differ significantly, further increasing the difficulty of long-term value identification and cross-layer credit allocation. Summary of the Invention
[0005] This invention addresses the problems of existing technologies by providing a process scheduling optimization method and system for quality feedback delays in intelligent manufacturing. It extracts macroscopic structural features reflecting cross-layer progress from original scheduling observations, obtains low-dimensional state representations with structural semantic consistency through contrastive representation learning, generates intrinsic driving signals for structural novelty in the latent space using online clustering pseudo-counting, and differentiates the global learning driving signals based on the contribution of each layer to structural progress using a cross-layer credit allocation mechanism. This method provides structurally sensitive and controllably decaying auxiliary learning driving support for scheduling strategy optimization under delayed feedback conditions without changing the external definition and evaluation criteria of quality feedback. It systematically solves the long-term value identification difficulties and cross-layer credit allocation problems under delayed quality feedback conditions.
[0006] To address the above technical problems, the present invention provides the following technical solution: a process scheduling optimization method for quality feedback delay in intelligent manufacturing, comprising the aforementioned steps S1 to S5.
[0007] Further, in step S1, the construction of the macroscopic structural feature vector is as follows: Assume there are M layers in the industrial network, the first... The task backlog of the layer at time t is denoted as Its task capacity limit is denoted as The stacking ratio is defined as follows: , Indicates the maximum capacity of the task; the first There are a total of 100 floors The first implementing entity, the first The load of each execution entity is denoted as Extract the average value Maximum value and standard deviation Global structural features include global backlog. and time phase ,in The maximum number of decision steps in a single round; the comprehensive macroscopic structural feature vector is:
[0008]
[0009] The macroscopic structural feature vector is used only as input to the learning-driven branch and does not replace the original observation input of the policy network.
[0010] Furthermore, in the aforementioned step S2, let time... The macroscopic structural characteristics are The structure encoder is Output implicit representation After normalization by the second norm, we obtain Latent space structural similarity is defined as... ; For anchor point samples Its positive samples are defined as ,in and negative sample set The loss is obtained by random sampling from the empirical buffer; the contrastive loss is in the form of normalized log-likelihood.
[0011]
[0012] in The temperature coefficient is used; the target encoder updates via momentum. , Let the momentum coefficient be ; in step S3, let the cluster center set be . , The initialization phase involves collecting initial interaction data, calculating the structural implicit representation set, performing initial clustering to obtain initial cluster centers, and initializing the number of visits to each cluster. During the online update phase, for each time step... The stable hidden representation is calculated by the target encoder. Cluster identifiers are determined according to the nearest center criterion. Update cluster access count And update the cluster center using incremental mean method. The intrinsic learning driving signal is .
[0013] Furthermore, in step S4 above, assuming there are M layers in the industrial network, the intrinsic learning driving signal of each layer is defined as follows: Introducing intrinsic value estimation for each layer. Construct the intrinsic temporal difference residuals:
[0014]
[0015] in As a termination marker, The discount factor is used; the stratum importance score is defined as the time mean of the residual magnitude. Cross-layer weight is defined as follows: , It is a very small positive number; in step S5, the first... The total reward signal for a layer is defined as:
[0016]
[0017] in The weighting coefficients are driven by intrinsic learning; the extrinsic quality feedback signal is triggered by milestones. , Milestone event indicator function The intensity of the reward at the milestone trigger moment. For system status, For scheduling actions, the policies at each layer are updated using a multi-agent proximal policy optimization method.
[0018] (The claims are finalized upon finalization, with the agent providing final additions.)
[0019] Compared with the prior art, the beneficial technical effects of the present invention using the above technical solution are as follows:
[0020] (1) Solving the problem of delayed quality feedback and improving long-term value recognition: In scenarios where quality feedback signals need to undergo multiple layers of processing and cross-layer transmission before they can be concentrated, traditional scheduling methods cannot obtain effective immediate value feedback for a large number of intermediate decisions. The strategies tend to be biased towards short-sighted behavior and ignore decision paths that require long-term cross-layer advancement to realize quality benefits. This invention obtains structurally sensitive and semantically consistent state representations in the representation space through macroscopic structural feature construction and comparative representation learning. In the latent space, it generates intrinsic auxiliary learning signals oriented towards structural novelty through online clustering pseudo-counting. This provides a continuously available and controllably decaying auxiliary learning driver for policy optimization under delayed feedback conditions. This enables the agent to effectively distinguish the long-term potential differences of different intermediate paths in the long prefix interval before the milestone quality feedback is triggered. It avoids misjudging early but insufficient procedural rewards as effective guidance and staying on suboptimal paths, which significantly improves the training efficiency and convergence stability in delayed feedback environments.
[0021] (2) Structure-sensitive adaptive novelty estimation with controllable signal decay that does not dominate the optimization objective: Existing intrinsic motivation methods mainly rely on pixel-level or feature-level metrics at the instantaneous observation level for novelty signals, lacking explicit encoding of the evolution trend of multi-layer structures. In cross-layer coupling scenarios, they are prone to generating noise-driven signals unrelated to milestone advancement and premature saturation. This invention focuses novelty estimation on changes in the morphology of the scheduling structure rather than disturbances from instantaneous observation noise by extracting macroscopic structural features. It replaces the fine-grained counting at the instantaneous state level with access frequency metrics at the structure cluster level, enhancing the semantic consistency and generalization stability of the signal. The intrinsic auxiliary learning signal constructed in the form of the square root reciprocal has monotonically decaying properties, allowing policy updates to gradually transition from structural coverage expansion to milestone-related structural focus as training progresses. Furthermore, the cumulative driving signal over all rounds satisfies the sublinear upper bound constraint, preventing continuous linear accumulation in the long time domain and suppressing the optimization of the external quality feedback objective.
[0022] (3) Differentiated cross-layer credit allocation enhances the targeting of strategy updates at each layer: In multi-layer manufacturing networks, quality milestone gains are jointly facilitated by the joint scheduling behavior of multiple layers and agents. Broadcasting the same global auxiliary learning signal indiscriminately to all layers will blur credit attribution and introduce inter-layer interference. This invention introduces an intrinsic value estimation network for each layer and calculates temporal difference residuals. The relative contribution of each layer to the formation of structural novelty is quantified by the time mean of the residual amplitude. Layers more strongly associated with key structural changes receive more sufficient auxiliary learning drive, while layers less associated with structural state changes are protected from excessive interference from irrelevant signals. Cross-layer credit allocation only affects the backpropagation and weighting of auxiliary learning signals without changing the definition and evaluation criteria of external quality feedback. It is embedded into the existing multi-agent optimization framework in a plug-in manner, significantly improving the attribution clarity and training stability of multi-layer multi-agent joint optimization under the condition of quality feedback delay. Attached Figure Description
[0023] Figure 1 This is a schematic diagram illustrating the overall process flow of a process scheduling optimization method for quality feedback delay in intelligent manufacturing.
[0024] Figure 2 This is a schematic diagram of the main principle of the method of the present invention. Detailed Implementation
[0025] To better understand the technical content of the present invention, specific embodiments are described below in conjunction with the accompanying drawings.
[0026] In this invention, various aspects of the invention are described with reference to the accompanying drawings, in which numerous illustrative embodiments are shown. Embodiments of the invention are not limited to those depicted in the drawings. It should be understood that the invention is implemented through any of the various concepts and embodiments described above, as well as the concepts and embodiments described in detail below, because the concepts and embodiments disclosed herein are not limited to any particular implementation. Furthermore, some aspects of the invention disclosed may be used alone or in any suitable combination with other aspects of the invention disclosed.
[0027] like Figure 1-2 As shown in the embodiment, this paper provides a process scheduling optimization method for quality feedback delay in intelligent manufacturing.
[0028] This method addresses the long-term difficulty in value identification and the short-sightedness of scheduling strategies caused by the severe lag in quality feedback signals throughout the entire "production-monitoring-testing" process of intelligent manufacturing. By elevating the target of the auxiliary learning-driven process from high-dimensional instantaneous observation to a generalizable scheduling structure state, a stable novelty estimate and intrinsic auxiliary learning signal are constructed in the structural representation space using online clustering pseudo-counting.
[0029] (1) Macro-structural feature construction stage: In a typical multi-layer intelligent manufacturing network, tasks need to be processed and transferred layer by layer along the workshop level, production line level, workstation level, etc. The quality feedback signal depends on the triggering of milestone events, resulting in a large number of intermediate scheduling decisions not being able to obtain immediate value feedback. The cross-layer aggregated structural description is extracted from the original scheduling observations, and the macro-structural features are divided into two categories: hierarchical features and global features. Hierarchical structural features include the backlog ratio of each layer (defined as the ratio of the current backlog to the capacity limit), the mean and extreme values and standard deviation of the execution subject load (used to capture potential bottleneck subjects and the degree of load balance). Global structural features include the global backlog obtained by summing the backlogs of all layers, and the time phase (defined as the ratio of the current step sequence to the total length of the round, used to distinguish the resource layout stage in the early stage of the round and the milestone sprint stage in the later stage of the round). The macro-structural feature vector is only used as the input of the learning-driven branch and does not replace the original observation input of the policy network, so that the novelty estimation focuses on the changes in the scheduling structure.
[0030] The structural description of cross-layer aggregation is extracted from the original scheduling observations, and a macro-structural feature vector is constructed. The features are divided into two categories: hierarchical structural features and global structural features. The hierarchical structural features include the backlog ratio of each layer, the average and extreme values of the execution subject load, and the load balance degree to describe the backlog degree and execution load pattern of each layer. The global structural features include the global backlog amount and time phase to describe the overall backlog and process phase under cross-layer coupling. This macro-structural feature vector is only used as the input of the subsequent learning-driven branch and does not replace the original observation input of the policy network. It provides a semantically consistent input basis that is insensitive to local noise for contrastive representation learning.
[0031] (2) Contrastive Representation Learning Stage: Contrastive representation learning encoders are constructed to map macroscopic structural features to a latent space with unified similarity semantics. Within the temporal neighborhood (radius) Samples from the empirical buffer are used as positive samples, and randomly sampled samples from the empirical buffer are used as negative samples. The encoder is trained using normalized log-likelihood loss, making structurally similar states closer in the latent space and separating states with greater structural differences. Latent space similarity is defined as the inner product of the normalized latent representations. The temperature coefficient τ in the contrastive loss is used to adjust the sharpness of the contrastive normalization term. A momentum-updated target encoder method is used to provide a stable and consistent latent representation for subsequent online clustering. The momentum coefficient m controls the update speed, avoiding clustering reference frame drift caused by frequent updates of the online encoder.
[0032] A contrastive representation learning encoder is constructed to map macroscopic structural features to a latent space with unified similarity semantics. Samples within the temporal neighborhood are used as positive samples, and samples randomly sampled from the empirical buffer are used as negative samples. The encoder is trained by normalized log-likelihood loss so that structurally similar states are closer in the latent space, while states with greater structural differences are more separated. The momentum update target encoder is used to provide a stable latent representation for subsequent online clustering.
[0033] (3) Online Clustering and Pseudo-Counting Stage: Online clustering is performed in the latent space. Let the set of cluster centers contain... In the initialization phase, initial cluster centers are obtained by clustering in the latent space using initial interaction data, and the number of visits to each cluster is initialized to zero. In the online update phase, a stable latent representation is calculated by the target encoder for each time step. Cluster identifiers are determined by calculating the Euclidean distance to each cluster center according to the nearest center criterion. Update the access count of this cluster sequentially. and updating cluster centers using incremental mean This allows the cluster center to adaptively migrate according to the distribution of interactive data. It is expressed as the reciprocal of the square root of the number of cluster visits. An intrinsic learning-driven signal is constructed that provides a stronger drive to newly emerging or less frequently accessed scheduling structures, while the signal gradually decays with increasing access frequency. Monotonic decay ensures that the signal naturally weakens when the strategy repeatedly accesses the same cluster of structures, and slow decay provides a significant drive to newly emerging clusters in the early stages of training to increase milestone triggering frequency. The cumulative intrinsic learning drive over all rounds satisfies… The sublinear upper bound constraint.
[0034] Subsequently, online clustering is performed in the latent space to construct pseudo-counting intrinsic learning driving signals. During the initialization phase, initial interaction data is used to cluster in the latent space to obtain initial cluster centers and initialize the access count of each cluster. During the online update phase, cluster centers and cluster access counts are incrementally updated after being assigned to the corresponding structural clusters according to the nearest center criterion. Intrinsic learning driving signals are constructed in the form of the reciprocal of the square root of the cluster access count, so that newly emerging or less frequently accessed scheduling structures receive stronger learning driving signals, while the signals gradually decay as the access count increases. Furthermore, a cross-layer credit allocation mechanism is constructed, introducing an intrinsic value estimation network for each layer. By calculating the temporal difference residual of the global intrinsic learning driving signal and taking its amplitude time mean as the layer importance score, the relative contribution of each layer to the formation of structural novelty is evaluated. After normalizing the contribution weights, the global intrinsic learning driving signal is weighted and allocated to obtain the differentiated intrinsic learning driving signals of each layer. Finally, the allocated intrinsic learning driving signals of each layer are linearly fused with the milestone external quality feedback signal in proportion to form a total reward signal, which is sent to the multi-agent proximal policy optimization update process to complete policy learning and realize end-to-end optimization of the scheduling policy under the condition of quality feedback delay.
[0035] (4) Cross-layer credit allocation and strategy optimization stage: Introduce an intrinsic value estimation network for each layer. By calculating the temporal differential residual of the globally intrinsically learned driving signal. Assess the relative contribution of each layer to the formation of structural novelty. The layer importance score is the time mean of the residual amplitude. A higher score indicates a greater error in characterizing the structural novelty signal for that layer, thus requiring more correction through drive signal feedback. After normalizing the cross-layer weights, the global intrinsic learning drive signal is weighted and allocated to obtain differentiated intrinsic learning drive signals for each layer. The allocated intrinsic learning drive signals for each layer are then proportionally combined with the milestone external quality feedback signal. Linear fusion forms the total reward signal, which is then fed into the multi-agent proximal policy optimization and update process to complete policy learning. Through these steps, end-to-end optimization of the scheduling policy can be achieved under conditions of quality feedback delay.
[0036] In the construction of macroscopic structural eigenvectors, let there be a total of Layered industrial network, first Layer at time The backlog of tasks is denoted as Its task capacity limit is denoted as The stacking ratio is defined as follows: Used to reflect the degree of overpacking in this layer; to characterize the load distribution of multiple entities within the layer, let the first... There are a total of 100 floors The first implementing entity, the first The load of each execution entity is denoted as Extract the average value Reflects the average load level within the floor, as well as the maximum load. Used to identify potential bottleneck entities.
[0037] As a preferred option, the macroscopic structural feature vector also extracts the intra-layer load standard deviation. Reflects the degree of load balancing; global structural characteristics include global backlog. and time phase ,in The maximum number of decision steps per round is used, and the time phase is used to distinguish between the resource allocation phase in the early stage of the round and the milestone sprint phase in the later stage of the round; the comprehensive macroscopic structural feature vector is:
[0038] This feature vector retains the most critical intra-layer morphological differences and cross-layer coupling characteristics in the scheduling structure without significantly increasing dimensionality, and is not used to replace the original observation input of the policy network but only as the input of the learning-driven branch.
[0039] As a preferred embodiment, during the training of the contrastive representation learning encoder, time step 1 is set to 1. The macroscopic structural characteristics are The structure encoder is denoted as Its output is implicit. After normalization by the second norm, we obtain To unify the similarity scale, the latent space structural similarity is defined as the inner product of the normalized latent representations. ; For anchor point samples Its positive samples are defined as ,in and , The time neighborhood radius is the set of negative samples. The loss is obtained by random sampling from the empirical buffer; the contrastive loss is in the form of normalized log-likelihood. ,in The temperature coefficient is used to adjust the sharpness of the contrast normalization term. The smaller the similarity, the more amplified the difference in similarity, and the more the training tends to widen the gap between positive and negative samples; the encoder updates parameters through gradient descent, while the target encoder updates them through momentum. , is the momentum coefficient.
[0040] As a preferred embodiment, in the construction of the intrinsic learning-driven signal for online clustering and pseudo-counting, let the set of cluster centers be... , The preset total number of clusters is used; during the initialization phase, a set of initial interaction data is collected to calculate the structural implicit representation set. This set is then used for initial clustering to obtain the initial cluster centers and to initialize the access count for each cluster. During the online update phase, for each time step... The stable hidden representation is calculated by the target encoder. Cluster identifiers are determined according to the nearest center criterion. Update cluster access counts sequentially. And update the cluster center using incremental mean method. This allows the cluster center to adaptively migrate with the distribution of interactive data to adapt to the changes in the strategy distribution as training iterations occur.
[0041] As a preferred approach, the intrinsic learning driving signal is defined based on the pseudo-counting definition. This form has two key properties: monotonically decaying and slowly decaying. Monotonically decaying ensures that when the policy repeatedly visits the same cluster of structures, the intrinsic learning driving signal naturally weakens, allowing policy updates to gradually shift from expanding structure coverage to focusing on milestone-related structures. Slow decaying provides significant driving force to newly emerging clusters of structures in the early stages of training to increase the frequency of milestone reward triggering, while retaining a moderate driving force in the middle and later stages of training to prevent the policy from becoming rigid too early. The cumulative intrinsic learning driving signal over all rounds satisfies the sublinear upper bound constraint and will not continuously accumulate linearly in the long time domain and dominate the optimization of the external quality feedback objective.
[0042] As a preferred option, construct the intrinsic temporal difference residual. ,in As a termination marker, The discount factor is used; the stratum importance score is defined as the time mean of the residual magnitude. This score describes the first The local information of a layer reflects its fitting bias and degree of unexplainedness to the global intrinsic learning-driven signal. A higher score indicates a greater error in the current characterization of the structural novelty signal by that layer, thus requiring additional correction and learning through backpropagation of the intrinsic learning-driven signal. Cross-layer weights are defined as follows: ,in It is a very small positive number to avoid numerical instability when the importance scores of all layers are close to zero at the same time.
[0043] As a preferred option, in the fusion of the total reward signal, the first... The total reward for policy updates in a layer is defined as a linear fusion of the external quality feedback signal and the internal learning-driven signal after assignment. ,in The intrinsic learning-driven weighting coefficients are used to balance the intensity of the learning drive with external goal constraints; the external quality feedback signal adopts a milestone-triggered form. ,in This is a milestone event indicator function, triggered when inter-layer dependencies are satisfied at a certain stage or when cross-layer collaborative progress reaches a structural condition. The intensity of the reward at the milestone trigger moment; For system status, To schedule actions, each layer of policies is updated using a multi-agent proximal policy optimization method. The total reward is integrated to replace the pure extrinsic reward for advantage estimation and policy update. This introduces a learning-driven signal feedback mechanism in a plug-in manner. Cross-layer credit allocation only affects the feedback and weighting of the intrinsic learning-driven signal without changing the definition and evaluation criteria of the milestone extrinsic reward.
[0044] The controllable decay of the intrinsically learned driving signal is guaranteed by the inherent mathematical characteristics of the pseudo-counting mechanism: for a length of The number of rounds is given. Let the total number of rounds at the end of the round be... The structural cluster, the first The cumulative number of visits to each cluster is And satisfy For any fixed structure cluster, its first The intrinsic learning-driving signal obtained during the second visit is Therefore, the cumulative intrinsic learning drive of this cluster throughout the entire round satisfies Summing over all structural clusters yields the total accumulated intrinsic learning drive that satisfies the given conditions. The upper bound indicates that although the intrinsic learning-driven signal constructed based on pseudo-counting can provide strong incentives for newly emerging structural clusters in the early stages of training, its cumulative size at most grows linearly with the number of interaction steps. It will not continuously accumulate linearly in the long time domain and dominate the optimization of the external quality feedback target, so that the policy update can gradually transition from structural coverage expansion to milestone-related structural focus as training progresses.
[0045] This invention also provides a process scheduling optimization system for intelligent manufacturing oriented towards quality feedback delay, used to implement a process scheduling optimization method for intelligent manufacturing oriented towards quality feedback delay, including:
[0046] The macroscopic feature extraction module is used to extract the cross-layer aggregation structure description from the original scheduling observations and construct the macroscopic structure feature vector; the contrastive representation learning module is used to map the macroscopic structure feature vector to the latent space, using samples in the temporal neighborhood as positive samples and samples randomly sampled from the experience buffer as negative samples, training the encoder through contrastive loss, and updating the target encoder using momentum.
[0047] An intrinsic driver generation module is used to perform online clustering in the latent space and construct an intrinsic learning driver signal by the inverse square root of the number of cluster visits.
[0048] The credit allocation module is used to introduce an intrinsic value estimation network for each layer. It determines the layer importance score by calculating the temporal difference residual of the global intrinsic learning driving signal, and performs weighted allocation on the global intrinsic learning driving signal to obtain differentiated intrinsic learning driving signals for each layer.
[0049] The policy optimization module is used to linearly fuse the allocated intrinsic learning driving signals of each layer with the extrinsic quality feedback signals of milestones to form a total reward signal, which is then fed into the multi-agent reinforcement learning update process to complete policy learning.
[0050] While the present invention has been described above with reference to preferred embodiments, it is not intended to limit the invention. Those skilled in the art can make various modifications and refinements without departing from the spirit and scope of the invention. Therefore, the scope of protection of the present invention shall be determined by the claims.
Claims
1. A process scheduling optimization method for quality feedback delay in intelligent manufacturing, characterized in that, Includes the following steps: Step S1: Extract cross-layer aggregated structural descriptions from the original scheduling observations of the entire "production-monitoring-testing" process, and construct macro-structural feature vectors. The macro-structural feature vectors include hierarchical structural features and global structural features. Step S2: Construct a contrastive representation learning encoder, map macroscopic structural feature vectors to the latent space, use samples in the temporal neighborhood as positive samples and samples randomly sampled from the experience buffer as negative samples, train the encoder through contrastive loss, and update the target encoder using momentum. Step S3: Perform online clustering in the latent space and construct the intrinsic learning driving signal using the reciprocal of the square root of the number of cluster visits; Step S4: Construct a cross-layer credit allocation mechanism, introduce an intrinsic value estimation network for each layer, determine the layer importance score by calculating the temporal difference residual of the global intrinsic learning driving signal, and perform weighted allocation of the global intrinsic learning driving signal to obtain differentiated intrinsic learning driving signals for each layer. Step S5: Linearly fuse the allocated intrinsic learning driving signals of each layer with the milestone extrinsic quality feedback signals to form a total reward signal, and send it into the multi-agent reinforcement learning update process to complete policy learning.
2. The process scheduling optimization method for quality feedback delay in intelligent manufacturing according to claim 1, characterized in that, In step S1: the hierarchical structure features include the backlog ratio of each layer, the average load of the executing entity, the maximum load, and the load balance; the global structure features include the global backlog and the time phase; the macroscopic structure feature vector is only used as the input of the learning-driven branch and does not replace the original observation input of the policy network.
3. The process scheduling optimization method for quality feedback delay in intelligent manufacturing according to claim 1, characterized in that, Suppose there are M layers in the industrial network, the first... The task backlog of the layer at time t is ; The overburden ratio is calculated as follows: , Indicates the maximum capacity of the task; Let the first There are a total of 100 floors The first implementing entity, the first The load of each execution entity is The average load is calculated as follows: ; The maximum load value is: ; The standard deviation of the load is calculated as follows: ; The global backlog is calculated as follows: The time phase is: ,in, This represents the maximum number of decision steps in a single round. Macroscopic structural eigenvectors are , Indicates the time phase.
4. The process scheduling optimization method for quality feedback delay in intelligent manufacturing according to claim 3, characterized in that, Let the macroscopic structural characteristics at time t be: The structure encoder is The output implicit representation is: After normalization to the second norm, we get , Latent space structural similarity is defined as: ; These represent the normalized implicit representations of two different samples (anchor sample and comparison sample), respectively. The contrast loss is expressed as normalized log-likelihood: ; The temperature coefficient is used to adjust the sharpness of the contrast normalization term. The target encoder updates via momentum. , is the momentum coefficient.
5. The process scheduling optimization method for quality feedback delay in intelligent manufacturing according to claim 4, characterized in that, In step S3, the set of cluster centers is: , This is the preset total number of clusters; During the initialization phase, initial interaction data is collected, the structured implicit representation set is calculated, initial clustering is performed to obtain initial cluster centers, and the number of visits to each cluster is initialized. and initialize the cluster identifiers. ; Indicates time step The stable latent representation is calculated by the target encoder; Update cluster access count ,in Indicates the current time step The cluster to which it belongs (i.e., the first) The number of visits to each cluster is calculated, and the cluster center is updated using an incremental average method. , The intrinsic learning driving signal is: .
6. The process scheduling optimization method for quality feedback delay in intelligent manufacturing according to claim 5, characterized in that, In step S4, based on the shared M-layer industrial network, the intrinsic learning driving signal of each layer is defined as: In the formula, Assign weights across layers. , A very small positive number is used to avoid numerical instability when the importance scores of all layers are close to zero simultaneously; Introduce intrinsic value estimation for each layer The intrinsic temporal difference residuals are constructed as follows: , in, As a termination marker, As a discount factor, Layer importance score is defined as the time mean of residual magnitude: .
7. The process scheduling optimization method for quality feedback delay in intelligent manufacturing according to claim 6, characterized in that, Step S5, the total reward signal of the l-th layer is defined as: , in, The intrinsic learning-driven weighting coefficient is used to balance the intensity of learning drive with external goal constraints; This indicates external quality feedback signals for milestones; The milestone external quality feedback signal is triggered in the following manner: ,in For milestone event indicator functions, The intensity of the reward at the milestone trigger moment. For system status, For scheduling actions; The strategies at each layer are updated using a multi-agent proximal strategy optimization method.
8. A process scheduling optimization system for intelligent manufacturing oriented towards quality feedback delay, used in the process scheduling optimization method for intelligent manufacturing oriented towards quality feedback delay as described in any one of claims 1-7, characterized in that, include: The macroscopic feature extraction module is used to extract the structural description of cross-layer aggregation from the original scheduling observations and construct the macroscopic structural feature vector. The contrastive representation learning module is used to map the macroscopic structural feature vector to the latent space, using samples in the temporal neighborhood as positive samples and samples randomly sampled from the experience buffer as negative samples. The encoder is trained by contrastive loss and the target encoder is updated by momentum. An intrinsic driver generation module is used to perform online clustering in the latent space and construct an intrinsic learning driver signal by the inverse square root of the number of cluster visits. The credit allocation module is used to introduce an intrinsic value estimation network for each layer. It determines the layer importance score by calculating the temporal difference residual of the global intrinsic learning driving signal, and performs weighted allocation on the global intrinsic learning driving signal to obtain differentiated intrinsic learning driving signals for each layer. The policy optimization module is used to linearly fuse the allocated intrinsic learning driving signals of each layer with the extrinsic quality feedback signals of milestones to form a total reward signal, which is then fed into the multi-agent reinforcement learning update process to complete policy learning.
9. The process scheduling optimization system for quality feedback delay in intelligent manufacturing according to claim 8, characterized in that, In the intrinsic driver generation module, the cumulative intrinsic learning driver signal over all rounds satisfies a sublinear upper bound: , Where T is the round length and K is the total number of structural clusters appearing in the round. For the number of visits to each cluster, This is a signal that drives intrinsic learning.
10. The process scheduling optimization system for quality feedback delay in intelligent manufacturing according to claim 8, characterized in that, In the credit allocation module, cross-layer credit allocation only applies to the feedback and weighting of the intrinsic learning-driven signal, without changing the definition and evaluation criteria of the milestone's extrinsic quality feedback signal.