Automatic boxing method based on reinforcement learning and dynamic search

By constructing a time-series observation layer and an oscillation cause map, combined with inertial anchoring budget and hysteresis mechanism, the problem of frequent switching of packing paths is solved, thereby improving the stability and efficiency of the packing process and extending the service life of the robotic arm.

CN121072918AActive Publication Date: 2025-12-05华商国际工程有限公司 +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511595970.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-04
Publication Date
2025-12-05
Estimated Expiration
2045-11-04

AI Technical Summary

Technical Problem

In existing automated packing processes, the frequent switching of packing paths under dynamic search conditions leads to low execution efficiency of the robotic arm, accelerated hardware fatigue, and affects system stability and service life.

Method used

By establishing a time-series observation layer, collecting data on container strategy scoring gradients, action withdrawal times, and energy consumption changes, an oscillation cause map is constructed. Inertial anchoring budget and hysteresis mechanisms are introduced to limit high-frequency path switching, generate shadow capture trajectories and withdrawal delay thresholds, activate damped response migration mechanisms, and form closed-loop stable control.

Benefits of technology

It significantly improves the continuity and stability of the packing process, reduces energy consumption and task failure rate, and extends the life of mechanical structures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121072918A_ABST
    Figure CN121072918A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic boxing method based on reinforcement learning and dynamic search, and relates to the technical field of intelligent manufacturing and logistics automation, and the method comprises the following steps: establishing a time sequence observation layer from decision to execution, collecting a strategy score gradient, an action withdrawal frequency and execution energy consumption, and generating baseline data representing boxing sequence oscillation; and performing causal coherent decomposition based on the baseline data, identifying an approximate optimal path switching cluster, calculating switching frequency, staying duration and withdrawing cost, and constructing an oscillation cause map. According to the method, through construction of scoring and execution whole process monitoring, causal decomposition, disturbance playback, energy barrier modeling, a buffer execution mechanism and an inverse phase shock suppression strategy, identification, suppression and extinguishing of boxing path shock are achieved, a strategy stable closed loop is formed, the stability and execution efficiency of the boxing process are effectively improved, and energy consumption and the failure rate are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent manufacturing and logistics automation, and particularly relates to an automated packing method based on reinforcement learning and dynamic search. BACKGROUND

[0002] The "automated packing based on reinforcement learning and dynamic search" refers to introducing the reinforcement learning algorithm and dynamic search mechanism of artificial intelligence in the automated packing process, so that the packing system can autonomously learn and optimize the packing strategy under the changing sizes, shapes and order demands of goods. Reinforcement learning trains the model through the "trial and error-feedback-optimization" way, so that it gradually masters how to maximize the loading rate and improve the stability in the limited box space; dynamic search constantly adjusts, corrects and optimizes the packing path and order according to the real-time arrangement of goods and the changes of the remaining space in the actual packing process. The combination of the two can not only avoid the low efficiency problem caused by the dependence of traditional packing methods on fixed rules, but also realize intelligent adaptation to complex and variable scenarios, thereby significantly improving the space utilization rate, reducing transportation costs and improving the overall efficiency of automated logistics.

[0003] The prior art has the following disadvantages: In the prior art, the automated packing process generally relies on the path optimization mechanism based on search algorithms, but under dynamic search conditions, when the evaluation results of multiple packing paths are close to the optimal, the algorithm often frequently switches between these paths, causing the packing order to produce shock. This shock phenomenon will make the mechanical arm of the execution mechanism in a state of repeated grabbing and withdrawing, causing a large amount of invalid actions to accumulate, not only significantly reducing the packing execution efficiency, but also causing hardware fatigue due to excessive start and stop of mechanical parts, accelerating the wear and failure of key components, thereby having a serious adverse effect on the stability and service life of the system.

[0004] The above information disclosed in the background section is only intended to enhance the understanding of the background of the present disclosure, and therefore it can include information that does not constitute the prior art known to those of ordinary skill in the art. SUMMARY

[0005] The purpose of the present application is to provide an automated packing method based on reinforcement learning and dynamic search to solve the problems in the background.

[0006] In order to achieve the above purpose, the present application provides the following technical solution: an automated packing method based on reinforcement learning and dynamic search, comprising the following steps: S001, establishing a time sequence observation layer from decision to execution, collecting policy score gradient, action withdrawal times and execution energy consumption, and generating baseline data representing packing order shock; S002, based on baseline data, causal coherence decomposition, identify the approximate optimal path switching cluster, calculate the switching frequency, stay duration and withdrawal cost, build the shock cause atlas; S003, through the shock cause atlas, introduce counterfactual replay chain, inject disturbance sequence to key time window, compare path evolution, identify steady state interval, calculate inertia anchoring budget and minimum stay threshold; S004, according to the anchoring budget and stay threshold, build the potential energy landscape in strategy generation, superimpose hysteresis mechanism and time consistency penalty, limit path high frequency switching, form smooth transition area; S005, establish execution buffer cone in smooth transition area, generate shadow grabbing trajectory and withdrawal delay threshold, link path evaluation module output single step execution action list; S006, based on the execution action list, dynamic regulation, inject inverse phase reward pulse in strategy time domain, activate damping response migration mechanism, extinguish container path shock, realize process stable closed loop.

[0007] Preferably, step S001 comprises: On the basis of establishing the time sequence observation layer between the decision process and the execution process, collect the container strategy score gradient, action withdrawal times and execution stage energy consumption change data; Based on the score gradient, record each container strategy score value, score timestamp, target goods number, target grabbing position coordinates and placement position coordinates, and establish a one-to-one mapping relationship between the score sequence and the action instruction; Based on the number of action withdrawals, record the action withdrawal behavior caused by strategy change or execution failure in each container task, collect the withdrawal reason, withdrawal path length, withdrawal duration and withdrawal grabbing and placement position coordinates; Based on the execution energy consumption change data, collect the current, voltage, speed and driving load change information in the execution task, and synchronize the score sequence and the withdrawal behavior according to the timestamp, and build the behavior baseline data structure for identifying container sequence shock.

[0008] Preferably, step S002 comprises: After collecting the score gradient, action withdrawal times and energy consumption change data and generating baseline data for representing container sequence shock, based on the continuity and fluctuation amplitude of the score value, divide the stable score stage, and extract the strategy records with similar score but different container behaviors in each score stage, build the strategy path switching cluster; Time series behavior tracking is performed on each strategy path switching cluster, the duration, switching frequency and withdrawal behavior corresponding to each switching event are extracted, and the average strategy stay duration and withdrawal cost of each strategy path switching cluster are calculated; Collecting energy consumption data corresponding to each strategy path switching cluster, analyzing the current, power, temperature rise and terminal vibration change amplitude before and after path switching, and identifying high energy consumption response type path switching cluster; Fusing the score change behavior, strategy path switching feature and energy consumption response data to construct a shock cause graph for reflecting the causes of packing order shock, and the graph covers the complete time sequence and path number evolution process.

[0009] Preferably, when constructing the shock cause graph, for each strategy path switching event, the switching start and end time, corresponding path number, switching frequency and energy consumption change data within five seconds after switching are labeled, and different colors are used to identify power abnormalities, terminal vibration enhancement and action withdrawal situations, so as to realize visual marking of high shock risk areas.

[0010] Preferably, step S003 comprises: After constructing the shock cause graph for reflecting the causes of packing order shock, selecting a key time window based on the score change gradient, path jump frequency and energy consumption fluctuation amplitude; Inserting score disturbance, packing order disturbance and target cargo disturbance in the key time window respectively, tracking the strategy path change, action withdrawal and energy consumption response after disturbance, and extracting the strategy behavior interval that remains stable or recovers to stability; Comparing the strategy evolution sequences before and after disturbance to construct the mapping relationship between disturbance and strategy stability, and identifying the boundary conditions affecting the strategy stability; Based on the stable behavior interval, calculating the strategy inertia anchoring budget and the minimum strategy stay threshold, which are used as control parameter inputs for subsequent construction of path switching inhibition mechanism.

[0011] Preferably, step S004 comprises: After calculating the inertia anchoring budget and the minimum strategy stay threshold of the strategy path, constructing a path potential energy graph, and determining the path score potential barrier strength based on the offset amount of the current score and the anchoring budget boundary; Introducing a hysteresis mechanism in the path score judgment process, setting the score jump radius and hysteresis period through the analysis of the score trend cumulative amplitude and the continuity of the sliding score window, and inhibiting the path switching caused by short-term score fluctuation; Based on the difference between the actual execution length of the strategy and the minimum strategy stay threshold, introducing a time consistency penalty term to impose a penalty value on the replacement path whose score does not reach the execution length, and avoiding early strategy jump; The path potential energy graph, hysteresis mechanism and time consistency penalty term work together to construct a strategy score smooth transition area, limit the strategy path switching frequency, and improve the stability of the packing path.

[0012] Preferably, step S005 comprises: When the policy score is in the smooth transition region, a buffer cone structure is constructed, with the current grasp point as the starting point and the target placement point as the direction, and the cone angle and length are set according to the score gradient and historical jump conditions; A shadow grasp trajectory is generated in the execution buffer cone, the trajectory meets the requirements of movement continuity, spatial obstacle avoidance and attitude consistency, and whether it is converted into an actual action is determined according to the score trend and hysteresis; A motion withdrawal delay threshold is set, the withdrawal waiting time is determined according to the score fluctuation period, the score difference and the executed time, and the immediate jump behavior is suppressed when the withdrawal condition is not met; After completing the path verification and delay evaluation, the score, path stability and execution state are re-integrated, and a structured single-step execution action list is output to guide the physical action execution of the current strategy.

[0013] Preferably, the shadow grasp trajectory is only allowed to be converted into an actual execution action when the continuous score update meets the hysteresis trigger condition and the path stability index reaches the set threshold, and the conversion into an actual execution action can only be started after the withdrawal delay timer ends, to ensure that the action stability verification before policy switching is sufficient.

[0014] Preferably, step S006 comprises: Based on the execution action list output by the execution buffer cone, the score fluctuation frequency and direction reversal characteristics are extracted within the score update period, and a high-frequency oscillation interval is identified; An inverse-phase reward pulse opposite to the score direction is injected into the score change stream in the identified oscillation interval, a score suppression window is constructed, and the policy switching behavior is delayed; When the score trend approaches the switching state, a damping response mechanism in the action execution phase is activated, the initial speed, acceleration and attitude change of the path are controlled, and the action is released slowly; The execution action list is dynamically updated based on the score trend, the policy execution state and the damping feedback result, a closed-loop control structure driven by the policy score, the rhythm control and the action feedback is constructed, and the adaptive extinguishing and stable execution of the score disturbance are realized.

[0015] Preferably, the injection period and amplitude of the inverse-phase reward pulse are dynamically adjusted according to the score change rate and the candidate path score difference, and a preset injection time point is set before the score trend approaches the anchor budget boundary, to form a flat top region of the score rise and delay the trigger of the policy jump behavior.

[0016] In the above technical solution, the technical effects and advantages provided by the present application are: The application realizes quantitative analysis and cause positioning of unstable strategy behavior by constructing a time sequence observation system for the whole process from decision to execution, accurately collecting score gradient, action withdrawal behavior and energy consumption change; combined with causal coherence decomposition and counterfactual disturbance playback chain, it first realizes the graph modeling and visual explanation of the causes of containerization strategy shock; further, by introducing inertia anchoring budget and minimum strategy residence threshold, a barrier energy landscape with behavior stickiness is constructed, supplemented by hysteresis mechanism and time consistency penalty term, which fundamentally suppresses the high-frequency strategy jump tendency; at the execution layer, an execution buffer cone is constructed and shadow grabbing trajectory and withdrawal delay threshold are deployed, so that the physical action has the absorption capacity and slow-release characteristics of score disturbance; finally, by injecting inverse phase reward pulses in the time and frequency domains of strategy response, the damping migration mechanism of execution behavior is effectively activated, the active suppression and rapid extinguishing of strategy shock are realized, thereby forming a closed-loop stable intelligent containerization control process. The overall scheme realizes multi-layer collaborative shock suppression from strategy learning, path evaluation to action execution, effectively improves the continuity, stability and mechanical structure life of the containerization process, and significantly reduces the execution energy consumption and task failure rate. BRIEF DESCRIPTION OF DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed in the embodiments. Obviously, the drawings described below are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art based on these drawings.

[0018] Figure 1 The method flowchart of the automatic containerization method based on reinforcement learning and dynamic search according to the present application. DETAILED DESCRIPTION

[0019] Example implementations will now be described more fully with reference to the accompanying drawings. Example implementations may, however, be implemented in many different forms and should not be construed as limited to the examples set forth herein; rather, these example implementations are provided so that this disclosure will be thorough and complete, and will fully convey the inventive aspects to those skilled in the art.

[0020] The present application provides an automatic containerization method based on reinforcement learning and dynamic search as shown in Figure 1 The method flowchart of the automatic containerization method based on reinforcement learning and dynamic search according to the present application. S001, a time sequence observation layer between the decision process and the execution process is established, the score gradient of the containerization strategy, the number of action withdrawals and the energy consumption change data of the execution stage are collected, and the baseline data for representing the containerization sequence shock are formed; To solve the problem of execution shock caused by frequent switching of the loading path in the approximate optimal region, a detailed timing observation process needs to be established from the whole process of strategy decision to physical execution. Through the multi-dimensional collection of three types of data: score change, action withdrawal, and energy consumption fluctuation, the baseline description structure of the loading shock is constructed, and the specific implementation steps are as follows: A full-process time observation path from strategy generation to action execution is constructed, which specifically includes recording the score value corresponding to the current strategy in real time when the reinforcement learning strategy makes each loading decision. The score value is the comprehensive evaluation index in the current loading state, and the numerical range is set to between 0 and 1. The closer to 1, the better the current loading result. The generation process of the score value is based on the weighted calculation of four specific quantitative dimensions: loading space utilization, target cargo posture stability, balance of residual space distribution after loading, and path execution length. At the same time, a millisecond-level timestamp is attached to each score generation to track the accurate time node of score generation. After each score generation, the target cargo number, three-dimensional coordinates of the target grabbing position (X, Y, Z axis positions represented in millimeters), and specific spatial positioning of the target placement point in the box (defined by three-dimensional coordinates of the length, width, and height of the box) are recorded synchronously. In addition, the strategy score is bound with the corresponding action number to establish a unique mapping table to ensure that there is a one-to-one correspondence between the score value and the subsequent action, thereby forming a continuous score sequence data chain. This sequence will be fused with other execution observation data in the subsequent steps to depict the internal relationship between strategy fluctuations and behavior fluctuations.

[0021] The detailed trajectory of each action withdrawal in the execution process is recorded in real time, which specifically includes the complete process of action termination due to reasons such as strategy re-evaluation, target posture recognition failure, cargo size exceeding the limit, or path blocked by obstacles after the loading robot receives the loading instruction and starts execution. Each action withdrawal event needs to record the following five parameters: (1) the trigger source of the withdrawal, such as whether it is due to score update, whether the target is detected to be mismatched, or whether the path verification fails; (2) the loading instruction number corresponding to the withdrawal action; (3) the three-dimensional coordinates of the grabbing point and the placement point of the current execution action; (4) the moving path trajectory length of the robot before withdrawal (measured in millimeters), which is obtained by integrating the position sensor; (5) the total time used for the withdrawal action (measured in milliseconds), which is used for subsequent evaluation of action redundancy and execution jitter frequency. The complete record of each withdrawal action is associated with the strategy score state in the current time period and marked on the score sequence. By statistically analyzing the intensity of withdrawal actions in certain time periods with large score fluctuations, the unstable regions of the strategy can be identified, and the initial symptoms of loading shock can be further deduced.

[0022] Continuous energy consumption observation is performed throughout the entire process of action execution, specifically including current input value, voltage stability, driving motor operating temperature, mechanical arm load change and end execution speed curve during the execution of the boxing instruction to the completion of the placement of the object. The energy consumption measurement adopts a high-frequency sampling method, collecting current and voltage parameters every 50 milliseconds, recording the load state of the driving unit and the linear speed change of the end effector every 200 milliseconds. For example, if a current execution is a grabbing task, the current is 1.8A and the speed is 120mm / s during the movement of the mechanical arm from the initial position to the target cargo position, and the current rises to 2.6A and the speed fluctuates to ±45mm / s during the withdrawal phase, it can be preliminarily judged that the task has energy consumption abnormalities caused by strategy shock. All energy consumption data are recorded with time stamps and classified and summarized according to the grabbing action number, and are connected with the corresponding score and withdrawal data. By constructing the "score change - withdrawal behavior - energy consumption curve" three data linkage relationship, it can be identified whether there is a significant energy consumption surge in the strategy frequent update interval, so as to convert the abstract strategy shock into a quantifiable physical cost measurement, and enhance the behavior transparency of the boxing system.

[0023] The score sequence, withdrawal trajectory data and energy consumption curve are data fused with the boxing task number as the index to construct a boxing behavior baseline data structure for shock identification. In specific implementation, based on the time synchronization mechanism, the three types of data are uniformly mapped to the same time axis, and a sliding time window is set to align and analyze different task periods. The width of each window is set to 5 seconds, and the score change frequency, the total number of withdrawal times, the average energy consumption and the standard deviation contained in the window are used as baseline feature vectors to construct the shock intensity index of each time period by feature superposition. For example, when the score change is more than 5 times, the withdrawal action is more than 3 times, and the energy consumption variance is higher than the preset threshold 0.08, the window is determined as a high-risk interval of shock. The baseline data is further assigned with the boxing sequence number and time stamp, which is used to track the stability of the boxing sequence at different stages, and provides complete, detailed and structured data support for subsequent shock cause identification, strategy modification guidance and execution path optimization.

[0024] S002, based on the collected baseline data, performing causal coherence decomposition to identify the approximate optimal boxing path switching cluster existing in the boxing sequence, calculating the switching frequency, strategy residence time and withdrawal cost of each switching cluster, and constructing a shock cause map reflecting the reasons for the boxing sequence shock; To further identify the specific causes of the shock behavior caused by frequent switching of strategies in the boxing process, it is necessary to construct a path switching cluster based on the baseline data, combined with the score sequence, action withdrawal trajectory and energy consumption change, extract shock indicators, and visualize the causal relationship, which includes the following steps: For the established score gradient time series, the whole score curve is divided into multiple non-overlapping score stages according to the continuity and fluctuation trend of the score value change, and each stage has a time span of not less than five seconds, and the score change is required to be not more than 0.02, so as to ensure that the score in the divided interval is relatively stable. For each score stage, extract all the packing decision records whose score values fall within the 0.05 interval but whose execution behaviors have significant differences, and the specific behavior differences are determined by the target cargo number, target grabbing point spatial coordinates, and target placement position three-dimensional label. When the score difference of any two strategies is less than 0.02 but the corresponding packing object numbers are inconsistent, or the grabbing point and the placement point positions differ by more than 20 mm, it is considered that the behavior difference is clear. On this basis, these strategy records are combined into strategy switching clusters, each switching cluster contains not less than three strategy units with similar scores but different behaviors, and the switching cluster number is sequentially incremented in time order for subsequent behavior tracking. The construction process breaks through the path selection method of existing technology relying only on score value sorting, introduces a new perspective of score and behavior matching, and accurately captures potential shock sources with similar scores but serious behavior fluctuations.

[0025] For each constructed strategy switching cluster, time series behavior tracking is performed to analyze the switching characteristics of the specific performance in the whole packing process. In the tracking process, first, the time point at which each strategy unit first appears and its corresponding strategy number are extracted, and then the duration of the strategy after being executed is recorded until it is switched to another strategy in the switching cluster. Each time the behavior jumps from the current strategy to another strategy is counted as a switch, and the switching time, starting strategy number, target strategy number, and jump direction are recorded. For example, from strategy A to strategy B, the duration of strategy A is 1.2 seconds, and the jump time point is the 18th second, then the corresponding record is "strategy A→strategy B, duration 1.2 seconds, switch at T=18s". At the same time, the cumulative number of times each strategy is activated in the switching cluster and the average duration are counted, and whether an action rollback behavior follows each switch is extracted. If an action rollback event occurs within five seconds after the strategy switching, it is considered that the strategy switching leads to unstable execution, and the total number of rollbacks is accumulated and marked in the switching record. Finally, the behavior indicator set of each switching cluster is output, including the switching frequency (the number of switches per unit time), the average strategy residence time (the average duration of each time), and the rollback cost after strategy switching (the average number of rollbacks per unit switching), as the core input of shock assessment. This method is different from the traditional method of analyzing the success rate of the strategy or the fluctuation of the strategy score in the existing packing system, and for the first time introduces the linkage analysis of behavior chain and physical feedback, fully reveals the dynamic coupling relationship between the strategy layer and the execution layer.

[0026] For each policy switching cluster corresponding to the packing behavior segment, all associated energy consumption data are extracted, and the direct impact of switching behavior on energy consumption is analyzed. In specific implementation, the execution current, voltage, drive shaft power, servo motor load torque and end speed change data in the complete time interval from policy activation to the replacement of the policy are extracted first, the sampling frequency is set to once every 50 milliseconds, and the duration corresponds to the complete length of the policy residence segment. With policy switching as the boundary, the average current value, power peak value, temperature rise amplitude and execution end vibration amplitude before and after switching are counted respectively, and the vibration amplitude is calculated by the three-axis speed standard deviation. If there is a significant difference before and after switching, such as power rise exceeding 10%, end vibration amplitude doubling, current fluctuation being severe and accompanied by abnormal temperature growth, it is marked as "high energy consumption response switching". After comparing the energy consumption performance of multiple switching clusters, it can be found that some clusters have a continuous energy consumption rising trend when the switching behavior is frequent, and others are relatively stable. Combined with the behavior characteristics extracted in the second step, high oscillation risk switching clusters and general path fluctuations can be effectively distinguished. This method not only establishes a clear relationship between policy switching and physical resource consumption, but also realizes the causal closed loop of inferring the reliability of the strategy layer from the action layer feedback, breaking the problem of lack of physical evaluation constraints in traditional score-driven decision making.

[0027] Based on the above score fluctuation behavior, policy switching statistical characteristics and energy consumption response indicators, an oscillation cause map is constructed to visually show the evolution process of policy oscillation behavior in the time and path dimensions. The horizontal axis of the map is the standard time sequence, with seconds as the unit, and the vertical axis is the strategy path number. Each policy switching cluster is identified by an independent color band, and the start and end time intervals of policy switching and the switching frequency are marked. Each path jump connection is represented by a dashed line. At each jump point, the energy consumption response within five seconds after the jump is superimposed and displayed, including current rise amplitude, end vibration change value and whether the action is withdrawn, and high-risk indicator areas are marked with different colors. For example, a red dot represents that the jump is accompanied by a high power peak and action withdrawal, a yellow mark indicates only power anomaly but no withdrawal, and green indicates stable switching. The overall structure of the map is a multi-dimensional superposition of the policy evolution path diagram, the behavior change diagram and the energy consumption reaction diagram, realizing visual, traceable and quantifiable expression of the whole process of oscillation behavior.

[0028] S003, based on the oscillation cause map, a counterfactual replay chain is introduced to insert a disturbance sequence into the selected key time window, compare the packing path evolution process under each disturbance condition, identify the stable evolution interval, and calculate the inertia anchoring budget and the minimum policy residence threshold; To effectively solve the problem of high-frequency strategy switching caused by close path scores, based on the unstable segments identified in the oscillation cause map, artificial disturbance sequences are introduced in the key time window, the strategy response evolution process is observed, the system stability boundary is identified, and the anchoring budget and stay duration required for strategy inertia execution are calculated, including the following steps: Based on the marked score sharp change segments and path frequent jump zones in the oscillation cause map, representative key time windows are selected for counterfactual testing. Each time window is extended by five seconds before and after, totaling ten seconds, ensuring coverage of a complete strategy switching cycle. In the selected time window, the following three conditions must be met: the score change gradient exceeds 0.03, the strategy path jumps at least twice within ten seconds, and there is at least one container action withdrawal between path jumps, accompanied by significant energy consumption increase (such as current fluctuation exceeding average value by more than 20%). In this window, the container behavior has shown obvious instability characteristics, making it suitable as a test basis for subsequent counterfactual disturbance. Unlike traditional methods that only evaluate strategy output results, this step synchronously binds strategy evaluation and execution, ensuring that counterfactual insertion has actual observable behavior response.

[0029] In the key time window, multiple disturbance sequences are constructed according to the preset scheme and inserted into the score gradient change nodes. The disturbance sequence consists of three types of disturbances: score disturbance, container sequence disturbance, and target cargo characteristic disturbance. Score disturbance adjusts the score value manually, such as adjusting the original score 0.83 to 0.87 or 0.79, simulating strategy selection changes after score advantage enhancement or weakening; container sequence disturbance changes the order of two container actions in the same window, such as reversing the container sequence of the 12th and 14th seconds, verifying whether the strategy jump is affected by the order; target cargo disturbance replaces the container object pointed to by the current strategy, such as changing the target cargo volume from small to medium size, or changing its edge shape, observing whether the strategy produces jump tendency. Each type of disturbance is inserted separately, with consistent control variables, and each type of disturbance is implemented more than three times. After each disturbance insertion, the strategy output path, action execution feedback, energy consumption change trend, and withdrawal behavior are tracked to evaluate the strategy's resistance to disturbance and stability critical point under disturbance conditions.

[0030] For each type of disturbance, the strategy evolution sequence is aligned with the original undisturbed sequence one by one. The comparison is made to determine whether the path switching is intensified, the strategy is returned to the original path, the withdrawal behavior is increased or decreased, and the energy consumption is fluctuated or continues to expand. The specific determination criteria are as follows: if the strategy returns to the pre-disturbance path within three seconds after the disturbance is inserted and does not trigger additional withdrawal behavior, it is considered to have high stability; if the path is replaced after the disturbance, accompanied by a new round of action withdrawal and abnormal increase of energy consumption, it indicates that the strategy has poor stability under this type of disturbance. To quantify the steady-state trend, the path duration, switching frequency, withdrawal frequency, and energy consumption average and standard deviation after disturbance are calculated to construct a disturbance-evolution behavior mapping table, which marks which disturbance combination is easy to induce unstable response and which path has convergence under different disturbance conditions. Through a large number of disturbance comparison and analysis, a steady-state evolution interval is extracted, i.e. the strategy behavior interval that remains stable or quickly returns to a stable state under the influence of disturbance. This analysis process is different from the conventional strategy performance evaluation, which not only measures the strategy output, but also focuses on the recovery ability of the strategy in abnormal disturbance scenarios, providing behavior boundary basis for subsequent shock suppression parameter design.

[0031] Based on the statistical results of all steady-state evolution intervals, the strategy inertia anchoring budget and the minimum strategy stay threshold are calculated respectively. The inertia anchoring budget is defined as the score floating range, path disturbance tolerance, and target change limit that the current strategy can maintain stable execution under disturbance conditions. For example, the score floating budget is ±0.02, the path jump tolerance is within once, and the target cargo replacement allows a size difference of no more than 10%. The minimum strategy stay threshold is the time length that the strategy needs to execute without switching intervention to ensure that the strategy has enough time to influence the execution layer before the disturbance is introduced, and is not misled by the instantaneous score fluctuation to produce unnecessary switching. This threshold is usually between 2.0 and 3.0 seconds and is dynamically adjusted according to the strategy historical withdrawal rate and path execution energy fluctuation curve. Finally, the inertia anchoring budget and the minimum stay threshold of each strategy path are packaged as a structured parameter set, which is used as the input basis for building the potential energy landscape and hysteresis control mechanism in the next stage, so that the strategy will not switch easily when approaching the score fluctuation threshold, thereby improving the overall path stability and suppressing the strategy shock behavior caused by perturbation.

[0032] S004, according to the inertia anchoring budget and the minimum strategy stay threshold, a potential energy landscape is constructed in the bin packing strategy generation link, a hysteresis mechanism and a time consistency penalty term are added, the switching frequency of the strategy path is controlled and limited, and a smooth transition area of the strategy is formed; To avoid the strategy frequently jumping near the score change threshold, causing the packing process to produce instruction shock and action retreat, the present application introduces an energy barrier structure based on inertia anchoring budget and minimum stay threshold in the strategy generation process. Through the interaction of barrier map, hysteresis mechanism and time consistency penalty, a smooth transition area that can buffer score fluctuations is constructed, which specifically includes the following steps: Combined with the extracted inertia anchoring budget and score change data of each packing path, a complete path barrier energy map is constructed. The map takes the offset between path score and score stable interval as the vertical axis, and the path index number as the horizontal axis, and numerically expresses the current score state of each path and its inertia anchoring threshold interval. In actual construction, the absolute difference between the current path score and the anchoring budget boundary needs to be calculated. For example, if the score of path P12 is 0.846 and the anchoring budget is ±0.015, then when the score changes to 0.832 or 0.861, it is within the allowed range and is considered a low barrier area. When the score changes beyond this interval, it enters the high barrier area. In the map, the barrier strength of different paths is identified using a color gradient method. The color gradient from light to dark indicates that the closer to the budget boundary, the more difficult it is to jump. In the strategy score update process, the map values are recalculated each time to ensure that the score selection has the latest energy gradient reference. Unlike the maximum value strategy decision mechanism of traditional score sorting, before selecting the path with the highest score, it is first determined whether the current score has crossed the barrier threshold. If not, the current strategy is maintained, even if the score is not optimal, it will not jump immediately, thereby forming a strategy execution inertia constrained by energy potential.

[0033] On the basis of the constructed barrier map, a path hysteresis mechanism is added to the strategy score judgment process, so that path switching is no longer directly triggered by a single score fluctuation, but is determined by the persistence and accumulation of score trends. The specific operation is as follows: set the jump radius and hysteresis judgment period for the current execution path, for example, the jump radius of the current execution path P8 is set to 0.018, and the hysteresis period is set to 1.2 seconds. During this period, if the cumulative change amplitude of the score in the downward direction does not exceed 0.018, even if the score is once lower than that of other paths, it will not be switched. Only when the score continues to decline and the cumulative change exceeds the jump radius, is the jump to the path with a better score allowed. The hysteresis mechanism uses a sliding score window detection method, which evaluates the score change direction every 0.2 seconds based on time. If the score change direction is downward for six consecutive times and finally breaks through the set jump radius, it is judged that the hysteresis switching condition is met. This mechanism introduces score trend analysis and behavior inertia parameters to ensure that the path jump has a "impulse trigger" feature, overcoming the existing score decision-making strategy structure, and has a significant shock suppression effect.

[0034] On the basis of the hysteresis mechanism, a time consistency penalty term is introduced, and the relationship between the actual execution time and the minimum stay threshold is used as the basis to impose a policy replacement penalty on the situation where the minimum execution time is not met and the new scored path is covered. For example, if path P3 has just been selected for execution for 1.1 seconds, and its minimum policy stay threshold is 2.7 seconds, the current scoring update selects P4 as the preferred path, then the score of P4 needs to be deducted by a penalty value proportional to the time difference, such as 0.012 points, so that it does not have enough advantage in the ranking. Only when the current path has been running for 2.7 seconds or more, the scoring update is considered valid. The penalty mechanism calculates according to the execution time difference and the path fluctuation frequency, which not only ensures that each path obtains a complete execution period, but also avoids the disconnection between the scoring mechanism and the execution response, further reducing the switching incentive caused by short-term scoring high fluctuation. This mechanism first establishes a dynamic link between scoring weight and execution time in policy strategy control, which is different from the traditional static policy output method, and enhances the logical consistency between scoring and behavior.

[0035] The barrier map, hysteresis mechanism and time penalty logic are combined to form a smooth transition area of the strategy, and switching rhythm control is applied in the area. The smooth transition area is centered on the current path score, and the anchor budget allowed area, the hysteresis trigger waiting area and the penalty release buffer area are extended upwards and downwards to form a dynamic scoring selection band. Within this scoring band, the strategy score is allowed to fluctuate within a certain range, but before the jump threshold is not met, the time penalty is released or the hysteresis trend is confirmed, the path does not jump, but the current strategy is output stably. The strategy scoring update period is uniformly set to every 0.2 seconds, and each update rejudges whether it is still within the smooth transition area according to the latest scoring state, execution time and historical scoring fluctuation trend. If it is still within the range, the path remains unchanged; if all three switching conditions are met, the path smooth transition is completed, and the new strategy is executed. This smooth transition area mechanism first realizes the joint restriction of the container strategy from the three dimensions of scoring fluctuation, execution persistence and behavior trend, so that the strategy path does not jump due to slight scoring fluctuation, significantly improves the physical stability and behavior consistency of strategy selection, and lays the foundation for subsequent execution action buffer control and reverse incentive mechanism.

[0036] S005, on the basis of the strategy smooth transition area, an execution buffer cone is established, a shadow grabbing trajectory and an action withdrawal delay threshold are generated combining the barrier energy landscape and the hysteresis mechanism, the path evaluation module recalculates the current container path score, and outputs the corresponding single-step execution action list; To ensure the execution delay and physical buffer capacity of the packing strategy path in the process of score fluctuation and score switching, an execution buffer cone is established on the basis of the strategy smooth transition area, combined with the previously constructed potential energy landscape and hysteresis mechanism, a shadow grasping trajectory with verification and hysteresis response characteristics is generated, and a specific action withdrawal delay threshold is set, and finally the complete single-step execution action list is output through the linkage of the path score reevaluation process, which specifically includes the following steps: According to the score state of the current packing strategy path and the historical path jump behavior, an execution buffer cone structure with directionality, time tolerance and path score tolerance is constructed in three-dimensional space. The starting point of the execution buffer cone is the grasping point coordinate defined in the current strategy path, and the direction extends towards the current target placement point direction, forming a spatial cone. The axis of the cone is the execution direction from grasping to placing, and the cone angle is determined by two parameters, one is the instantaneous gradient value of the score change, and the other is the stability score of the last stage path in the same score interval. For example, in the case of slow score change and low frequency of historical path jump, the cone angle is set to within 15 degrees; if the score fluctuates rapidly and the path frequently switches in the near future, the cone angle is expanded to 25 to 30 degrees to expand the buffer zone. The length of the cone is set to the maximum movement distance corresponding to the expected action duration of the current execution path, which is usually between 0.8 meters and 1.2 meters. The entire buffer cone is the behavior tolerance space under the influence of path score disturbance, which can absorb small amplitude path jump requests within the score range, making the strategy behavior continuous. This structure is significantly different from the rigid instruction logic in the traditional execution path, and it provides a buffer zone for the conversion between strategy and physical action, increasing the elasticity of the packing decision to absorb score fluctuations.

[0037] Inside the buffer cone, a shadow trajectory is generated according to the current policy score trend and hysteresis interval judgment state, which is used for pre-execution simulation before the actual action is issued. The shadow trajectory is completely inside the buffer cone in space, its starting point coincides with the current grasp point, and its ending point is the target placement position of the candidate path, and the trajectory path is adjusted according to the score trend direction. The following three dimensions are considered in the trajectory planning process: first, the continuity of the moving speed, to ensure that the trajectory does not appear to decelerate and stop in the middle of the way; second, the space obstacle avoidance judgment, through the collision detection between the pre-play path and the boundary of the goods or container, to ensure the accessibility of the path; finally, the consistency of the posture, that is, to ensure that the change between the grasp angle and the placement angle does not exceed the set maximum posture difference threshold, for example, the change of the gripper rotation angle does not exceed 30 degrees. The generated shadow trajectory is not immediately executed, but is waiting according to whether the score continues to enter the policy jump interval and whether the hysteresis judgment meets the switching condition. If the above conditions are continuously met, the shadow trajectory is converted into a real action trajectory; if the score trend falls back or the path judgment fails, the trajectory is immediately invalidated, and the current policy continues to be maintained. Through this "pre-judgment-verification-transformation" path buffer mechanism, the risk of execution interruption due to incomplete path preparation is effectively avoided, and the physical implementability of policy jumping is improved.

[0038] After the shadow trajectory is constructed and enters the score continuous observation state, a corresponding action withdrawal delay threshold is set to add a delay protection time for the current policy action when the score trend has not been confirmed and the path verification has not been completed. The length of the withdrawal delay threshold is determined by the score fluctuation period, the policy score difference and the current policy execution time. For example, if the score trend duration is less than 1 second, the policy score change amplitude is less than 0.02, and the current path has been executed for more than 1.5 seconds but has not reached the minimum policy stay threshold of 2.5 seconds, the withdrawal delay time is set to 600 milliseconds, that is, even if the candidate path score rises, the withdrawal action is not initiated within 600 milliseconds. The delay timer is refreshed periodically, and whether the withdrawal condition is met is judged every 200 milliseconds. If the score falls back before the delay time ends, or the candidate path fails in the shadow trajectory verification process, the withdrawal action is automatically cancelled, and the current path continues to execute. The introduction of this delay mechanism is a kind of "physical patience" design for the packing action layer, which changes the policy judgment from the jump mode of "select and withdraw" to the stable process of "evaluation-confirmation-execution", enhances the resistance of the execution action to the score disturbance, significantly reduces the unnecessary withdrawal operation, and reduces the energy consumption and wear and tear problems caused by frequent start and stop of the robot arm.

[0039] After the completion of the buffer cone execution path filtering, shadow trajectory completion stability verification, and withdrawal delay threshold time evaluation, the current strategy path score is re-evaluated, the latest score, path stability parameters, and action execution conditions are summarized, and a complete single-step execution action list is output. The list includes the following: (1) strategy path number and current score value; (2) grasping object identification, including cargo number, size, and shape classification; (3) actual grasping point three-dimensional coordinates, including X, Y, Z values and clamping angle; (4) target placement point position coordinates and placement direction; (5) path score position in the potential energy diagram and whether it enters the high potential barrier area; (6) whether the current path is the original path or a new candidate path, whether it has passed the shadow trajectory verification; (7) withdrawal delay timer state, remaining time; (8) strategy execution time, remaining time, and whether the minimum stay threshold has been exceeded; (9) the expected energy consumption and path movement distance of this action. Through this list, the execution mechanism can receive structured, clear, and fully evaluated action instructions at once, avoiding execution failures or instruction conflicts caused by missing score information, unprocessed strategy jumps, and insufficient path verification.

[0040] S006, based on the execution action list output by the execution buffer cone, real-time dynamic control is performed, inverse phase reward pulses are injected in the time and frequency dimensions of the packing strategy, and a damping response migration mechanism is activated to extinguish the oscillation behavior during the packing path switching process and maintain the stability of the packing process closed loop; To realize the dynamic stability response of the packing execution action list based on the execution buffer cone output under high-frequency disturbance conditions, inverse phase reward pulses are actively injected in the time and frequency dimensions of the strategy score response, and combined with the damping migration response mechanism of the action layer, the rhythm traction and behavior release of strategy jump are realized, and finally an adaptive oscillation extinction closed loop control structure of the packing execution process is constructed, including the following steps: According to the output execution action list, the strategy score fluctuation characteristics are monitored in real time within the scoring response update cycle, the time rhythm parameters and frequency characteristic indexes of the score change are extracted, and the critical conditions of high-frequency oscillation occurrence are identified. The score monitoring cycle is set to 100 milliseconds, the score update times in the past 500 milliseconds time window, the maximum score floating amplitude and the frequency of score direction reversal are counted. For example, if the score is updated more than 6 times in 500 milliseconds, the maximum fluctuation exceeds 0.015, and the number of direction reversals exceeds 3, it is marked as a "score high-frequency fluctuation interval". At the same time, the duration of the current path in the last three strategy switches is counted, if the three durations are all less than 2 seconds, and are accompanied by score jump and action withdrawal, it can be determined that the strategy execution has entered the unstable section. Under this condition, instead of directly selecting the path based on the highest score strategy, the inverse phase damping process is prepared to provide the basis for the rhythm perception of subsequent strategy rhythm adjustment. Compared with the traditional score priority decision-making method, this step first introduces the score rhythm judgment logic, which expands the score behavior from static numerical sorting to dynamic rhythm identification, and builds a time-driven anchor point for subsequent pulse compensation.

[0041] In the score time period identified as the oscillation critical interval, an inverse phase reward pulse is actively injected into the strategy score change stream, an inhibition signal band is constructed before the strategy jump path is generated by using the reverse feedback of the score direction, and a targeted rhythm buffer is formed. The injection amplitude of the inverse phase pulse is set according to the current score change speed and the score difference of the candidate path, if the score change rate is higher than 0.02 / second, and the score difference between the candidate path and the current path is less than 0.01, the amplitude of the inverse phase pulse is set to 0.007, the injection period is every 200 milliseconds, and the injection direction is opposite to the original score change direction. In the strategy score stream, when the current score trend approaches the anchor budget boundary, the pulse injection action will be initiated 300 milliseconds in advance, forming a score inhibition window, to avoid the strategy score crossing the barrier boundary and entering the jump state. For example, if the current strategy score is 0.832, the anchor upper limit is 0.850, and the candidate strategy score is 0.848, the inverse phase reward pulse will automatically intervene when the score exceeds 0.845, slightly reducing the candidate strategy score, so that it does not immediately constitute a jump advantage. In the score curve, the pulse appears as a rhythmic interference signal, making the score rise appear a "flat top" section, thereby delaying the strategy switching. The rhythm intervention mechanism has no public literature reported in the prior art, and the traditional path strategy relies on score hard switching without a strategy stable period, which is easy to cause oscillation. The present application significantly prolongs the strategy residence time by the inverse phase reward injection mechanism, realizes rhythm traction and score stability improvement.

[0042] On the basis of the rhythm regulation of the score response to the counter-phase reward pulse pair, the damping response migration mechanism of the action layer is further activated, the speed is slowly started, the acceleration is gradually increased, and the posture transition curve is introduced in the execution path generation stage to realize the physical slow release of the packing behavior. In actual deployment, once the current strategy enters the score jump waiting to switch state, the execution action does not start at full speed immediately, but a "slow start window" is set, the slow start window time is generally 300 to 500 milliseconds, the starting speed is reduced to 60% of the normal starting speed, the acceleration control is within 0.4 m / s 2, and the action initial segment suppression mechanism is introduced to make the grabbing trajectory keep stable direction and slow growth speed in the first 20% of the path. For example, the original starting speed of the path is 300 mm / s, which is adjusted to 180 mm / s when the damping response mechanism is activated, and a linear slow change curve is forced to be executed within the range of 20 cm in the front segment of the trajectory to prevent position from changing dramatically at the critical point of strategy switching. At the same time, during the starting process of the mechanical arm gripper, the posture angle adjustment range is limited within 10 degrees, the continuous posture output is maintained, and the violent angle adjustment is avoided. The core of the damping response mechanism is to physically buffer the execution jump risk caused by the score instability of the strategy layer, to absorb and convert the impact of strategy shock on the action into slow release behavior, and to greatly reduce the impact load of the execution end caused by the change of the packing path selection. Unlike the instant response instruction triggering method in the prior art, the present application constructs an execution response control mechanism driven by the strategy score and fed back by the action inertia, realizes the time coupling of the strategy layer and the execution layer, and has a significant stability improvement effect.

[0043] On the basis of the composite control of the score reverse phase pulse and the action damping response, a dynamic closed loop control structure covering score monitoring, strategy evaluation, path verification, action execution and feedback correction is constructed, continuously tracking the score flow and the execution path behavior, and updating the execution action list in real time. The closed loop structure takes the update time of the execution action list as the core beat point, and the score change, strategy state and action execution state are synchronized and aligned, and the list generation rhythm is dynamically adjusted according to the fluctuation trend of the score curve: when the score fluctuation amplitude is less than 0.008 for three consecutive periods and the path switching does not occur, the list refresh period is shortened to every 500 milliseconds; when the score fluctuation intensifies and the number of switching paths increases, the refresh period is extended to 800 milliseconds, so as to increase the strategy response and execution verification buffer time. In addition, the current score trend, the remaining time of strategy execution, the remaining length of action buffer interval, whether it is in the shadow track stage, whether the delay withdrawal mechanism is triggered are dynamically collected when the list is refreshed, and the current action list is re-calculated whether it enters the "controlled stable area" after comprehensive evaluation. If the evaluation result is yes, the current path execution is preferentially maintained; if the evaluation result is no, the list is re-generated, and the new action content is output according to the current score sorting, stability index and damping response state. Finally, the strategy score behavior, action generation path and actual execution behavior form a stable feedback loop, realizing the closed loop control chain of score driving-rhythm adjustment-action buffer-execution confirmation, so that the whole packing behavior has the ability of self-sensing, self-correction and self-inhibition to score disturbance. Compared with the traditional score dominant strategy structure, the closed loop mechanism integrates score regulation, rhythm response and physical feedback, has high stability, self-adaptability and continuous control ability, and is a key component of the invention to realize the high robustness intelligent packing strategy.

[0044] The present application realizes the quantitative analysis and cause positioning of the unstable behavior of the strategy by constructing a time sequence observation system for the whole process from decision to execution, accurately collecting the score gradient, action withdrawal behavior and energy consumption change; combined with the causal coherence decomposition and counterfactual disturbance playback chain, the present application first realizes the graph modeling and visual interpretation of the shock cause of the packing strategy; further, by introducing the inertia anchoring budget and the minimum strategy residence threshold, a barrier energy landscape with behavior stickiness is constructed, supplemented by a hysteresis mechanism and a time consistency penalty term, which fundamentally suppresses the high-frequency strategy jump tendency; at the execution layer, an execution buffer cone is constructed and a shadow grabbing trajectory and a delay threshold are deployed, so that the physical action has the absorption capacity and buffer characteristics of score disturbance; finally, by injecting reverse phase reward pulses in the time and frequency domains of the strategy response, the damping migration mechanism of the execution behavior is effectively activated, the active suppression and rapid extinguishing of the strategy shock are realized, and a closed loop stable intelligent packing control process is formed. The overall scheme realizes multi-layer collaborative shock suppression from strategy learning, path evaluation to action execution, effectively improves the continuity, stability and mechanical structure life of the packing process, and significantly reduces the execution energy consumption and task failure rate.

[0045] Certain exemplary embodiments of the present application have been described above by way of illustration, and it is to be understood that equivalent alterations and modifications will occur to those skilled in the art in view of the foregoing without departing from the spirit and scope of the present application. Therefore, it is the intent that each element of the specification be incorporated by reference and that reference to a document filed prior to the priority date of this application be permitted, even though an item can not have been specifically referenced in the above description. In regard to the processes, methods, and / or algorithms disclosed, those skilled in the art will recognize that the functions required to be performed can be carried out by a variety of hardware and / or software means. Certain portions of the detailed description have been presented in terms of algorithms and symbolic representations of operations on data bits within a computer memory. These algorithmic descriptions and representations are the means used by those skilled in the data processing arts to most effectively convey the

Claims

1. An automated case packing method based on reinforcement learning and dynamic search, characterized in that, The method comprises the following steps: S001, establish a time sequence observation layer from decision to execution, collect policy score gradient, action withdrawal times and execution energy consumption, and generate baseline data representing container loading sequence shock; S002, based on the baseline data, perform causal coherence decomposition, identify the approximate optimal path switching cluster, calculate the switching frequency, residence time and withdrawal cost, and construct the shock cause atlas; S003, through the shock cause atlas, introduce the counterfactual playback chain, inject disturbance sequence to the key time window, compare the path evolution, identify the steady state interval, calculate the inertia anchoring budget and the minimum residence threshold; S004, according to the anchoring budget and the residence threshold, construct the potential energy landscape in the policy generation, superimpose the hysteresis mechanism and the time consistency penalty, limit the high frequency switching of the path, and form the smooth transition area; S005, establish an execution buffer cone in the smooth transition area, generate shadow grabbing trajectory and withdrawal delay threshold, and link the path evaluation module to output the single-step execution action list; S006, based on the execution action list, perform dynamic regulation, inject inverse phase reward pulse in the policy time-frequency domain, activate the damping response migration mechanism, and extinguish the container loading path shock.

2. The automated bin packing method based on reinforcement learning and dynamic search according to claim 1, wherein, Step S001 comprises: On the basis of establishing the time sequence observation layer between the decision process and the execution process, the container strategy score gradient, the action withdrawal times and the energy consumption change data in the execution stage are collected; Based on the score gradient, the score value, the score timestamp, the target cargo number, the target grabbing position coordinates and the placing position coordinates of each container loading strategy are recorded, and a one-to-one mapping relationship between the score sequence and the action instruction is established; Based on the action withdrawal times, the action withdrawal behavior caused by strategy change or execution failure in each container loading task is recorded, and the withdrawal reason, the withdrawal path length, the withdrawal duration and the withdrawal grabbing and placing position coordinates are collected; Based on the execution energy consumption change data, the current, voltage, speed and driving load change information in the execution task are collected, and the score sequence and the withdrawal behavior are synchronized according to the timestamp, and the behavior baseline data structure for identifying container loading sequence shock is constructed.

3. The automated bin packing method based on reinforcement learning and dynamic search of claim 1, wherein, Step S002 comprises: After collecting the score gradient, the action withdrawal times and the energy consumption change data and generating the baseline data representing the container loading sequence shock, the stable score stage is divided based on the continuity and fluctuation amplitude of the score value, and the strategy records with similar scores but different container loading behaviors are extracted in each score stage to construct the strategy path switching cluster; The time sequence behavior tracking is performed on each strategy path switching cluster, the duration, switching frequency and withdrawal behavior corresponding to each switching event are extracted, and the average strategy residence time and withdrawal cost of each strategy path switching cluster are calculated; The energy consumption data corresponding to each strategy path switching cluster is collected, the current, power, temperature rise and end vibration change amplitude before and after path switching are analyzed, and the high energy consumption response type path switching cluster is identified; The score change behavior, strategy path switching feature and energy consumption response data are fused to construct the shock cause atlas reflecting the cause of container loading sequence shock, and the atlas covers the complete time sequence and path number evolution process.

4. The automated bin packing method based on reinforcement learning and dynamic search according to claim 3, wherein, In constructing the oscillation cause map, for each strategy path switching event, the switching start and end time, corresponding path number, switching frequency and energy consumption change data within five seconds after switching are marked, and different colors are used to identify power abnormalities, end vibration enhancement and action withdrawal situations to realize the visual marking of high oscillation risk areas.

5. The automated bin packing method based on reinforcement learning and dynamic search of claim 1, wherein, Step S003 includes: After constructing the oscillation cause map reflecting the oscillation reasons of the packing sequence, based on the scoring change gradient, path jump frequency and energy consumption fluctuation amplitude, a key time window is selected; Insert score disturbance, packing sequence disturbance and target cargo disturbance into the key time window respectively, track the strategy path change, action withdrawal and energy consumption response after disturbance, and extract the strategy behavior interval that remains stable or recovers; Compare the strategy evolution sequence before and after disturbance to construct the mapping relationship between disturbance and strategy stability, and identify the boundary conditions affecting strategy stability; Based on the stable behavior interval, calculate the strategy inertia anchoring budget and the minimum strategy stay threshold, which are used as control parameter inputs for subsequent construction of path switching inhibition mechanism.

6. The automated bin packing method based on reinforcement learning and dynamic search of claim 1, wherein, Step S004 includes: After calculating the inertia anchoring budget and the minimum strategy stay threshold of the strategy path, construct the path potential barrier energy map, and determine the path score potential barrier strength based on the offset of the current score and the anchoring budget boundary; Introduce a hysteresis mechanism in the path score judgment process, set the score jump radius and hysteresis period by analyzing the score trend accumulation amplitude and the continuity of the sliding score window, and inhibit the path switching caused by short-term score fluctuations; Based on the difference between the actual execution time length of the strategy and the minimum strategy stay threshold, introduce a time consistency penalty term to impose a penalty value on the replacement path whose score does not reach the execution time length, to avoid premature strategy jumping; Combine the path potential barrier energy map, the hysteresis mechanism and the time consistency penalty term to construct a smooth transition region of strategy score, limit the frequency of strategy path switching, and improve the stability of the packing path.

7. The automated bin packing method based on reinforcement learning and dynamic search of claim 1, wherein, Step S005 includes: When the strategy score is in the smooth transition region, construct an execution buffer cone structure with the current grabbing point as the starting point and the target placement point as the direction, and set the cone angle and length according to the score gradient and historical jump; Generate a shadow grabbing trajectory in the execution buffer cone, which meets the requirements of movement continuity, spatial obstacle avoidance and attitude consistency, and dynamically determines whether to convert into actual action according to the score trend and hysteresis; Set an action withdrawal delay threshold, determine the withdrawal waiting time based on the score fluctuation period, score difference and executed time, and inhibit immediate jumping behavior when the withdrawal condition is not met; After completing path verification and delay evaluation, re-integrate the score, path stability and execution state to output a structured single-step execution action list for guiding the physical action execution of the current strategy.

8. The automated bin packing method based on reinforcement learning and dynamic search according to claim 7, characterized in that, The shadow grabbing trajectory is only allowed to be converted into actual execution action when it meets the hysteresis trigger condition in continuous score update and the path stability index reaches the set threshold, and the conversion into actual execution action can only be started after the withdrawal delay timer ends, to ensure sufficient action stability verification before strategy switching.

9. The automated bin packing method based on reinforcement learning and dynamic search of claim 1, wherein, Step S006 includes: Based on the execution action list output by the execution buffer cone, the score fluctuation frequency and direction reversal characteristics are extracted within the score update cycle, and the high-frequency oscillation interval is identified; Within the identified oscillation interval, an inverse-phase reward pulse opposite to the score direction is injected into the score change stream, a score suppression window is constructed, and the strategy switching behavior is delayed; When the score trend approaches the switching state, the damping response mechanism of the action execution stage is activated, the initial speed, acceleration and attitude change of the path are controlled, and the action release is realized; Based on the score trend, the strategy execution state and the damping feedback result, the execution action list is dynamically updated, the closed-loop control structure of score-driven, rhythm-controlled and action-feedback is constructed, and the adaptive extinguishing and stable execution of score disturbance are realized.

10. The automated bin packing method based on reinforcement learning and dynamic search according to claim 9, wherein, The injection cycle and amplitude of the inverse-phase reward pulse are dynamically adjusted according to the score change rate and the score difference of the candidate path, and a preset injection point in advance is set before the score trend approaches the anchor budget boundary, which is used to form a flat-top area of score rise and delay the triggering of strategy jump behavior.

Citation Information

Patent Citations

  • Boxing method based on deep reinforcement learning

    CN111695700A

  • Deep reinforcement learning network system

    CN112884126A

  • Automatic boxing method based on reinforcement learning and dynamic search

    CN114548855A

  • Multi-vehicle collaborative boxing method based on sequence-to-sequence strategy network deep reinforcement learning model

    CN116957437A

  • Logistics robot scheduling method based on deep reinforcement learning

    CN118536783A