An automated case packing method based on reinforcement learning and dynamic search

By constructing a time-series observation layer and a potential barrier energy landscape during the automated packing process, the causes of packing oscillations are identified. An inertial anchoring budget and hysteresis mechanism are introduced to generate shadow tracking trajectories, solving the problem of frequent packing path switching, achieving stable control of the packing process, and improving efficiency and mechanical structure lifespan.

CN121072918BActive Publication Date: 2026-01-23华商国际工程有限公司 +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511595970.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-04
Publication Date
2026-01-23
Estimated Expiration
2045-11-04

AI Technical Summary

Technical Problem

In existing automated packing processes, the frequent switching of packing paths under dynamic search conditions leads to low execution efficiency of the robotic arm, accelerated hardware fatigue, and affects system stability and lifespan.

Method used

By constructing a time-series observation layer from decision-making to execution, and using container strategy scoring gradients, action withdrawal counts, and energy consumption changes, the causes of oscillations in the packing sequence are identified. An inertial anchoring budget and minimum strategy dwell threshold are introduced to construct a barrier energy landscape. By superimposing hysteresis mechanisms and time consistency penalties, shadow capture trajectories and withdrawal delay thresholds are generated to achieve stable control of the packing path.

Benefits of technology

It significantly improves the continuity and stability of the packing process, reduces energy consumption and task failure rate, and extends the life of mechanical structures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121072918B_ABST
    Figure CN121072918B_ABST
Patent Text Reader

Abstract

The application discloses an automatic packing method based on reinforcement learning and dynamic search, and relates to the technical field of intelligent manufacturing and logistics automation, and comprises the following steps: a time sequence observation layer from decision to execution is established, strategy score gradient, action withdrawal times and execution energy consumption are collected, and baseline data representing packing sequence oscillation is generated; based on the baseline data, causal coherence decomposition is carried out, an approximately optimal path switching cluster is identified, switching frequency, residence time and withdrawal cost are calculated, and an oscillation cause atlas is constructed. Through the construction of score and execution whole process monitoring, causal decomposition, disturbance playback, energy barrier modeling, buffer execution mechanism and inverse phase shock suppression strategy, the application realizes the identification, suppression and extinguishing of the packing path oscillation, forms a stable closed loop of strategy, effectively improves the packing process stability and execution efficiency, and reduces the energy consumption and failure rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent manufacturing and logistics automation technology, specifically to an automated packing method based on reinforcement learning and dynamic search. Background Technology

[0002] "Automated packing based on reinforcement learning and dynamic search" refers to the introduction of artificial intelligence reinforcement learning algorithms and dynamic search mechanisms into the automated packing process. This enables the packing system to autonomously learn and optimize packing strategies in response to constantly changing cargo sizes, shapes, and order requirements. Reinforcement learning trains the model through a "trial and error-feedback-optimization" approach, gradually teaching it how to maximize loading efficiency and improve stability within limited container space. Dynamic search, on the other hand, continuously adjusts, corrects, and optimizes the packing path and sequence based on the real-time arrangement of goods and changes in remaining space during the actual packing process. The combination of these two methods not only avoids the inefficiency caused by traditional packing methods relying on fixed rules but also enables intelligent adaptation to complex and changing scenarios, thereby significantly improving space utilization, reducing transportation costs, and enhancing the overall efficiency of automated logistics.

[0003] The existing technology has the following shortcomings:

[0004] In existing technologies, automated packing processes generally rely on path optimization mechanisms based on search algorithms. However, under dynamic search conditions, when the evaluation results of multiple packing paths are close to optimal, the algorithm often switches frequently between these paths, causing oscillations in the packing order. This oscillation phenomenon causes the robotic arm of the actuator to be in a state of repeated grasping and retraction, resulting in the accumulation of a large number of invalid movements. This not only significantly reduces the efficiency of packing execution but also causes hardware fatigue due to excessive start-stop of mechanical components, accelerating the wear and failure of key components, thus having a serious adverse impact on the stability and service life of the system.

[0005] The information disclosed in the background section is only intended to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0006] The purpose of this invention is to provide an automated bin packing method based on reinforcement learning and dynamic search to solve the problems mentioned in the background art.

[0007] To achieve the above objectives, the present invention provides the following technical solution: an automated bin packing method based on reinforcement learning and dynamic search, comprising the following steps:

[0008] S001, establish a time-series observation layer from decision-making to execution, collect strategy scoring gradient, action withdrawal count and execution energy consumption, and generate baseline data representing packing order oscillation;

[0009] S002, based on baseline data, perform causal coherence decomposition, identify near-optimal path switching clusters, calculate switching frequency, dwell time and withdrawal cost, and construct an oscillation cause map;

[0010] S003 introduces a counterfactual replay chain through the oscillation causation map, injects a disturbance sequence into the key time window, compares the path evolution, identifies the steady-state interval, and calculates the inertial anchoring budget and minimum dwell threshold.

[0011] S004, based on the anchoring budget and dwell threshold, constructs a barrier energy landscape during strategy generation, superimposes a hysteresis mechanism and time consistency penalty, restricts high-frequency path switching, and forms a smooth transition area;

[0012] S005, establish an execution buffer cone in the smooth transition area, generate the shadow capture trajectory and withdrawal delay threshold, and the linkage path evaluation module outputs a list of single-step execution actions.

[0013] S006 dynamically adjusts the execution action list, injects a reverse-phase reward pulse into the time-frequency domain of the strategy, activates the damped response migration mechanism, extinguishes the packing path oscillation, and achieves a stable closed loop process.

[0014] Preferably, step S001 includes:

[0015] Based on the establishment of a time-series observation layer between the decision-making process and the execution process, data on container strategy scoring gradient, number of action withdrawals, and energy consumption changes during the execution phase are collected.

[0016] Based on the scoring gradient, the scoring value, scoring timestamp, target cargo number, target grab position coordinates and placement position coordinates of each packing strategy are recorded, and a one-to-one mapping relationship between the scoring sequence and the action instructions is established.

[0017] Based on the number of action withdrawals, record the action withdrawal behavior caused by strategy changes or execution failures in each packing task, and collect the withdrawal reason, withdrawal path length, withdrawal duration, and coordinates of the withdrawal grab and placement positions.

[0018] Based on the energy consumption change data, information on changes in current, voltage, speed and drive load during the execution of tasks is collected, and the scoring sequence and withdrawal behavior are synchronized according to the timestamp to construct a behavioral baseline data structure for identifying packing sequence oscillations.

[0019] Preferably, step S002 includes:

[0020] After collecting data on scoring gradient, number of action withdrawals, and energy consumption changes, and generating baseline data to characterize packing order oscillations, stable scoring phases are divided based on the continuity and fluctuation amplitude of the scoring values. Within each scoring phase, strategy records with similar scores but different packing behaviors are extracted to construct a strategy path switching cluster.

[0021] Perform time-series behavior tracking for each strategy path switching cluster, extract the duration, switching frequency and post-switching withdrawal behavior corresponding to each switching event, and calculate the average strategy dwell time and withdrawal cost for each strategy path switching cluster.

[0022] Collect energy consumption data corresponding to each strategy path switching cluster, analyze the changes in current, power, temperature rise and end vibration before and after path switching, and identify high energy consumption response path switching clusters.

[0023] By integrating scoring change behavior, strategy path switching characteristics, and energy consumption response data, an oscillation cause map is constructed to reflect the causes of packing order oscillations. The map covers the complete time series and path number evolution process.

[0024] Preferably, when constructing the oscillation cause map, for each strategy path switching event, the start and end time of the switching, the corresponding path number, the switching frequency, and the energy consumption change data within five seconds after the switching are marked. Different colors are used to mark power anomalies, end-point vibration enhancement, and action withdrawal, so as to achieve visual marking of high oscillation risk areas.

[0025] Preferably, step S003 includes:

[0026] After constructing an oscillation causal map to reflect the reasons for packing sequence oscillations, key time windows are selected based on the gradient of score changes, the frequency of path jumps, and the amplitude of energy consumption fluctuations.

[0027] Within the critical time window, scoring disturbances, packing order disturbances, and target cargo disturbances are inserted respectively. The changes in strategy path, action withdrawal, and energy consumption response after the disturbances are tracked, and the strategy behavior intervals that maintain or restore stability are extracted.

[0028] By comparing the policy evolution sequences before and after the perturbation, a mapping relationship between the perturbation and policy stability is constructed, and the boundary conditions affecting policy stability are identified.

[0029] The inertial anchoring budget and minimum policy dwell threshold are calculated based on the stable behavior interval and used as control parameter inputs for the subsequent construction of path switching suppression mechanisms.

[0030] Preferably, step S004 includes:

[0031] After calculating the inertial anchoring budget and minimum policy dwell threshold of the policy path, a path barrier energy map is constructed, and the path score barrier strength is determined based on the offset between the current score and the anchoring budget boundary.

[0032] A hysteresis mechanism is introduced in the path scoring process. By analyzing the cumulative magnitude of the scoring trend and the continuity of the sliding scoring window, the scoring jump radius and hysteresis period are set to suppress path switching caused by short-term scoring fluctuations.

[0033] Based on the difference between the actual execution time of the strategy and the minimum strategy dwell threshold, a time consistency penalty term is introduced to apply a penalty value to the replacement path whose score has not reached the execution time, so as to prevent the strategy jump from happening prematurely.

[0034] By combining the path barrier energy map, hysteresis mechanism, and time consistency penalty term, a smooth transition region for policy scoring is constructed, which limits the frequency of policy path switching and improves the stability of packing path.

[0035] Preferably, step S005 includes:

[0036] When the strategy score is in the smooth transition region, an execution buffer cone structure is constructed, with the current grab point as the starting point and the target placement point as the direction. The cone angle and length are set according to the score gradient and historical jumps.

[0037] A shadow grab trajectory is generated within the execution buffer cone. The trajectory meets the requirements of movement continuity, spatial obstacle avoidance and posture consistency. Based on the scoring trend and hysteresis judgment, it is dynamically determined whether to convert it into an actual action.

[0038] Set an action withdrawal delay threshold and determine the withdrawal waiting time based on the score fluctuation cycle, score difference and execution time. Suppress immediate jump behavior when the withdrawal conditions are not met.

[0039] After completing path verification and latency assessment, the scoring, path stability, and execution status are re-integrated to output a structured list of single-step execution actions to guide the physical execution of the current strategy.

[0040] Preferably, the shadow capture trajectory is only allowed to be converted into an actual action when the hysteresis trigger condition is met in the continuous score update and the path stability index reaches the set threshold. Furthermore, the conversion into an actual action can only be started after the withdrawal delay timer has expired, so as to ensure that the action stability verification before the strategy switch is sufficient.

[0041] Preferably, step S006 includes:

[0042] Based on the list of execution actions output by the execution buffer cone, the frequency and direction reversal features of score fluctuations are extracted within the score update cycle to identify high-frequency oscillation intervals;

[0043] Injecting a reverse-phase reward pulse, opposite to the scoring direction, into the scoring change stream within the identified oscillation range creates a scoring suppression window, delaying strategy switching behavior.

[0044] When the scoring trend approaches the switching state, the damping response mechanism of the action execution phase is activated to control the initial velocity, acceleration and attitude changes of the path, thereby achieving action mitigation.

[0045] The execution action list is dynamically updated based on the scoring trend, strategy execution status and damping feedback results. A closed-loop control structure of strategy scoring-driven, rhythm regulation and action feedback is constructed to achieve adaptive extinguishing of scoring disturbances and stable execution.

[0046] Preferably, the injection cycle and amplitude of the reverse phase reward pulse are dynamically adjusted according to the rate of change of the score and the score difference of the candidate path, and a pre-injection point is preset before the score trend approaches the anchor budget boundary to form a score rise plateau area and delay the triggering of the strategy jump behavior.

[0047] The technical effects and advantages provided by the present invention in the above technical solution are as follows:

[0048] This invention constructs a time-series observation system covering the entire process from decision-making to execution, accurately collecting scoring gradients, action withdrawal behaviors, and energy consumption changes to achieve quantitative analysis and causal localization of unstable policy behavior. Combining causal coherence decomposition and counterfactual perturbation replay chains, it achieves for the first time a graphical modeling and visual explanation of the causes of packing strategy oscillations. Furthermore, by introducing an inertial anchoring budget and a minimum policy dwell threshold, it constructs a potential barrier energy landscape with behavioral viscosity, supplemented by a hysteresis mechanism and a time consistency penalty term, fundamentally suppressing high-frequency policy jump tendencies. At the execution layer, it constructs an execution buffer cone and deploys shadow-grabbing trajectories and withdrawal delay thresholds, enabling physical actions to absorb and mitigate scoring perturbations. Finally, by injecting inverse-phase reward pulses into the time and frequency domains of the policy response, it effectively activates the damping migration mechanism of execution behavior, achieving proactive suppression and rapid extinguishing of policy oscillations, thus forming a closed-loop stable intelligent packing control process. The overall solution achieves multi-layered collaborative vibration suppression from strategy learning and path evaluation to action execution, effectively improving the continuity, stability and mechanical structure life of the packing process, and significantly reducing execution energy consumption and task failure rate. Attached Figure Description

[0049] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.

[0050] Figure 1 This is a flowchart of an automated bin packing method based on reinforcement learning and dynamic search according to the present invention. Detailed Implementation

[0051] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that the description of this disclosure will be more complete and fully convey the concept of the exemplary embodiments to those skilled in the art.

[0052] This invention provides, for example Figure 1 An automated bin packing method based on reinforcement learning and dynamic search is shown, comprising the following steps:

[0053] S001, establish a time-series observation layer between the decision-making process and the execution process, and collect data on container strategy scoring gradient, number of action withdrawals and energy consumption changes during the execution phase to form baseline data for characterizing the oscillation of the packing sequence;

[0054] To address the execution oscillation issue caused by frequent switching of packing paths within near-optimal regions, a detailed time-series observation process needs to be established throughout the entire process from strategy decision-making to physical execution. This involves multi-dimensional data collection of three types: score changes, action withdrawals, and energy consumption fluctuations, to construct a baseline description structure for packing oscillations. The specific implementation steps are as follows:

[0055] A full-process time observation path from policy generation to action execution is constructed. Specifically, this involves recording the score value corresponding to each packing decision made by the reinforcement learning policy in real time. This score value is a comprehensive evaluation index for that packing state, with a value range between 0 and 1, where a value closer to 1 indicates a better packing result. The score value generation process is based on a weighted calculation of four specific quantitative dimensions: packing space utilization, target cargo posture stability, post-packing space remaining distribution balance, and path execution length. A millisecond-level timestamp is appended with each score generation to track the accurate time node of score generation. Immediately after each score generation, the target cargo number determined by the policy, the three-dimensional coordinates of the target grasping position (X, Y, Z axis positions in millimeters), and the specific spatial location of the target placement point within the container (defined by the three-dimensional coordinates of the container's length, width, and height) are recorded synchronously. Furthermore, the policy score is bound to the corresponding action number, establishing a unique mapping table to ensure a one-to-one correspondence between the score value and subsequent actions, thus forming a continuous score sequence data chain. This sequence will be fused with other execution observation data in subsequent steps to characterize the intrinsic relationship between policy fluctuations and behavioral fluctuations.

[0056] The detailed trajectory of each action withdrawal during the execution process is recorded in real time. Specifically, this includes the complete process of the action being stopped due to reasons such as strategy re-evaluation, target posture recognition failure, cargo size exceeding limits, or the gripping path being blocked by obstacles after the packing robot receives the packing instruction and starts execution. Each action withdrawal event needs to clearly record the following five parameters: (1) the trigger source for the withdrawal, such as whether it is due to score update, whether target mismatch is detected, or whether path verification fails; (2) the packing instruction number corresponding to the withdrawal action; (3) the three-dimensional coordinates of the gripping point and placement point of the current execution action; (4) the length of the robot's movement path trajectory before withdrawal (in millimeter), obtained by integration through the position sensor; (5) the total time used for the withdrawal action (in milliseconds), used for subsequent evaluation of the degree of action redundancy and execution jitter frequency. The complete record of each withdrawal action will be associated with the strategy score status in the current time period and marked on the score sequence. By statistically analyzing the density of withdrawal actions in certain time periods with large score fluctuations, it is possible to identify the unstable area of ​​the strategy and further deduce the initial symptoms of packing oscillation.

[0057] Continuous energy consumption monitoring is conducted throughout the entire process of the action execution, specifically including current input values, voltage stability, drive motor operating temperature, robotic arm load changes, and end effector speed curves from the start of the packing instruction to the completion of object placement. Energy consumption measurement employs a high-frequency sampling method, acquiring current and voltage parameters every 50 milliseconds and recording the load state of the drive unit and the linear velocity change of the end effector every 200 milliseconds. For example, if a grasping task is being performed, and the robotic arm's current is 1.8A and speed is 120 mm / s during its movement from the initial position to the target cargo position, but the current rises to 2.6A and the speed fluctuates to ±45 mm / s during the withdrawal phase, it can be preliminarily determined that the task exhibits an energy consumption anomaly caused by strategy oscillation. All energy consumption data is simultaneously timestamped and categorized according to the grasping action number, and linked to the corresponding score and withdrawal data. By constructing a triple data linkage relationship of "rating change - withdrawal behavior - energy consumption curve", it is possible to identify whether there is a significant surge in energy consumption during the period of frequent policy updates, thereby transforming the abstract policy oscillation into a quantifiable physical cost measurement and enhancing the transparency of the packing system behavior.

[0058] The scoring sequence, withdrawal trajectory data, and energy consumption curve are fused using the packing task number as an index to construct a baseline data structure for oscillation identification in packing behavior. Specifically, a time synchronization mechanism maps the three types of data onto the same time axis, and a sliding time window is set to perform alignment analysis for different task cycles. Each window is set to a width of 5 seconds, and the scoring change frequency, total number of withdrawals, average energy consumption, and standard deviation contained within it serve as baseline feature vectors. Oscillation intensity indicators for each time period are constructed through feature overlay. For example, when the scoring changes more than 5 times, the withdrawal action exceeds 3 times, and the energy consumption variance exceeds a preset threshold of 0.08, the window is identified as a high-risk oscillation zone. The baseline data is further assigned a packing sequence number and a timestamp to track the stability of the packing sequence at different stages, providing complete, detailed, and structured data support for subsequent oscillation cause identification, strategy correction guidance, and execution path optimization.

[0059] S002, based on the collected baseline data, perform causal coherence decomposition to identify the near-optimal packing path switching clusters in the packing sequence, calculate the switching frequency, strategy dwell time and withdrawal cost of each switching cluster, and construct an oscillation causal map to reflect the cause of packing sequence oscillations;

[0060] To deeply identify the specific causes of oscillating behavior due to frequent strategy switching during the packing process, it is necessary to construct path switching clusters based on baseline data, combined with scoring sequences, action withdrawal trajectories, and energy consumption changes, extract oscillation indicators, and visualize the causal relationships. This includes the following steps:

[0061] For the established scoring gradient time series, the entire scoring curve is divided into multiple non-overlapping scoring stages based on the continuity and fluctuation trend of scoring value changes. Each stage has a time span of no less than five seconds, and the scoring change is required to be no more than 0.02 to ensure relative stability of the scoring within the divided intervals. For each scoring stage, all packing decision records with scoring values ​​falling within the 0.05 interval but with significant differences in execution behavior are extracted. Specific behavioral differences are determined by the target cargo number, the spatial coordinates of the target grabbing point, and the 3D annotation of the target placement location. When the difference between any two strategy scores is less than 0.02 but the corresponding packing object numbers are inconsistent, or the difference between the grabbing point and the placement point exceeds 20 mm, it is considered a clear behavioral difference. Based on this, these strategy records are combined into strategy switching clusters. Each switching cluster contains no less than three strategy units with similar scores but different behaviors. The switching cluster numbers are sequentially increased in time for subsequent behavior tracking. This construction process breaks through the existing path selection method that relies solely on scoring value sorting, introducing a new perspective of scoring and behavior matching, accurately capturing potential oscillation sources where scores are similar but behavior fluctuations are severe.

[0062] For each constructed strategy switching cluster, time-series behavior tracking is performed to analyze its switching characteristics throughout the entire packing process. During tracking, the time point of first appearance of each strategy unit and its corresponding strategy number are extracted first. Then, the duration of the strategy after execution is recorded until it is switched to another strategy within the switching cluster. Each jump from the current strategy to another strategy is counted as a switch, and the switch occurrence time, starting strategy number, target strategy number, and jump direction are recorded. For example, a jump from strategy A to strategy B, where strategy A lasts 1.2 seconds and the jump time is at 18 seconds, is recorded as "Strategy A → Strategy B, lasting 1.2 seconds, switched at T=18s". Simultaneously, the cumulative number of activations and average duration of each strategy within the switching cluster are counted, and whether an action withdrawal event immediately follows each switch is extracted. If an action withdrawal event occurs within five seconds after a strategy switch, it is considered an execution instability caused by that strategy switch, and the total number of withdrawals is accumulated and marked in the switching record. The final output is a set of behavioral metrics for each switching cluster, including switching frequency (number of switches per unit time), average policy dwell time (average duration of each switch), and withdrawal cost after policy switching (average number of withdrawals triggered per switch), which serve as the core input for oscillation assessment. This method differs from the traditional approach of existing binning systems that only analyze policy success rate or policy score fluctuations. It introduces for the first time the linkage analysis of behavioral chains and physical feedback, comprehensively revealing the dynamic coupling relationship between the policy layer and the execution layer.

[0063] For each strategy switching cluster, all associated energy consumption data is extracted for the corresponding packing behavior segment, and the direct impact of switching behavior on energy consumption is analyzed. Specifically, data on execution current, voltage, drive shaft power, servo motor load torque, and end-effector speed changes are extracted within the complete time interval from strategy activation to strategy replacement. The sampling frequency is set to once every 50 milliseconds, and the duration corresponds to the complete length of the strategy dwell segment. Using strategy switching as the boundary, the average current value, peak power, temperature rise, and end-effector vibration amplitude are statistically analyzed before and after switching. The vibration amplitude is calculated using the standard deviation of the three-axis speed. If significant differences occur before and after switching, such as a power increase exceeding 10%, a doubling of end-effector vibration amplitude, or severe current fluctuations accompanied by abnormal temperature increases, it is marked as a "high-energy-consumption response-type switching." Comparing the energy consumption performance of multiple switching clusters reveals that some clusters exhibit a continuous upward trend in energy consumption when switching behavior is frequent, while others remain relatively stable. Combining this with the behavioral characteristics extracted in the second step, high-oscillation-risk switching clusters can be effectively distinguished from general path fluctuations. This method not only establishes a clear relationship between strategy switching and physical resource consumption, but also realizes a causal closed loop that infers the reliability of the strategy layer from the feedback of the action layer, thus overcoming the problem of traditional score-driven decision-making lacking physical evaluation constraints.

[0064] Based on the above-mentioned scoring fluctuation behavior, strategy switching statistical characteristics, and energy consumption response indicators, an oscillation cause map is constructed to intuitively display the evolution of strategy oscillation behavior in the time and path dimensions. The horizontal axis of the map represents the standard time series in seconds, and the vertical axis represents the strategy path number. Each strategy switching cluster is identified by an independent color band, and the start and end time intervals and switching frequency of the strategy switching are marked. Dashed lines represent each path jump connection. At each jump point, the energy consumption response within five seconds after the jump is overlaid, including the current increase, the change value of the terminal vibration, and whether the action withdrawal occurred. High-risk indicator areas are marked with different colors. For example, a red dot represents a jump accompanied by a high power peak and action withdrawal, a yellow mark indicates only power anomaly without withdrawal, and green indicates a smooth switch. The map as a whole constitutes a multi-dimensional overlay structure of strategy evolution path map, behavior change map, and energy consumption response map, realizing a visible, traceable, and quantifiable expression of the entire process of oscillation behavior.

[0065] S003 introduces a counterfactual replay chain based on the oscillation cause spectrum, inserts a disturbance sequence into the selected key time window, compares the evolution process of the packing path under each disturbance condition, identifies the stable evolution interval, and calculates the inertial anchoring budget and the minimum strategy dwell threshold.

[0066] To effectively address the high-frequency policy switching problem caused by path score proximity, based on the unstable segments identified in the oscillation causation map, an artificial perturbation sequence is introduced within a critical time window to observe the policy response evolution process, identify the system stability boundary, and thereby calculate the anchoring budget and dwell time required for policy inertial execution. The specific steps include:

[0067] Based on the segments of rapid score changes and frequent path jumps marked in the oscillation causation map, representative key time windows are selected for counterfactual testing. Each time window extends five seconds before and after, lasting a total of ten seconds to ensure coverage of a complete policy switching cycle. Within the selected time window, the following three conditions must be met: the score change gradient exceeds 0.03, the policy path experiences at least two non-repeating jumps within ten seconds, and at least one binning action is withdrawn between path jumps, accompanied by a significant increase in energy consumption (e.g., current fluctuations exceeding the average by more than 20%). Within this window, binning behavior exhibits obvious unstable characteristics, making it suitable as a test base for subsequent counterfactual perturbations. Unlike traditional methods that only evaluate policy output results, this step synchronously binds policy evaluation with execution dynamics, ensuring that counterfactual insertions have actual observable behavioral responses.

[0068] Within the critical time window, multiple perturbation sequences are constructed according to a preset scheme and inserted into the scoring gradient change nodes. The perturbation sequences consist of three types of perturbations: scoring perturbation, packing order perturbation, and target cargo characteristic perturbation. Scoring perturbation simulates strategy changes after a score advantage is enhanced or weakened by manually adjusting the scoring value, for example, changing the original score from 0.83 to 0.87 or 0.79. Packing order perturbation verifies whether strategy switching is affected by the order by changing the sequence of two packing actions within the same window, for example, swapping the packing order at seconds 12 and 14. Target cargo perturbation observes whether the strategy exhibits a switching tendency by replacing the packing object targeted by the current strategy, for example, changing the target cargo volume from small to medium size or altering its edge shape. Each type of disturbance is inserted separately, while the control variables remain consistent. Each type of disturbance must be implemented at least three times. After each disturbance insertion, the policy output path, action execution feedback, energy consumption change trend, and withdrawal behavior are fully tracked to evaluate the policy's immunity and stability critical point under disturbance conditions.

[0069] For the strategy evolution sequences generated after various disturbances, a one-to-one alignment analysis is performed with the original undisturbed sequences to compare whether path switching intensifies, whether the strategy reverts to the original path, whether withdrawal behavior increases or decreases, and whether energy consumption fluctuates, converges, or continues to amplify. Specific judgment criteria are as follows: if the strategy recovers to the pre-disturbance path within three seconds after the disturbance insertion without triggering additional withdrawal behavior, the strategy is considered to have high stability; if a path change occurs after the disturbance, accompanied by a new round of withdrawal actions and an abnormal increase in energy consumption, the strategy is considered to have poor stability under this type of disturbance. To quantify the steady-state trend, the duration of the path after the disturbance, the number of switching events, the withdrawal frequency, and the average and standard deviation of energy consumption are calculated to construct a disturbance-evolution behavior mapping table, marking which disturbance combinations are prone to inducing unstable responses and which paths have convergence under different disturbance conditions. Through extensive disturbance comparison analysis, a set of steady-state evolution intervals is extracted, i.e., the intervals of strategy behavior that remain stable or quickly return to a stable state under the influence of disturbances. This analysis process differs from conventional strategy performance evaluation. It not only measures the quality of the strategy output, but also focuses on the strategy's recovery capability in abnormal disturbance scenarios, providing behavioral boundary basis for subsequent seismic damping parameter design.

[0070] Based on the statistical results of all steady-state evolution intervals, the strategy inertial anchoring budget and minimum strategy dwell threshold are calculated. The inertial anchoring budget is defined as the score fluctuation range, path disturbance tolerance, and target change limit that the current strategy can maintain stable execution under disturbance conditions. For example, the score fluctuation budget is ±0.02, the path jump tolerance is less than once, and the allowable size difference for target cargo replacement does not exceed 10%. The minimum strategy dwell threshold is the minimum execution time required by the strategy without switching intervention, ensuring that the strategy has sufficient time to influence the execution layer before the introduction of disturbance, and is not misled by instantaneous score fluctuations to generate unnecessary switching. This threshold is usually between 2.0 and 3.0 seconds and is dynamically adjusted according to the strategy's historical withdrawal rate and path execution energy consumption fluctuation curve. Finally, the inertial anchoring budget and minimum dwell threshold of each strategy path are encapsulated into a structured parameter set, which serves as the input basis for the next stage of constructing the potential barrier energy landscape and hysteresis control mechanism. This ensures that the strategy will not easily switch when approaching the score fluctuation threshold, thereby improving the overall path stability and suppressing strategy oscillation behavior caused by perturbations.

[0071] S004, based on the inertial anchoring budget and the minimum policy dwell threshold, constructs a potential barrier energy landscape in the packing policy generation stage, superimposes a hysteresis mechanism and a time consistency penalty term, applies regulation and restriction to the policy path switching frequency, and forms a policy smooth transition region;

[0072] To avoid frequent policy jumps when approaching the score change threshold, which could lead to instruction oscillations and action reversals during the packing process, this invention introduces an energy barrier structure based on inertial anchoring budget and minimum dwell threshold during policy generation. Through the interaction of barrier spectrum, hysteresis mechanism, and time consistency penalty, a smooth transition region that can buffer score fluctuations is constructed. Specifically, the following steps are included:

[0073] By combining the extracted inertial anchoring budget and score change data for each packing path, a complete path barrier energy map is constructed. This map uses the offset between the path score and the score stability interval as the vertical axis and the path index number as the horizontal axis, numerically representing the current score state of each path and its inertial anchoring threshold interval. In actual construction, the absolute difference between the current path score and the anchoring budget boundary needs to be calculated. For example, if the score of path P12 is 0.846 and its anchoring budget is ±0.015, then the score change to 0.832 or 0.861 is within the allowable range and is considered a low barrier region; changes exceeding this range indicate a high barrier region. In the map, a color gradient is used to indicate the barrier strength of different paths, with lighter to darker shades representing greater difficulty in transitioning closer to the budget boundary. During the strategy score update process, the map values ​​are recalculated each time to ensure that the score selection has the latest energy gradient reference. This graph differs from the traditional maximum value strategy decision-making mechanism of scoring and ranking. Before selecting the path with the highest score, it first determines whether the current score has crossed the potential barrier threshold. If it has not crossed it, the current strategy remains unchanged. Even if the score is not optimal, it does not jump immediately, thus forming a strategy execution inertia constrained by energy potential.

[0074] Based on the constructed barrier map, a path hysteresis mechanism is superimposed on the strategy scoring and judgment process. This ensures that path switching is no longer directly triggered by a single score fluctuation, but is determined by the persistence and accumulation of the scoring trend. Specifically, a jump radius and hysteresis judgment period are set for the current execution path. For example, for the current execution path P8, the jump radius is set to 0.018, and the hysteresis period is set to 1.2 seconds. During this period, if the cumulative change in the direction of score decline does not exceed 0.018, even if the score is once lower than other paths, switching will not occur. Only when the score continues to decline and accumulates beyond the jump radius is a jump to a higher-scoring path allowed. The hysteresis mechanism uses a sliding scoring window detection method, evaluating the direction of score change every 0.2 seconds based on time. If six consecutive score changes show a downward trend and eventually exceed the set jump radius, the hysteresis switching condition is met. This mechanism, by introducing scoring trend analysis and behavioral inertia parameters, ensures that path switching has an "impulse-triggered" characteristic, overcoming the existing strategy structure of immediate score switching and exhibiting a significant shock-absorbing effect.

[0075] Building upon the hysteresis mechanism, a time consistency penalty is introduced. Based on the relationship between the actual execution time of the strategy and the minimum dwell threshold, a strategy replacement penalty is applied when a new scoring path covers a path before the minimum execution time is met. For example, if path P3 has just been selected and executed for 1.1 seconds, and its minimum strategy dwell threshold is 2.7 seconds, and the current scoring update recommends P4 as the preferred path, then P4's score will be deducted a penalty value proportional to the time difference, such as 0.012 points, to prevent it from having a sufficient advantage in the ranking. The scoring update is only considered valid when the current path has run for 2.7 seconds or more. The penalty mechanism calculates based on both the execution time difference and the path fluctuation frequency, ensuring that each path receives a complete execution cycle and preventing a disconnect between the scoring mechanism and the execution response, further reducing the incentive for switching due to short-term high score fluctuations. This mechanism is the first in path strategy control to dynamically link scoring weights with execution time, differing from traditional static strategy output methods and enhancing the logical consistency between scoring and behavior.

[0076] The barrier graph, hysteresis mechanism, and time penalty logic work together to form a smooth transition zone for the strategy, within which the switching rhythm is regulated. The smooth transition zone is centered on the current path score, extending upwards and downwards to form an anchored budget allowance zone, a hysteresis trigger waiting zone, and a penalty release buffer zone, constituting a dynamic score selection band. Within this scoring band, the strategy score is allowed to fluctuate within a certain range, but before meeting the three conditions of a jump threshold, time penalty release, or hysteresis trend confirmation, no path jump will occur; instead, the current strategy will maintain stable output. The strategy score update cycle is uniformly set to once every 0.2 seconds. Each update reassesses whether the strategy is still within the smooth transition zone based on the latest score status, execution duration, and historical score fluctuation trends. If it is still within the range, the path remains unchanged; if all three switching conditions are met, the smooth transition is complete, and the new strategy is executed. This smooth transition region mechanism is the first to jointly restrict the packing strategy from three dimensions: score fluctuation, execution persistence, and behavioral trend. This prevents the strategy path from jumping due to slight score fluctuations, significantly improving the physical stability and behavioral consistency of strategy selection, and laying the foundation for buffering and regulating subsequent execution actions and reverse incentive mechanisms.

[0077] S005, an execution buffer cone is established on the basis of the smooth transition area of ​​the strategy. The shadow grabbing trajectory and action withdrawal delay threshold are generated by combining the potential barrier energy landscape and the hysteresis mechanism. The linkage path evaluation module recalculates the current packing path score and outputs the corresponding single-step execution action list.

[0078] To ensure that the packing strategy path has sufficient execution latency and physical buffering capacity during score fluctuations and score switching, an execution buffer cone is established on the basis of the smooth transition region of the strategy. Combined with the previously constructed potential barrier energy landscape and hysteresis mechanism, a shadow grabbing trajectory with verification and hysteresis response characteristics is generated, and specific action withdrawal latency thresholds are set. Finally, the linkage path score re-evaluation process outputs a complete list of single-step execution actions, which includes the following steps:

[0079] Based on the current scoring status of the packing strategy path and historical path jump behavior, an execution buffer cone structure with directionality, temporal tolerance, and path scoring tolerance is constructed in three-dimensional space. The starting point of this execution buffer cone is the coordinate of the grab point defined in the current strategy path, extending towards the current target placement point, forming a spatial cone. The axis of this cone is the execution direction from grab to placement. The cone angle is determined by two parameters: the instantaneous gradient value of the scoring change and the stability score of the previous stage path within the same scoring range. For example, when scoring changes slowly and historical path jumps are infrequent, the cone angle is set to within 15 degrees; if the scoring fluctuates rapidly and recent path switching is frequent, the cone angle is expanded to 25 to 30 degrees to increase the buffer zone. The cone length is set to the maximum movement distance corresponding to the expected duration of the action in the current execution path, typically between 0.8 meters and 1.2 meters. The entire buffer cone serves as a behavioral tolerance space under the influence of path scoring disturbances, capable of absorbing small-amplitude path jump requests generated within this scoring range, ensuring the continuity of strategy behavior. This structure differs significantly from the rigid instruction logic in traditional execution paths. It provides a buffer zone for the transition between strategy and physical actions, increasing the resilience of binning decisions to score fluctuations.

[0080] Inside the buffer cone, a shadow grasping trajectory is generated based on the current strategy scoring trend and the hysteresis interval judgment state. This trajectory is used for pre-execution simulation before the actual action is issued. The shadow trajectory lies entirely within the buffer cone, with its starting point coinciding with the current grasping point and its ending point at the target placement position of the alternative path. The trajectory path is adjusted according to the scoring trend direction. The trajectory planning process considers the following three dimensions: first, the continuity of movement speed, ensuring the trajectory does not decelerate or stall during its journey; second, spatial obstacle avoidance judgment, using collision detection between the pre-simulated path and the boundaries of the goods or containers to ensure path accessibility; and finally, attitude consistency, ensuring that the change between the grasping angle and the placement angle does not exceed the set maximum attitude difference threshold, for example, the change in the gripper rotation angle does not exceed 30 degrees. The generated shadow trajectory is not executed immediately but waits based on whether the scoring continues to enter the strategy transition zone and whether the hysteresis judgment meets the switching conditions. If the above conditions are continuously met, the shadow trajectory becomes the actual action trajectory; if the scoring trend declines or the path judgment fails, the trajectory is immediately discarded, and the current strategy continues. This "prediction-verification-transformation" path buffering mechanism effectively avoids the risk of execution interruption due to incomplete path preparation, and improves the physical feasibility of strategy transitions.

[0081] After the shadow trajectory is constructed and enters the continuous scoring observation state, a corresponding action withdrawal delay threshold is set. This threshold adds a delay protection time to the current strategy action when the scoring trend has not been confirmed and path verification has not been completed. The length of this withdrawal delay threshold is determined by the scoring fluctuation period, the strategy scoring difference, and the execution time of the current strategy. For example, if the scoring trend lasts for less than 1 second, the strategy scoring change is less than 0.02, and the current path has been executed for more than 1.5 seconds but has not reached the minimum strategy dwell threshold of 2.5 seconds, then the withdrawal delay time is set to 600 milliseconds. This means that even if the candidate path score increases, no withdrawal action will be initiated within 600 milliseconds. This delay timer is refreshed periodically, checking every 200 milliseconds whether the withdrawal conditions are met. If the score drops before the delay time ends, or the candidate path fails during shadow trajectory verification, the withdrawal action is automatically canceled, and the current path continues to execute. The introduction of this delay mechanism is a "physical patience" design for the packing action layer. It transforms the strategy judgment from a "select and withdraw" jump mode to a stable process of "evaluation-confirmation-execution", which enhances the resistance of the executed action to scoring disturbances, significantly reduces unnecessary withdrawal operations, and reduces the energy consumption and wear problems caused by frequent start-stop of the robotic arm.

[0082] After the buffer cone completes path filtering, the shadow trajectory completes stability verification, and the withdrawal delay threshold completes time evaluation, the current strategy path score is re-evaluated. The latest score, path stability parameters, and action execution conditions are summarized to output a complete list of single-step execution actions. The list includes the following: (1) Strategy path number and current score value; (2) Grab object identifier, including cargo number, size, and shape classification; (3) Three-dimensional coordinates of the actual grab point, including X, Y, Z values ​​and clamping angle; (4) Target placement point location coordinates and placement direction; (5) The position of the path score in the barrier energy map and whether it has entered the high-barrier area; (6) Whether the current path is the original path or a new candidate path, and whether it has passed the shadow trajectory verification; (7) Withdrawal delay timer status and remaining time; (8) Strategy execution time, remaining time, and whether the minimum dwell threshold has been exceeded; (9) Expected energy consumption and path movement distance for this action. This checklist enables implementing agencies to receive structured, clear, and fully evaluated action instructions in one go, avoiding execution failures or instruction conflicts caused by missing scoring information, unprocessed policy changes, or insufficient path verification.

[0083] S006, based on the list of execution actions output by the execution buffer cone, performs real-time dynamic control, injects inverse phase reward pulses into the time and frequency dimensions of the packing strategy, activates the damping response migration mechanism of the execution action, so as to extinguish the oscillation behavior during the packing path switching process and maintain the stability closed loop of the packing process.

[0084] To achieve a dynamically stable response of the bin-packing action list based on the output of the execution buffer cone under high-frequency disturbance conditions, inverse-phase reward pulses are actively injected into the time and frequency dimensions of the policy scoring response. Combined with the damped migration response mechanism of the action layer, rhythmic guidance and behavioral mitigation of policy jumps are achieved. Finally, an adaptive oscillation extinguishing closed-loop control structure for the bin-packing execution process is constructed, which includes the following steps:

[0085] Based on the output list of execution actions, the system monitors the fluctuation characteristics of the strategy score in real time within the score response update cycle, extracts the time rhythm parameters and frequency characteristic indicators of score changes, and identifies the critical conditions for high-frequency oscillations. The score monitoring cycle is set to 100 milliseconds, and the number of score updates, the maximum score fluctuation amplitude, and the frequency of score direction reversals are counted within the past 500 millisecond time window. For example, if the score is updated more than 6 times within 500 milliseconds, the maximum fluctuation exceeds 0.015, and the number of direction reversals exceeds 3 times, it is marked as a "high-frequency fluctuation interval in the score." Simultaneously, the duration of the current path in the three historical strategy switches is statistically analyzed. If all three durations are less than 2 seconds and are accompanied by score jumps and action withdrawals, it can be determined that the strategy execution has entered an unstable segment. Under this condition, path selection is no longer directly based on the strategy with the highest score, but rather a reverse phase oscillation suppression process is initiated, providing a rhythm-aware basis for subsequent strategy rhythm adjustments. Compared with the traditional score-first decision-making method, this step introduces score rhythm judgment logic for the first time, extending the scoring behavior from static numerical sorting to dynamic rhythm recognition, and constructing a time-driven anchor point for subsequent pulse compensation.

[0086] During the scoring period identified as the critical oscillation range, a reverse-phase reward pulse is actively injected into the strategy scoring change stream. This reverse feedback of the scoring direction constructs a suppression signal band before the strategy jump path occurs, forming a targeted rhythm buffer. The injection amplitude of the reverse-phase pulse is set based on the difference between the current scoring rate and the candidate path scoring. If the scoring rate is higher than 0.02 / second and the difference between the candidate path and the current path scoring is less than 0.01, the reverse-phase pulse amplitude is set to 0.007, the injection period is once every 200 milliseconds, and the injection direction is opposite to the original scoring change direction. In the strategy scoring stream, whenever the current scoring trend approaches the anchored budget boundary, a pulse injection action is initiated 300 milliseconds in advance, forming a scoring suppression window to prevent the strategy scoring from crossing the barrier boundary and entering a jump state. For example, if the current strategy scoring is 0.832, the anchored upper bound is 0.850, and the candidate strategy scoring is 0.848, the reverse-phase reward pulse will automatically intervene when the scoring exceeds 0.845, slightly reducing the candidate strategy scoring to prevent it from immediately forming a jump advantage. In the scoring curve, this pulse manifests as a rhythmic interference signal, causing a "flat-top" segment in the score rise, thereby delaying the strategy switch. This rhythmic intervention mechanism is not reported in existing literature. Traditional path strategies often rely on hard score switching, lacking a strategy stabilization period and prone to oscillations. This invention significantly extends the strategy dwell time through a reverse-phase reward injection mechanism, achieving rhythmic traction and improved score stability.

[0087] Building upon the rhythm regulation of the scoring response by the inverse-phase reward pulse, the damping response migration mechanism of the action layer is further activated. During the execution path generation phase, speed easing, acceleration gradualization, and attitude transition curves are introduced to achieve physical mitigation of the packing behavior. In actual deployment, once the current strategy enters the scoring transition waiting state, the execution action does not immediately start at full speed. Instead, a "gradual start window" is set, typically 300 to 500 milliseconds. The initial speed is reduced to 60% of the normal start speed, and the acceleration is controlled within 0.4 m / s². An initial action suppression mechanism is also introduced to maintain a stable direction and slow growth speed in the first 20% of the path. For example, a path with a planned start speed of 300 mm / s is adjusted to 180 mm / s when the damping response mechanism is activated, and a linear gradual change curve is forcibly executed within the first 20 cm of the path to prevent drastic position changes at the strategy switching critical point. Simultaneously, during the start-up of the robotic arm gripper, the attitude angle adjustment range is limited to within 10 degrees to maintain continuous attitude output and avoid drastic angle adjustments. The core of the damping response mechanism lies in physically buffering the execution jump risk caused by the instability of the strategy layer score, absorbing the impact of strategy oscillations on actions and transforming it into a mitigating behavior, thus significantly reducing the impact load on the execution end caused by changes in packing path selection. Unlike the immediate response command triggering method in existing technologies, this invention constructs an execution response control mechanism driven by strategy score and fed back by action inertia, realizing time coupling between the strategy layer and the execution layer, and exhibiting a significant stability improvement effect.

[0088] Based on a composite control system consisting of a scoring inverse-phase pulse and an action damping response, a dynamic closed-loop control structure is constructed, covering scoring monitoring, strategy evaluation, path verification, action execution, and feedback correction. This structure continuously tracks the scoring flow and execution path behavior, updating the action list in real time. The closed-loop structure uses the update time of the action list as its core beat point, synchronizing scoring changes, strategy states, and action execution states. The list generation rhythm is dynamically adjusted according to the fluctuation trend of the scoring curve: when the scoring fluctuation amplitude is less than 0.008 for three consecutive periods and no path switching occurs, the list refresh cycle is shortened to every 500 milliseconds; when scoring fluctuations intensify and the number of path switching increases, the refresh cycle is extended to 800 milliseconds to increase the buffer time for strategy response and execution verification. Furthermore, during list refresh, the current scoring trend, remaining strategy execution time, remaining length of the action buffer interval, whether it is in the shadow trajectory stage, and whether a delayed withdrawal mechanism has been triggered are dynamically collected. After comprehensive evaluation, the current action list is recalculated to determine whether it has entered the "controlled stable zone." If the evaluation result is yes, the current path is maintained first; if the evaluation result is no, the list is regenerated, and new actions are output based on the current score ranking, stability indicators, and damping response status. Ultimately, the strategy scoring behavior, action generation path, and actual execution behavior form a stable feedback loop, realizing a closed-loop control chain of score-driven, rhythm adjustment, action mitigation, and execution confirmation. This enables the entire packing process to be self-aware, self-correcting, and self-suppressive of score disturbances. Compared to traditional score-dominated strategy structures, this closed-loop mechanism integrates score regulation, rhythm response, and physical feedback, possessing high stability, adaptability, and continuous control capabilities. It is a key component of this invention's highly robust intelligent packing strategy.

[0089] This invention constructs a time-series observation system covering the entire process from decision-making to execution, accurately collecting scoring gradients, action withdrawal behaviors, and energy consumption changes to achieve quantitative analysis and causal localization of unstable policy behavior. Combining causal coherence decomposition and counterfactual perturbation replay chains, it achieves for the first time a graphical modeling and visual explanation of the causes of packing strategy oscillations. Furthermore, by introducing an inertial anchoring budget and a minimum policy dwell threshold, it constructs a potential barrier energy landscape with behavioral viscosity, supplemented by a hysteresis mechanism and a time consistency penalty term, fundamentally suppressing high-frequency policy jump tendencies. At the execution layer, it constructs an execution buffer cone and deploys shadow-grabbing trajectories and withdrawal delay thresholds, enabling physical actions to absorb and mitigate scoring perturbations. Finally, by injecting inverse-phase reward pulses into the time and frequency domains of the policy response, it effectively activates the damping migration mechanism of execution behavior, achieving proactive suppression and rapid extinguishing of policy oscillations, thus forming a closed-loop stable intelligent packing control process. The overall solution achieves multi-layered collaborative vibration suppression from strategy learning and path evaluation to action execution, effectively improving the continuity, stability and mechanical structure life of the packing process, and significantly reducing execution energy consumption and task failure rate.

[0090] The foregoing has only described certain exemplary embodiments of the present invention by way of illustration. Undoubtedly, those skilled in the art can modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the foregoing drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.

Claims

1. An automated bin packing method based on reinforcement learning and dynamic search, characterized in that, Includes the following steps: S001, establish a time-series observation layer from decision-making to execution, collect strategy scoring gradient, action withdrawal count and execution energy consumption, and generate baseline data representing packing order oscillation; S002, based on baseline data, perform causal coherence decomposition, identify near-optimal path switching clusters, calculate switching frequency, dwell time and withdrawal cost, and construct an oscillation cause map; S003 introduces a counterfactual replay chain through the oscillation causation map, injects a disturbance sequence into the key time window, compares the path evolution, identifies the steady-state interval, and calculates the inertial anchoring budget and minimum dwell threshold. S004, based on the anchoring budget and dwell threshold, constructs a barrier energy landscape during strategy generation, superimposes a hysteresis mechanism and time consistency penalty, restricts high-frequency path switching, and forms a smooth transition area; S005, establish an execution buffer cone in the smooth transition area, generate the shadow capture trajectory and withdrawal delay threshold, and the linkage path evaluation module outputs a list of single-step execution actions. S006, based on the list of execution actions, performs dynamic control, injects inverse phase reward pulses into the time-frequency domain of the strategy, activates the damped response migration mechanism, and extinguishes the packing path oscillation.

2. The automated bin packing method based on reinforcement learning and dynamic search according to claim 1, characterized in that, Step S001 includes: Based on the establishment of a time-series observation layer between the decision-making process and the execution process, data on container strategy scoring gradient, number of action withdrawals, and energy consumption changes during the execution phase are collected. Based on the scoring gradient, the scoring value, scoring timestamp, target cargo number, target grab position coordinates and placement position coordinates of each packing strategy are recorded, and a one-to-one mapping relationship between the scoring sequence and the action instructions is established. Based on the number of action withdrawals, record the action withdrawal behavior caused by strategy changes or execution failures in each packing task, and collect the withdrawal reason, withdrawal path length, withdrawal duration, and coordinates of the withdrawal grab and placement positions. Based on the energy consumption change data, information on changes in current, voltage, speed and drive load during task execution is collected, and the scoring sequence and withdrawal behavior are synchronized by timestamp to construct a behavioral baseline data structure for identifying packing sequence oscillations.

3. The automated bin packing method based on reinforcement learning and dynamic search according to claim 1, characterized in that, Step S002 includes: After collecting data on scoring gradient, number of action withdrawals, and energy consumption changes, and generating baseline data to characterize packing order oscillations, stable scoring phases are divided based on the continuity and fluctuation amplitude of the scoring values. Within each scoring phase, strategy records with similar scores but different packing behaviors are extracted to construct a strategy path switching cluster. Perform time-series behavior tracking for each strategy path switching cluster, extract the duration, switching frequency and post-switching withdrawal behavior corresponding to each switching event, and calculate the average strategy dwell time and withdrawal cost for each strategy path switching cluster. Collect energy consumption data corresponding to each strategy path switching cluster, analyze the changes in current, power, temperature rise and end vibration before and after path switching, and identify high energy consumption response path switching clusters. By integrating scoring change behavior, strategy path switching characteristics, and energy consumption response data, an oscillation cause map is constructed to reflect the causes of packing order oscillations. The map covers the complete time series and path number evolution process.

4. The automated bin packing method based on reinforcement learning and dynamic search according to claim 3, characterized in that, When constructing the oscillation cause map, for each strategy path switching event, the start and end time of the switching, the corresponding path number, the switching frequency, and the energy consumption change data within five seconds after the switching are marked. Different colors are used to mark power anomalies, end-point vibration enhancement, and action withdrawal, so as to achieve visual marking of high oscillation risk areas.

5. The automated bin packing method based on reinforcement learning and dynamic search according to claim 1, characterized in that, Step S003 includes: After constructing an oscillation causal map to reflect the reasons for packing sequence oscillations, key time windows are selected based on the gradient of score changes, the frequency of path jumps, and the amplitude of energy consumption fluctuations. Within the critical time window, scoring disturbances, packing order disturbances, and target cargo disturbances are inserted respectively. The changes in strategy path, action withdrawal, and energy consumption response after the disturbances are tracked, and the strategy behavior intervals that maintain or restore stability are extracted. By comparing the policy evolution sequences before and after the perturbation, a mapping relationship between the perturbation and policy stability is constructed, and the boundary conditions affecting policy stability are identified. The inertial anchoring budget and minimum policy dwell threshold are calculated based on the stable behavior interval and used as control parameter inputs for the subsequent construction of path switching suppression mechanisms.

6. The automated bin packing method based on reinforcement learning and dynamic search according to claim 1, characterized in that, Step S004 includes: After calculating the inertial anchoring budget and minimum policy dwell threshold of the policy path, a path barrier energy map is constructed, and the path score barrier strength is determined based on the offset between the current score and the anchoring budget boundary. A hysteresis mechanism is introduced in the path scoring process. By analyzing the cumulative magnitude of the scoring trend and the continuity of the sliding scoring window, the scoring jump radius and hysteresis period are set to suppress path switching caused by short-term scoring fluctuations. Based on the difference between the actual execution time of the strategy and the minimum strategy dwell threshold, a time consistency penalty term is introduced to apply a penalty value to the replacement path whose score has not reached the execution time, so as to prevent the strategy jump from happening prematurely. By combining the path barrier energy map, hysteresis mechanism, and time consistency penalty term, a smooth transition region for policy scoring is constructed, which limits the frequency of policy path switching and improves the stability of packing path.

7. The automated bin packing method based on reinforcement learning and dynamic search according to claim 1, characterized in that, Step S005 includes: When the strategy score is in the smooth transition region, an execution buffer cone structure is constructed, with the current grab point as the starting point and the target placement point as the direction. The cone angle and length are set according to the score gradient and historical jumps. A shadow grab trajectory is generated within the execution buffer cone. The trajectory meets the requirements of movement continuity, spatial obstacle avoidance and posture consistency. Based on the scoring trend and hysteresis judgment, it is dynamically determined whether to convert it into an actual action. Set an action withdrawal delay threshold and determine the withdrawal waiting time based on the score fluctuation cycle, score difference and execution time. Suppress immediate jump behavior when the withdrawal conditions are not met. After completing path verification and latency assessment, the scoring, path stability, and execution status are re-integrated to output a structured list of single-step execution actions to guide the physical execution of the current strategy.

8. The automated bin packing method based on reinforcement learning and dynamic search according to claim 7, characterized in that, The shadow capture trajectory is only allowed to be converted into an actual action when the hysteresis trigger condition is met in the continuous score update and the path stability index reaches the set threshold. Furthermore, the conversion into an actual action can only be started after the withdrawal delay timer has expired, so as to ensure that the action stability verification is sufficient before the strategy switch.

9. The automated bin packing method based on reinforcement learning and dynamic search according to claim 1, characterized in that, Step S006 includes: Based on the list of execution actions output by the execution buffer cone, the frequency and direction reversal features of score fluctuations are extracted within the score update cycle to identify high-frequency oscillation intervals; Injecting a reverse-phase reward pulse, opposite to the scoring direction, into the scoring change stream within the identified oscillation range creates a scoring suppression window, delaying strategy switching behavior. When the scoring trend approaches the switching state, the damping response mechanism of the action execution phase is activated to control the initial velocity, acceleration and attitude changes of the path, thereby achieving action mitigation. The execution action list is dynamically updated based on the scoring trend, strategy execution status and damping feedback results. A closed-loop control structure of strategy scoring-driven, rhythm regulation and action feedback is constructed to achieve adaptive extinguishing of scoring disturbances and stable execution.

10. The automated bin packing method based on reinforcement learning and dynamic search according to claim 9, characterized in that, The injection cycle and amplitude of the reverse phase reward pulse are dynamically adjusted according to the rate of change of the score and the score difference of the candidate path. Furthermore, a pre-injection point is preset before the score trend approaches the anchor budget boundary to form a flat-top zone of score rise and delay the triggering of the strategy jump behavior.

Citation Information

Patent Citations

  • Boxing method based on deep reinforcement learning

    CN111695700A

  • Deep reinforcement learning network system

    CN112884126A