Intelligent warehouse management system and method based on Internet of Things

Through the IoT intelligent warehousing management system, the intelligent body strategy is optimized using multi-source sensor synchronous sampling and conflict detection, solving the problem of cold air loss in cold chain warehousing, and achieving efficient refrigeration management and energy consumption optimization.

CN120338670AActive Publication Date: 2025-07-18SHANGHAI ZHONGTONG YUNCHANG TECH CO LTD

Patent Information

Application Number
CN202510797113.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-07-18
Estimated Expiration
2045-06-16

AI Technical Summary

Technical Problem

In cold chain warehousing, independent regulation of multiple control equipment leads to the loss of cold air, increasing the difficulty of warehousing management, especially when the cold storage door is frequently opened, refrigeration efficiency and energy consumption are difficult to optimize.

Method used

The intelligent warehousing management system based on the Internet of Things is adopted, and the synchronization of sampling, conflict detection and coordinated decision-making through multi-source sensors, combined with reinforcement learning, optimized intelligent body strategies, realize global coordinated control, and reduce energy consumption and cargo losses.

Benefits of technology

It significantly improves the coordination efficiency of each subsystem in cold chain warehousing, reduces energy consumption and cargo losses, and ensures the continuous and stable operation of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338670A_ABST
    Figure CN120338670A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent warehousing, and discloses an intelligent warehousing management system and method based on the Internet of Things, and the method comprises the following steps: monitoring a cold chain warehousing environment and the state of each agent in real time, and collecting related data; wherein each agent is responsible for controlling one independent physical device, and the types of the independent physical devices comprise a generator, a motor and a sensor; each agent evaluates own state and shares state information so as to enhance a global view angle; and identifying potential strategy conflicts among all the paired agents by adopting a conflict detection algorithm. According to the method, synchronous sampling is carried out through a multi-source sensor, a global collaborative view angle is enhanced based on conflict function modeling and dynamic influence weight distribution, and a multi-agent strategy conflict is solved by combining conflict threshold updating of Q function iteration and action optimization driven by coordination factors; the collaborative efficiency of all subsystems in cold chain storage can be remarkably improved, energy consumption and cargo loss are reduced, and continuous and stable operation of the system is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent warehousing, and more specifically, to an intelligent warehousing management system and method based on the Internet of Things. Background Art

[0002] Cold chain warehousing is a specialized warehousing system that maintains a low-temperature environment through refrigeration technology and equipment. It is mainly used for storing temperature-sensitive commodities such as food and medicine, and is the core link of cold chain logistics.

[0003] In cold chain warehousing, since the regulation of multiple control devices in the cold storage internal environment is controlled independently, when optimizing the refrigeration unit and independently scheduling and managing the AGV, the frequent opening of the cold storage door will cause cold air loss, increasing the difficulty of warehousing management adjustment. Summary of the Invention

[0004] The present invention provides an intelligent warehousing management system and method based on the Internet of Things to solve the technical problems in the related art.

[0005] The present invention provides an intelligent warehousing management method based on the Internet of Things, including the following steps:

[0006] S100, Data collection: Real-time monitor the cold chain warehousing environment and the status of each agent, and collect relevant data;

[0007] Each agent is responsible for controlling an independent physical device, and the types of independent physical devices include generators, motors, and sensors;

[0008] S200, Status evaluation and information sharing: Each agent evaluates its own status and shares the status information to enhance the global perspective;

[0009] S300, Conflict detection: Use a conflict detection algorithm to identify potential policy conflicts between all pairs of agents;

[0010] S400, Coordination decision-making: Through a coordination mechanism, adjust the action selection of policy conflicts to achieve global optimization;

[0011] S500, Control execution: Apply the coordinated policies of each agent to the actual control system and adjust the device operation parameters;

[0012] S600, Feedback learning and adjustment: According to the adjusted device operation parameters and environmental feedback, update the policy parameters of each agent, and use the reward mechanism in reinforcement learning for self-optimization;

[0013] S700, Policy deployment and closed-loop management: Apply the updated policy parameters to the actual cold chain warehousing environment, monitor the regulation effect, and form a continuous optimization closed-loop.

[0014] Furthermore, S100 specifically includes the following steps:

[0015] S110, multi-source sensor synchronous sampling: Collect the spatial temperature field through a distributed optical fiber temperature measurement network, obtain the thermal inertia parameters of goods using an RFID shelf reader, and collect the cold storage door status and compressor power;

[0016] S120, agent state synchronization: Subscribe to the state vectors published by each agent through ROS nodes, record the path planning parameters of the AGV scheduling agent, and collect the control parameters of the refrigeration unit agent;

[0017] S130, heterogeneous data preprocessing: Implement sliding window calibration for temperature data, perform wavelet denoising on power data, and conduct spatio-temporal alignment processing;

[0018] S140, metadata encapsulation and transmission: Encapsulate the processed data into a tuple in a unified format and transmit it to the edge computing node through the OPC UA protocol;

[0019] S150, output of the data collection stage: Finally generate a time series data set through a tuple in a unified format.

[0020] Furthermore, S200 specifically includes the following steps:

[0021] S210, define the conflict function: Each agent calculates the degree of conflict between the local goal and the global goal based on its own state and external input;

[0022] The calculation formula of the conflict function is as follows:

[0023] ;

[0024] Where represents the conflict value of agent at time , represents the current state vector of agent , represents the global goal state vector, represents the neighbor set of agent , , represents the weight coefficient in the conflict function, > 0, > 0, ;

[0025] S220, calculate the influence weight: Dynamically allocate the influence weight of the agent according to the conflict value and historical performance;

[0026] S230, State Information Sharing: Update the local state based on the influence weight and broadcast it to adjacent agents;

[0027] S240, Global Consistency Check: Check whether the deviation between the states of all agents and the global goal converges.

[0028] Furthermore, S300 specifically includes the following steps:

[0029] S310, Parameter Predefinition: Initialize the agent action set, define the conflict determination threshold, policy network parameters, learning rate, and discount factor;

[0030] S320, Conflict Function Modeling: Define a binary conflict function to quantify the policy contradiction between two actions;

[0031] The calculation formula of the binary conflict function is as follows:

[0032] ;

[0033] Where represents the binary conflict function, represents the action 's influence weight on the global, is the state-action value function;

[0034] S330, Agent Action Pair Traversal: Traverse all agent action pairs to avoid duplicate calculations;

[0035] S340, Dynamic Conflict Threshold Update: Iteratively update the conflict determination threshold based on the Q function, synchronized with policy optimization;

[0036] S350, Total Conflict Count Calculation: Statistically calculate the conflict results of all agent pairs and output the total conflict count.

[0037] Furthermore, the agent action pairs in S330 are ;

[0038] Where , represents the total number of agents.

[0039] Furthermore, S400 specifically includes the following steps:

[0040] S410, Initialize the Coordination Factor: Set the coordination factor to balance the weight of its own Q value and external influence;

[0041] S420, Calculate the Action Influence Function of Other Agents: For each non-current agent, quantify the influence value of its action on the agent;

[0042] S430, Comprehensive Q-value and External Influence: Sum the self Q-value and the weighted influence of other agents to generate an adjusted evaluation value;

[0043] S440, Select the Optimal Action Based on the Adjusted Value: Determine the new action by maximizing the adjusted evaluation value;

[0044] S450, Conflict-Aware Q-value Correction: If the new action causes a conflict, attenuate the Q-value update amplitude proportionally.

[0045] Furthermore, S500 specifically includes the following steps:

[0046] S510, Policy Mapping and System Integration: Map the coordinated policy to the physical parameters of the actual control system;

[0047] S520, Real-Time Parameter Adaptive Adjustment: Dynamically correct the device parameters based on the feedback signal to make the control objective consistent with the policy;

[0048] S530, Conflict Resolution and Stability Verification: Detect multi-device cooperation conflicts and verify the effectiveness of the parameters through stability criteria.

[0049] Furthermore, S600 specifically includes the following steps:

[0050] S610, Calculate the Target Q-value: Construct the target value based on the current state, action, immediate reward, and the maximum expected Q-value of the next state;

[0051] S620, Calculate the Temporal Difference Error: Quantify the policy optimization direction through the difference between the target value and the current Q-value;

[0052] S630, Update the Q-function Parameters: Adjust the parameters of the Q-function according to the error , and gradually approach the optimal policy.

[0053] Furthermore, S700 specifically includes the following steps:

[0054] S710, Policy Issuance and Device Synchronization: Issue the updated policy parameters to each device controller to ensure multi-device policy synchronization;

[0055] S720, Real-Time Regulation Execution and State Monitoring: The device adjusts the operating parameters based on the new policy and collects the state data in real time;

[0056] S730, Multi-Device Cooperation Effect Verification: Detect device cooperation conflicts and verify the system stability;

[0057] S740, Performance Feedback and Dynamic Readjustment: Trigger adaptive adjustment and reinforcement learning update according to the real-time performance;

[0058] S750, Closed-loop learning cycle: Store the device operation data in the experience pool, periodically trigger the Q-function update of S600, and form a closed loop of "policy optimization - execution - feedback".

[0059] The present invention also provides an intelligent warehousing management system based on the Internet of Things, which executes the aforementioned intelligent warehousing management method based on the Internet of Things, and includes:

[0060] Data acquisition and synchronization module: Execute the tasks in the S100 stage, including multi-source sensor synchronization, agent state synchronization, heterogeneous data processing, metadata encapsulation, and time-series dataset generation;

[0061] Multi-agent collaborative decision-making module: Integrate the S200 - S400 stages to achieve state evaluation, conflict detection, and coordinated decision-making;

[0062] Device control and execution module: Corresponding to the S500 stage, responsible for policy mapping, parameter adaptive adjustment, and conflict resolution verification;

[0063] Reinforcement learning feedback optimization module: Cover the S600 - S700 stages to achieve policy parameter update and closed-loop deployment;

[0064] Global monitoring and verification module: Run through the processes in S100 - S700, and execute multi-device collaborative verification, performance feedback, and stability monitoring.

[0065] The beneficial effects of the present invention are as follows:

[0066] By synchronously sampling multi-source sensors, enhancing the global collaborative perspective based on conflict function modeling and dynamic influence weight allocation, combining the conflict threshold update of Q-function iteration and action optimization driven by coordination factors to solve the multi-agent policy conflict, the present invention can significantly improve the collaborative efficiency of each subsystem in cold chain warehousing, reduce energy consumption and cargo loss, and ensure the continuous and stable operation of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Figure 1 is a flowchart of an intelligent warehousing management method based on the Internet of Things proposed by the present invention;

[0068] Figure 2 is a structural block diagram of an intelligent warehousing management system based on the Internet of Things proposed by the present invention.

[0069] In the figure: 101, data acquisition and synchronization module; 102, multi-agent collaborative decision-making module; 103, device control and execution module; 104, reinforcement learning feedback optimization module; 105, global monitoring and verification module. DETAILED DESCRIPTION OF THE INVENTION

[0070] Reference will now be made to example embodiments to discuss the subject matter described herein. It should be understood that discussing these embodiments is only to enable those skilled in the art to better understand and thus implement the subject matter described herein, and that changes can be made to the functions and arrangements of the elements discussed without departing from the scope of protection of the content of this specification. Each example can omit, substitute, or add various processes or components as needed. Additionally, features described relative to some examples can also be combined in other examples.

[0071] As Figure 1 shown, an Internet of Things-based intelligent warehousing management method includes the following steps:

[0072] S100, Data collection: Monitor the cold chain warehousing environment and the status of each intelligent agent in real time, and collect relevant data;

[0073] In one embodiment of the present invention, the specific steps are as follows:

[0074] S110, Multi-source sensor synchronous sampling: Collect the spatial temperature field through a distributed optical fiber temperature measurement network ( , is the sensor number); Obtain the thermal inertia parameters of the goods using an RFID shelf reader ( , is the shelf number); Collect the status of the cold storage door (door opening duration) and the compressor power ;

[0075] Its calculation formula is as follows:

[0076] ;

[0077] ;

[0078] ;

[0079] Wherein represents the temperature error, represents the true temperature, , wherein represents the temperature measurement noise variance, wherein represents the th shelf's goods heat capacity, represents the power error, wherein , represents the Laplace distribution, wherein and are the temperature change amount and the time interval respectively, wherein represents the power measurement noise intensity parameter, wherein Represents the compressor reference voltage, where Represents the set temperature;

[0080] S120, Agent State Synchronization: Subscribe to the state vectors published by each agent through the ROS node ( is the agent number); Record the path planning parameters of the AGV scheduling agent ; Collect the control parameters of the refrigeration unit agent ;

[0081] Where Represents the maximum speed of the AGV device, Represents the maximum acceleration of the AGV device, Represents the AGV safety distance, where Represents the compressor start-stop hysteresis bandwidth;

[0082] State Synchronization Model:

[0083] ;

[0084] Where Represents the observed state of the th agent at the moment, Represents the previous moment action vector, Represents the previous moment reward value, Represents the control parameter of the th agent at the moment, Represents the state of the th agent at the moment;

[0085] S130, Heterogeneous Data Preprocessing:

[0086] Implement sliding window calibration on temperature data:

[0087] ;

[0088] Where Represents the drift coefficient of the th sensor, Represents the calibrated temperature data, Represents the temperature error;

[0089] Perform wavelet denoising on power data:

[0090] ;

[0091] Where Represents the compressor power data after denoising, Indicates the total number of wavelet decomposition levels (usually taken as 5 - 8), Indicates the filtering function, Indicates the number of wavelet decomposition levels;

[0092] Spatio - temporal alignment processing (to solve clock offset at the 5 - ms level):

[0093] ;

[0094] where Indicates the timestamp after synchronization, Indicates the original timestamp, Indicates the network clock synchronization error, Indicates the amount of data for aligned data;

[0095] S140, Metadata encapsulation and transmission:

[0096] Encapsulate the processed data into a tuple in a unified format:

[0097] ;

[0098] Transmit it to the edge computing node through the OPC UA protocol:

[0099] ;

[0100] where Indicates the encoding length of single - temperature data (fixed at 32 bits), Indicates the network bandwidth, Indicates the protocol stack processing delay;

[0101] S150, Output in the data collection stage: Finally generate a time - series data set through a tuple in a unified format:

[0102] ;

[0103] where is the data acquisition period, meeting the real - time requirements of cold chain control, Indicates the initial time of acquisition, Indicates the end time of acquisition.

[0104] S200, State evaluation and information sharing: Each agent evaluates its own state and shares the state information to enhance the global perspective;

[0105] In an embodiment of the present invention, the specific steps are as follows:

[0106] S210, Define the conflict function: Each agent calculates the degree of conflict between the local goal and the global goal based on its own state and external inputs;

[0107] Its calculation formula is as follows:

[0108] ;

[0109] Where represents the conflict value of the agent at time , represents the current state vector of the agent , represents the global target state vector, which is provided by the data collection phase, represents the neighbor set of the agent , , represents the weight coefficient in the conflict function ( > 0, > 0, and needs to satisfy ).

[0110] S220, calculate the influence weight: Dynamically allocate the influence weight of the agent according to the conflict value and historical performance;

[0111] Its calculation formula is as follows:

[0112] ;

[0113] Where represents the influence weight of the agent at time , represents the decay coefficient of the conflict value ( ), represents the historical performance gain coefficient ( ), where represents the historical performance score of the agent , represents the total number of agents, where represents the agent at time of the conflict value.

[0114] S230, state information sharing: Update the local state based on the influence weight and broadcast it to adjacent agents;

[0115] Its calculation formula is as follows:

[0116] ;

[0117] Where represents the updated state vector of the agent , represents the neighbor agent The current state vector, where represents the agent at time influence weight.

[0118] S240, Global Consistency Check: Check whether the deviation between the states of all agents and the global goal converges;

[0119] Its calculation formula is as follows:

[0120] ;

[0121] Termination condition: If (threshold ), it is determined to converge, otherwise trigger parameter adjustment.

[0122] S300, Conflict Detection: Use a conflict detection algorithm to identify potential policy conflicts between all pairs of agents;

[0123] In an embodiment of the present invention, the specific steps are as follows:

[0124] S310, Parameter Predefinition: Initialize the agent action set , define the conflict determination threshold , policy network parameters , learning rate and discount factor ;

[0125] ;

[0126] Where represents the action of the i-th agent, , represents the total number of agents.

[0127] S320, Conflict Function Modeling: Define a binary conflict function to quantify the policy contradiction between two actions.

[0128] Its calculation formula is as follows:

[0129] ;

[0130] Where represents the influence weight of action on the global, is the state-action value function.

[0131] S330, Agent Action Pair Traversal: Traverse all agent action pairs , where , avoiding repeated calculations, .

[0132] S340, Dynamic conflict threshold update: Iteratively update the conflict determination threshold based on the Q function , ensuring synchronization with policy optimization;

[0133] Its calculation formula is as follows:

[0134] ;

[0135] Where is the current state, is the next state, is the candidate action, represents the conflict determination threshold of the next state, represents the current conflict determination threshold.

[0136] S350, Total conflict count calculation: Statistically count the conflict results of all agent pairs and output the total conflict count.

[0137] Its calculation formula is as follows:

[0138] ;

[0139] Where is the integer total conflict count, and its range is .

[0140] S400, Coordination decision-making: Through the coordination mechanism, adjust the conflict action selection to achieve global optimization, and its calculation formula is expressed as:

[0141] ;

[0142] Where represents the adjusted action of the i-th agent, is the coordination factor, is the influence of the actions of other agents on the current agent.

[0143] S410, Initialize the coordination factor: Set the coordination factor ( ) to balance the weight of its own Q value and external influence;

[0144] Where represents the coordination factor (a scalar parameter that needs to be tuned through experiments);

[0145] S420, Calculate the action influence function of other agents: For each non-current agent ( ), quantify the influence value of its action on the agent ;

[0146] Its calculation formula is as follows:

[0147] ;

[0148] Where represents the conflict value of agent j, represents the interaction distance (such as Euclidean distance) between agents i and j, represents a minimum value ( ≠0), represents a conflict indication function (taking 1 when there is a conflict, otherwise taking 0);

[0149] S430, Integrate Q value and external influence: Sum the agent's own Q value and the weighted influence of other agents to generate an adjusted evaluation value;

[0150] Its calculation formula is as follows:

[0151] ;

[0152] Where represents the action value based on state and policy network parameters , represents the global influence term weighted by the coordination factor;

[0153] S440, Select the optimal action based on the adjusted value: Determine the new action by maximizing the adjusted evaluation value;

[0154] Its calculation formula is as follows:

[0155] ;

[0156] Where represents the action space;

[0157] S450, Conflict-aware Q value correction: If the new action causes a conflict, attenuate the Q value update amplitude proportionally;

[0158] Its calculation formula is as follows:

[0159] ;

[0160] Where represents the corrected Q value, represents the learning rate, represents the current maximum conflict value, represents the preset conflict threshold;

[0161] S500, Control execution: Apply the coordinated strategies of each agent to the actual control system to adjust the device operation parameters;

[0162] In one embodiment of the present invention, the specific steps are as follows:

[0163] S510, Policy Mapping and System Integration: Map the coordinated policies (such as optimization weights, priority rules) to the physical parameters of the actual control system.

[0164] Let the policy parameters generated in step 3 be (dynamic weight), (priority coefficient), and the constraint boundary generated in step 4 be , then the target value of device parameter adjustment is:

[0165] ;

[0166] Where represents the dynamic weight, , represents the priority coefficient, , represents the constraint boundary, .

[0167] S520, Real-time Parameter Adaptive Adjustment: Dynamically correct device parameters based on feedback signals to ensure the consistency between the control target and the policy.

[0168] Formula: Let the real-time feedback error be , and the mean value of historical errors be , then the parameter correction amount is:

[0169] ;

[0170] Updated parameter:

[0171] ;

[0172] Where represents the updated parameter, represents the learning rate, represents the decay factor, represents the instantaneous error, represents the mean value of historical errors.

[0173] S530, Conflict Resolution and Stability Verification: Detect multi-device collaboration conflicts and verify the effectiveness of parameters through stability criteria.

[0174] Where the conflict index is as follows:

[0175] ;

[0176] Where the stability criterion is as follows:

[0177] If And , it is determined to be stable;

[0178] Among them represents the conflict weight, , is the reference benchmark value, is the conflict threshold, represents the error change rate threshold.

[0179] S600, Feedback learning and adjustment: According to the feedback of adjusting device operation parameters and environment, update the policy parameters of each agent, and use the reward mechanism in reinforcement learning for self-optimization;

[0180] Its calculation formula is expressed as:

[0181] ;

[0182] Among them is the reward, is the discount factor.

[0183] S610, Calculate the target Q-Value: Based on the current state , action , immediate reward , and the maximum expected value of the next state , construct the target value .

[0184] Its calculation formula is as follows:

[0185] ;

[0186] Among them represents the target value, represents the discount factor, , represents the Q function parameter, represents the optional action for the next state.

[0187] S620, Calculate the Temporal Difference Error (TD Error): Quantify the policy optimization direction through the difference between the target value and the current Q value.

[0188] Its calculation formula is expressed as:

[0189] ;

[0190] Among them is the error.

[0191] S630, Update Q - function parameters (Parameter Update): Adjust the parameters of the Q - function according to the error to gradually approximate the optimal strategy.

[0192] Its calculation formula is expressed as:

[0193] ;

[0194] where represents the learning rate, ;

[0195] S700, Policy Deployment and Closed - loop Management: Apply the updated policy parameters to the cold - chain storage environment, monitor the regulation effect, and form a continuous optimization closed - loop;

[0196] In an embodiment of the present invention, the specific steps are as follows:

[0197] S710, Policy Issuance and Device Synchronization: Send the updated policy parameters (such as the optimized function parameters , dynamic weights ) to each device controller to ensure multi - device policy synchronization;

[0198] Its calculation formula is expressed as:

[0199] ;

[0200] where represents the updated parameter, , is the total number of devices, represents the parameter of the k - th device.

[0201] S720, Real - time Regulation Execution and Status Monitoring: The device adjusts the operating parameters (such as power, speed) based on the new policy and collects status data in real - time (such as , ).

[0202] The calculation formula for action selection is as follows:

[0203] ;

[0204] where the previously mentioned , , , are used.

[0205] S730, Verification of Multi - device Collaboration Effect: Detect device collaboration conflicts and verify the system stability (using the previously mentioned conflict index and stability criterion).

[0206] Its calculation formula is as follows:

[0207] ;

[0208] S740, Performance feedback and dynamic readjustment: According to real-time performance (such as error , reward ), trigger the aforementioned adaptive adjustment or the aforementioned reinforcement learning update.

[0209] If , recalculate ; where represents the error threshold.

[0210] S750, Closed-loop learning cycle: Store the device operation data in the experience pool, and periodically trigger the aforementioned function update to form a "policy optimization - execution - feedback" closed loop.

[0211] It should be added that each agent is responsible for controlling an independent physical device (such as the device numbered k, corresponding to ).

[0212] The types of independent physical devices include, but are not limited to, generators, motors, and sensors.

[0213] Among them, the agent issues policy parameters (such as ) through the device controller to directly drive the device to execute actions , where the state includes the real-time parameters of the device, and the reward is calculated based on the device performance index.

[0214] Among them, multiple agents need to meet system-level constraints to avoid physical quantity conflicts between devices (such as power overload and frequency instability).

[0215] As Figure 2 shown, the present invention also discloses and provides an Internet of Things-based intelligent warehouse management system that executes the above-mentioned Internet of Things-based intelligent warehouse management method, including the following modules:

[0216] Data acquisition and synchronization module 101: Execute the tasks in the S100 stage, including multi-source sensor synchronization, agent state synchronization, heterogeneous data processing, metadata encapsulation, and time-series dataset generation;

[0217] Multi-agent collaborative decision-making module 102: Integrate the S200 - S400 stages to achieve state evaluation, conflict detection, and coordinated decision-making;

[0218] Device Control and Execution Module 103: Corresponding to the S500 stage, responsible for policy mapping, parameter adaptive adjustment, and conflict resolution verification;

[0219] Reinforcement Learning Feedback Optimization Module 104: Covering the S600 - S700 stages, realizing policy parameter update and closed-loop deployment;

[0220] Global Monitoring and Verification Module 105: Running through the processes in S100 - S700, performing multi-device collaborative verification, performance feedback, and stability monitoring.

[0221] The present invention also discloses a storage medium storing non-transitory computer-readable instructions for executing one or more steps in the foregoing intelligent warehouse management method based on the Internet of Things.

[0222] The computer program can be stored / distributed on a suitable medium, such as an optical storage medium or a solid-state medium supplied together with or as part of other hardware, but can also be distributed in other forms, such as via the Internet or other wired or wireless telecommunication systems. Any reference signs in the claims shall not be construed as limiting the scope.

[0223] The above has described the embodiments of this example, but this example is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative and not restrictive. Under the inspiration of this example, those of ordinary skill in the art can also make many forms, all of which fall within the protection scope of this example.

Claims

1. An intelligent warehouse management method based on the Internet of Things, characterized in that, It includes the following steps: S100, Data collection: Monitor the cold chain storage environment and the status of each agent in real time, and collect relevant data; Each agent is responsible for controlling an independent physical device, and the types of independent physical devices include generators, motors, and sensors; S200, Status evaluation and information sharing: Each agent evaluates its own status and shares the status information to enhance the global perspective; S300, Conflict detection: Use a conflict detection algorithm to identify potential policy conflicts between all pairs of agents; S400, Coordination decision-making: Through a coordination mechanism, adjust the action selection of policy conflicts to achieve global optimization; S500, Control execution: Apply the coordinated policies of each agent to the actual control system and adjust the device operation parameters; S600, Feedback learning and adjustment: According to the adjusted device operation parameters and environmental feedback, update the policy parameters of each agent, and use the reward mechanism in reinforcement learning for self-optimization; S700, Policy deployment and closed-loop management: Apply the updated policy parameters to the actual cold chain storage environment, monitor the control effect, and form a continuous optimization closed loop.

2. The intelligent warehousing management method based on the Internet of Things according to claim 1, wherein, Specifically, in S100, it includes the following steps: S110, Multi-source sensor synchronous sampling: Collect the spatial temperature field through a distributed optical fiber temperature measurement network, use an RFID shelf reader to obtain the thermal inertia parameters of goods, and collect the cold storage door status and compressor power; S120, Agent status synchronization: Subscribe to the status vectors published by each agent through ROS nodes, record the path planning parameters of the AGV scheduling agent, and collect the control parameters of the refrigeration unit agent; S130, Heterogeneous data preprocessing: Implement sliding window calibration for temperature data, perform wavelet denoising on power data, and perform spatio-temporal alignment processing; S140, Metadata encapsulation and transmission: Package the processed data into a tuple of a unified format and transmit it to the edge computing node through the OPC UA protocol; S150, Output of the data collection stage: Finally generate a time series data set through a tuple of a unified format.

3. The intelligent warehousing management method based on the Internet of Things according to claim 2, characterized in that, Specifically, in S200, it includes the following steps: S210, Define the conflict function: Each agent calculates the conflict degree between the local goal and the global goal based on its own status and external input; The calculation formula of the conflict function is as follows: ; wherein represents the agent at time conflict value, represents the agent current state vector, represents the global target state vector, represents the agent neighbor set, , represents the weight coefficient in the conflict function, > 0, > 0, ; S220, Calculate the influence weight: Dynamically allocate the influence weight of the agent according to the conflict value and historical performance; S230, Status information sharing: Update the local status based on the influence weight and broadcast it to adjacent agents; S240, Global consistency check: Check whether the deviation between the status of all agents and the global goal converges.

4. The intelligent warehousing management method based on the Internet of Things according to claim 3, characterized in that, Specifically, in S300, it includes the following steps: S310, Parameter predefinition: Initialize the agent action set, define the conflict determination threshold, policy network parameters, learning rate, and discount factor; S320, Conflict function modeling: Define a binary conflict function to quantify the policy contradiction between two actions; The calculation formula of the binary conflict function is as follows: ; Among them represents a binary conflict function represents an action the impact weight on the global is the state - action value function; S330, Traversal of agent action pairs: Traverse all agent action pairs to avoid repeated calculations; S340, Dynamic Conflict Threshold Update: Iteratively update the conflict determination threshold based on the Q function, synchronized with policy optimization; S350, Total Conflict Count Calculation: Statistically analyze the conflict results of all agent pairs and output the total number of conflicts.

5. The intelligent warehouse management method based on the Internet of Things according to claim 4, characterized in that Among them, the agent action pair in S330 is ; Among them , represents the total number of agents.

6. The intelligent warehousing management method based on the Internet of Things according to claim 5, wherein, Specifically, S400 includes the following steps: S410, Initializing the Coordination Factor: Set the coordination factor to balance the weight of its own Q value and external influence; S420, Calculating the Action Influence Function of Other Agents: For each non-current agent, quantify the influence value of its action on the agent; S430, Integrating the Q Value and External Influence: Sum the agent's own Q value and the weighted influence of other agents to generate an adjusted evaluation value; S440, Selecting the Optimal Action Based on the Adjusted Value: Determine the new action by maximizing the adjusted evaluation value; S450, Conflict-Aware Q Value Correction: If the new action causes a conflict, attenuate the Q value update amplitude proportionally.

7. An intelligent warehouse management method based on the Internet of Things according to claim 6, characterized in that, Specifically, S500 includes the following steps: S510, Policy Mapping and System Integration: Map the coordinated policy to the physical parameters of the actual control system; S520, Real-Time Parameter Adaptive Adjustment: Dynamically correct the device parameters based on the feedback signal to make the control target consistent with the policy; S530, Conflict Resolution and Stability Verification: Detect multi-device collaboration conflicts and verify the effectiveness of the parameters through stability criteria.

8. An intelligent warehouse management method based on the Internet of Things according to claim 7, characterized in that Specifically, S600 includes the following steps: S610, Calculating the Target Q Value: Based on the current state, action, immediate reward, and the maximum expected Q value of the next state, construct the target value; S620, Calculating the Temporal Difference Error: Quantify the policy optimization direction through the difference between the target value and the current Q value; S630, Update the Q - function parameters: Adjust the parameters of the Q - function according to the error , and gradually approach the optimal strategy.

9. An intelligent warehousing management method based on the Internet of Things according to claim 8, characterized in that Specifically, S700 includes the following steps: S710, Policy Distribution and Device Synchronization: Distribute the updated policy parameters to each device controller to ensure multi-device policy synchronization; S720, Real-Time Regulation Execution and Status Monitoring: The device adjusts the operating parameters based on the new policy and collects status data in real time; S730, Multi-Device Collaboration Effect Verification: Detect device collaboration conflicts and verify the system stability; S740, Performance Feedback and Dynamic Readjustment: Trigger adaptive adjustment and reinforcement learning update according to the real-time performance; S750, Closed-Loop Learning Cycle: Store the device operation data in the experience pool and periodically trigger the Q function update in S600 to form a "policy optimization - execution - feedback" closed loop.

10. An intelligent warehouse management system based on the Internet of Things, characterized in that, Implementing an Internet of Things-based intelligent warehouse management method as described in any one of claims 1-9, including: Data Acquisition and Synchronization Module: Execute the tasks in the S100 stage, including multi-source sensor synchronization, agent state synchronization, heterogeneous data processing, metadata encapsulation, and temporal dataset generation; Multi-Agent Collaborative Decision-Making Module: Integrate the S200-S400 stages to achieve state evaluation, conflict detection, and coordinated decision-making; Device Control and Execution Module: Corresponding to the S500 stage, responsible for policy mapping, parameter adaptive adjustment, and conflict resolution verification; Reinforcement Learning Feedback Optimization Module: Cover the S600-S700 stages to achieve policy parameter update and closed-loop deployment; Global Monitoring and Verification Module: Throughout the processes in S100 - S700, it performs multi - device collaborative verification, performance feedback, and stability monitoring.

Citation Information

Patent Citations

  • Intelligent agent task allocation method based on deep reinforcement learning

    CN114638339A

  • Multi-AGV driving control method, device and equipment and storage medium

    CN117519215A

  • AGV forklift cooperative scheduling method and system for intelligent warehouse management

    CN118863472A

  • SPS trolley material distribution method and system based on digital twinning

    CN119090378A

  • Community digital grid autonomous management method based on multi-agent system

    CN119168229A

Cited By

  • Intelligent gas storage scheduling method and system based on 5G network

    CN120525379A

  • Networked loom intelligent control method and system based on process knowledge software

    CN120972828A

  • Warehousing system, robot, obstacle avoidance method of robot and computer readable storage medium

    CN121209506A

  • Factory management and control system based on artificial intelligence and big data

    CN121581515A