Method, System and Storage Medium for Intelligent Control of On-vehicle Cold Chain Temperature and Humidity
By arranging multiple sensors in the carriage to collect data, using a three-way decision model and a temperature and humidity control system built with a multi-agent deep reinforcement learning, the problems of hysteresis and high energy consumption in traditional methods are solved, and accurate and stable collaborative optimization and adaptive control are achieved.
Patent Information
- Application Number
- CN202510458238.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-04-14
AI Technical Summary
The existing vehicle-mounted cold chain temperature and humidity control methods lack prospectiveness and cannot coordinately optimize temperature and humidity control, resulting in lag control, high energy consumption and lack of adaptive learning capabilities, making it difficult to cope with complex and changeable transportation environments and abnormal situations.
Data is collected through multiple temperature and humidity sensors, a three-way decision model and long-term short-term memory network are used to build a temperature and humidity trend analysis model, and a control decision system is built with multi-agent deep reinforcement learning to realize collaborative optimization control of temperature and humidity, and adaptive updates are carried out.
It improves the accuracy, stability and energy utilization efficiency of temperature and humidity control in vehicle-mounted cold chains, can predict and deal with temperature and humidity abnormalities in advance, adapt to different working conditions, reduce energy consumption and improve response speed.
Smart Images

Figure CN119987469B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of vehicle temperature control, and particularly to a method, system, and storage medium for intelligently controlling the temperature and humidity of vehicle cold chains. Background Art
[0002] Currently, cold chain logistics is increasingly widely used in fields such as food, medicine, and fresh produce, and vehicle cold chain temperature and humidity control is a key link in cold chain logistics. Traditional vehicle cold chain temperature and humidity control methods mainly rely on simple threshold control strategies, that is, when the temperature or humidity exceeds the preset threshold, the corresponding refrigeration or dehumidification equipment is started, and when the temperature and humidity return to the preset range, the equipment is turned off. This control method is simple to implement and can meet basic requirements in a stable environment. With the development of technology, some improved control methods such as proportional-integral-derivative (PID) control and fuzzy control have also been applied to vehicle cold chain temperature and humidity control, improving the accuracy and stability of temperature and humidity control through more complex control algorithms. In addition, remote monitoring systems based on Internet of Things technology have gradually been applied to cold chain logistics, realizing real-time monitoring and abnormal alarm of the temperature and humidity status of vehicle cold chains.
[0003] However, the existing vehicle cold chain temperature and humidity control methods still have many deficiencies. First, traditional methods such as threshold control and PID control lack foresight and cannot adjust the control strategy in advance according to the future temperature and humidity change trends, resulting in control lag and difficulty in coping with complex and changeable transportation environments. Second, these methods usually regard temperature and humidity as independent control objects, ignoring the coupling relationship between the two, and it is difficult to achieve coordinated optimization control of temperature and humidity. Third, traditional control methods do not adequately consider energy consumption optimization, often leading to frequent start and stop of the refrigeration system, increasing energy consumption and reducing the equipment life. In addition, existing methods generally lack the ability of adaptive learning and cannot continuously optimize the control strategy based on historical operation data, making it difficult to meet the special requirements of different cargo types and transportation environments. Finally, most existing methods lack intelligent means for handling abnormal situations. Once equipment failures or environmental mutations occur, manual intervention is often required, making it difficult to ensure the continuity and reliability of cold chain quality. Summary of the Invention
[0004] This application provides a method, system, and storage medium for intelligently controlling the temperature and humidity of vehicle cold chains, which is used to predict the future temperature and humidity change trends based on historical data and real-time status, realize coordinated optimization control of temperature and humidity, and have the ability of self-learning and continuous optimization, effectively improving the accuracy, stability, and energy utilization efficiency of cold chain temperature and humidity control.
[0005] In a first aspect, the present application provides a method for intelligently controlling the temperature and humidity of a vehicle-mounted cold chain. The method for intelligently controlling the temperature and humidity of a vehicle-mounted cold chain includes: collecting the temperature and humidity data inside and outside the vehicle compartment through multiple temperature and humidity sensors inside the vehicle compartment, and processing the temperature and humidity data to obtain preprocessed temperature and humidity data; based on the preprocessed temperature and humidity data, using a three-way decision-making model to divide the temperature and humidity state inside the vehicle compartment to obtain a state evaluation result of the temperature and humidity inside the vehicle compartment; using the preprocessed temperature and humidity data and the state evaluation result, constructing a temperature and humidity trend analysis model based on a long short-term memory network and an attention mechanism to obtain a temperature and humidity change trend; according to the state evaluation result and the temperature and humidity change trend, using multi-agent deep reinforcement learning to construct a control decision-making system to obtain a temperature and humidity control decision; based on the temperature and humidity control decision, dividing the control decision into multiple operation modes, and combining the state evaluation result and the temperature and humidity change trend to execute a control strategy to obtain an execution effect; according to the execution effect, adaptively updating the control decision-making system to obtain an optimized temperature and humidity control system.
[0006] In a second aspect, the present application provides a system for intelligently controlling the temperature and humidity of a vehicle-mounted cold chain. The system for intelligently controlling the temperature and humidity of a vehicle-mounted cold chain includes:
[0007] A processing module, configured to collect the temperature and humidity data inside and outside the vehicle compartment through multiple temperature and humidity sensors inside the vehicle compartment, and process the temperature and humidity data to obtain preprocessed temperature and humidity data;
[0008] A partitioning module, configured to divide the temperature and humidity state inside the vehicle compartment using a three-way decision-making model based on the preprocessed temperature and humidity data to obtain a state evaluation result of the temperature and humidity inside the vehicle compartment;
[0009] An analysis module, configured to construct a temperature and humidity trend analysis model based on a long short-term memory network and an attention mechanism using the preprocessed temperature and humidity data and the state evaluation result to obtain a temperature and humidity change trend;
[0010] A construction module, configured to construct a control decision-making system using multi-agent deep reinforcement learning according to the state evaluation result and the temperature and humidity change trend to obtain a temperature and humidity control decision;
[0011] A control module, configured to divide the control decision into multiple operation modes based on the temperature and humidity control decision, and combine the state evaluation result and the temperature and humidity change trend to execute a control strategy to obtain an execution effect;
[0012] An update module, configured to adaptively update the control decision-making system according to the execution effect to obtain an optimized temperature and humidity control system.
[0013] In a third aspect, there is provided a device for intelligently controlling the temperature and humidity of a vehicle-mounted cold chain, including: a memory and at least one processor, wherein instructions are stored in the memory; the at least one processor invokes the instructions in the memory so that the device for intelligently controlling the temperature and humidity of the vehicle-mounted cold chain executes the above-mentioned method for intelligently controlling the temperature and humidity of the vehicle-mounted cold chain.
[0014] In a fourth aspect, there is provided a computer-readable storage medium, in which instructions are stored, and when it runs on a computer, it causes the computer to execute the above-mentioned method for intelligently controlling the temperature and humidity of the vehicle-mounted cold chain.
[0015] In the technical solution provided by this application, the present invention collects temperature and humidity data through multiple temperature and humidity sensors arranged in the carriage and performs preprocessing, constructing a comprehensive and accurate data basis, avoiding the one-sidedness of data caused by single-point sampling, and ensuring the reliability of subsequent analysis and decision-making. The three-way decision-making model is used to divide the temperature and humidity state in the carriage, and the state is accurately classified into the positive domain, negative domain, and boundary domain, overcoming the limitations of traditional binary judgment, and being able to more carefully identify the critical situations of the temperature and humidity state, providing a clear state representation for precise control. The temperature and humidity trend analysis model constructed based on the long short-term memory network and attention mechanism effectively captures the temporal characteristics and long-term dependence relationships of temperature and humidity data, automatically identifies key historical data points through the attention mechanism, improves the prediction accuracy of temperature and humidity changes in complex environments, enables the control system to have foresight, and can anticipate and respond to temperature and humidity anomalies in advance. The control decision-making system is constructed by multi-agent deep reinforcement learning, decomposing temperature control and humidity control into sub-tasks that work together. Through the cooperation mechanism between agents and the design of the global reward function, the coupling problem in temperature and humidity control is solved, and the coordinated optimization control of temperature and humidity is realized. The control decision is divided into multiple operation modes and the control strategy is executed based on the state evaluation result and the temperature and humidity change trend, improving the adaptability of the control system to different working conditions, being able to adopt the optimal control strategy for different demand scenarios such as normal, rapid temperature adjustment, rapid humidity adjustment, energy saving, and emergency, improving the control accuracy and response speed, and at the same time reducing energy consumption. The mechanism for adaptively updating the control decision-making system enables the system to have the ability of continuous learning and self-optimization, being able to mine the optimization space from historical operation data, and continuously adjusting and improving the control strategy according to the execution effect, improving the robustness and long-term performance of the system. The present invention combines artificial intelligence algorithms with the cold chain physical model for the specific application field of vehicle-mounted cold chain temperature and humidity control, giving full play to the advantages of deep reinforcement learning in complex decision-making problems, the expertise of the long short-term memory network in processing time-series data, and the ability of the attention mechanism in identifying key information. At the same time, by introducing physical constraints and a multi-mode control architecture, it ensures that the algorithm output conforms to the actual physical laws and control requirements, realizes the substantial contribution of algorithm features to the solution, solves the complex scenario control problem that is difficult to handle by traditional methods, and significantly improves the accuracy, stability, foresight, and energy utilization efficiency of vehicle-mounted cold chain temperature and humidity control. Brief Description of the Drawings
[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for description in the embodiments. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0017] Figure 1 It is a schematic diagram of an embodiment of the method for intelligently controlling the temperature and humidity of a vehicle-mounted cold chain in an embodiment of the present application;
[0018] Figure 2 It is a schematic diagram of an embodiment of the system for intelligently controlling the temperature and humidity of a vehicle-mounted cold chain in an embodiment of the present application;
[0019] Figure 3 It is a structural schematic block diagram of the device for intelligently controlling the temperature and humidity of a vehicle-mounted cold chain in an embodiment of the present invention. Detailed implementation manners
[0020] The embodiments of the present application provide a method, a system and a storage medium for intelligently controlling the temperature and humidity of a vehicle-mounted cold chain. Terms such as "first", "second", "third", "fourth", etc. (if any) in the specification, claims and above-mentioned drawings of the present application are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" or "having" and any variation thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0021] For ease of understanding, the specific process of the embodiments of the present application will be described below. Please refer to Figure 1 , an embodiment of the method for intelligently controlling the temperature and humidity of a vehicle-mounted cold chain in an embodiment of the present application includes:
[0022] Step S101: Collect the temperature and humidity data inside and outside the carriage through multiple temperature and humidity sensors inside the carriage, and process the temperature and humidity data to obtain the preprocessed temperature and humidity data;
[0023] Step S102: Based on the preprocessed temperature and humidity data, use a three-way decision-making model to divide the temperature and humidity state inside the carriage to obtain the state evaluation result of the temperature and humidity inside the carriage;
[0024] Step S103: Use the preprocessed temperature and humidity data and the state evaluation result to construct a temperature and humidity trend analysis model based on a long short-term memory network and an attention mechanism to obtain the temperature and humidity change trend;
[0025] Step S104: According to the state evaluation result and the temperature and humidity change trend, use multi-agent deep reinforcement learning to construct a control decision-making system to obtain the temperature and humidity control decision;
[0026] Step S105: Based on the temperature and humidity control decision, divide the control decision into multiple operation modes, and combine the state evaluation result and the temperature and humidity change trend to execute the control strategy to obtain the execution effect;
[0027] Step S106: According to the execution effect, adaptively update the control decision system to obtain an optimized temperature and humidity control system.
[0028] It can be understood that the execution subject of this application can be a system for intelligently controlling the temperature and humidity of a vehicle-mounted cold chain, or a terminal or a server. Specifically, it is not limited here. This embodiment of the application takes the server as the execution subject as an example for illustration.
[0029] Specifically, the method for intelligently controlling the temperature and humidity of a vehicle-mounted cold chain according to the present invention first collects the temperature and humidity data inside and outside the carriage through multiple temperature and humidity sensors inside the carriage, and then performs data preprocessing. In a specific implementation, 4 sensor nodes are set in the top area inside the carriage, 6 sensor nodes are set in the middle area, and 4 sensor nodes are set in the bottom area to form a three-dimensional monitoring network. The data collected by these sensors includes the real-time temperature values and humidity values at each point inside the carriage, as well as information such as the environmental temperature, environmental humidity, vehicle speed, and door switch state outside the carriage. When preprocessing the collected raw data, first use the local outlier factor algorithm to detect outliers. This algorithm identifies outliers by calculating the local density ratio of a data point to its neighboring points. Then, use the time series interpolation algorithm to complement the detected abnormal data and missing data, and perform interpolation based on the valid data at adjacent time points. Finally, use the moving average method to smooth the data to reduce the influence of random noise.
[0030] Based on the preprocessed temperature and humidity data, use the three-way decision model to divide the temperature and humidity state inside the carriage to obtain the state evaluation result of the temperature and humidity inside the carriage. The three-way decision model is based on the three-way decision theory and divides the temperature and humidity state inside the carriage into three states: the positive domain (POS), the negative domain (NEG), and the boundary domain (BND). According to the optimal storage conditions of the transported goods, determine the temperature threshold and humidity threshold. When the temperature and humidity at all points inside the carriage are within the target range, the system state belongs to the positive domain, indicating an ideal state; when the temperature or humidity of a point exceeds the range by a certain value, the system state belongs to the negative domain, indicating an abnormal state; other situations belong to the boundary domain, indicating a critical state that needs to be regulated. At the same time, calculate the spatial uniformity index of the temperature and humidity inside the carriage, including the standard deviation of the temperature and humidity at each measurement point, to obtain the temperature and humidity distribution uniformity evaluation value. Combine the temperature and humidity state with the uniformity evaluation value to generate a comprehensive state evaluation result as the input for subsequent control decisions.
[0031] Using the preprocessed temperature and humidity data and the state evaluation results, a temperature and humidity trend analysis model is constructed based on the long short-term memory network and the attention mechanism to obtain the temperature and humidity change trends. The model first divides the historical data according to the time window to form input sequences containing multiple consecutive sampling points. Each sequence contains features such as temperature vectors, humidity vectors, ambient temperature, ambient humidity, vehicle speed, and door status. After normalizing these time-series feature data, they are input into a neural network composed of three layers of LSTM for processing. The first layer contains 128 neurons, the second layer contains 64 neurons, and the third layer contains 32 neurons. Dropout layers are added between each layer to prevent overfitting. The time-series features extracted by the LSTM network are weighted by the attention mechanism to automatically identify historical data points that have a greater impact on the prediction results. The weighted features are mapped to the output space through a fully connected layer to obtain the temperature and humidity change trends at multiple future time points. At the same time, a compartment heat transfer model is added to the model as a physical constraint to ensure that the prediction results conform to physical laws.
[0032] According to the state evaluation results and the temperature and humidity change trends, a control decision-making system is constructed using multi-agent deep reinforcement learning to obtain temperature and humidity control decisions. The system decomposes the control task into two subtasks: temperature control and humidity control, which are collaboratively completed by the temperature control agent and the humidity control agent respectively. The temperature control agent is responsible for adjusting the compressor power, fan speed, and circulation damper opening; the humidity control agent is responsible for adjusting the desiccant wheel speed and the regeneration heater power. The state space of the control system includes the current temperature and humidity distribution inside the compartment, ambient temperature and humidity, vehicle speed, door status, state evaluation results, and temperature and humidity change trends. The reward functions of each agent are designed considering temperature and humidity deviation penalties, uniformity penalties, and energy consumption penalties. To achieve cooperation between agents, a global reward function is designed, combining the rewards of each agent and the cooperation reward term. The agents adopt the Actor-Critic network architecture and are trained through the proximal policy optimization algorithm to generate the optimal control strategy.
[0033] Based on the temperature and humidity control decisions, the control decisions are divided into multiple operation modes, and the control strategy is executed in combination with the state evaluation results and the temperature and humidity change trends to obtain the execution effect. The operation modes include normal mode, rapid temperature adjustment mode, rapid humidity adjustment mode, energy-saving mode, and emergency mode. Each operation mode corresponds to different control objectives and actuator control strategy matrices. The actuator control adopts a hierarchical architecture. The high-level controller generates the control objectives of each subsystem according to the selected operation mode and the control actions output by the agent; the low-level controller adopts a proportional-integral control algorithm to ensure that each actuator accurately tracks the control objective. After the system executes the control action, it monitors the control effect in real time, compares the real-time temperature and humidity change data with the predicted temperature and humidity change trends, calculates the temperature and humidity deviation value, determines the execution effect of the control strategy, and uses it as the basis for subsequent adaptive updates.
[0034] According to the execution effect, the control decision-making system is adaptively updated to obtain an optimized temperature and humidity control system. This step first deeply mines the collected historical operation data, and uses the hierarchical clustering algorithm to divide the historical operation data into multiple typical operation modes according to similarity. For each operation mode, key performance indicators are calculated, including temperature adjustment rate, humidity adjustment rate, and energy utilization efficiency. Based on these performance evaluation data, working condition points with significant performance differences are identified, and the operation rules and potential optimization space of the control system are extracted. The temperature and humidity trend analysis model uses an incremental learning strategy to update parameters, adding newly collected data to the training set. The control decision-making system updates parameters by maintaining an experience buffer pool to store the state-action-reward sequence during operation. When the data volume in the buffer pool reaches the threshold, policy distillation technology is used for parameter update. At the same time, the system also uses an anomaly detection method based on the isolation forest algorithm to monitor the operation parameters of the cold chain system in real time, and promptly discovers and handles abnormal situations.
[0035] Taking the transportation of frozen meat as an example, the target temperature range in the carriage is set from -18°C to -15°C, and the target humidity range is from 85% to 90%. During a transportation process, the system detects that the temperature in the middle area of the carriage gradually rises to -14°C, which is higher than the upper target limit, and the state evaluation result is classified as the boundary domain. The temperature and humidity trend analysis model predicts that if no measures are taken, the temperature in this area will rise to -12°C within half an hour. Based on this prediction, the multi-agent control system automatically selects the rapid temperature adjustment mode, increases the compressor power output, and at the same time adjusts the fan speed and the opening degree of the circulation air damper to make more cold air flow to the middle area. After 20 minutes of execution, the temperature in the middle area drops to -16°C, returning to the target range. The system records the state-action-reward sequence of this temperature adjustment process, updates the control strategy, and can control the temperature within the target range more quickly and energy-efficiently the next time a similar situation occurs.
[0036] In the embodiments of the present application, the present invention collects and preprocesses temperature and humidity data by arranging multiple temperature and humidity sensors in the carriage, constructs a comprehensive and accurate data basis, avoids the one-sidedness of data caused by single-point sampling, and ensures the reliability of subsequent analysis and decision-making. The three-way decision model is used to classify the temperature and humidity states in the carriage, and the states are accurately classified into the positive domain, negative domain, and boundary domain, overcoming the limitations of traditional binary judgment, and being able to more carefully identify the critical situations of temperature and humidity states, providing a clear state representation for precise control. The temperature and humidity trend analysis model constructed based on the long short-term memory network and attention mechanism effectively captures the temporal characteristics and long-term dependence relationships of temperature and humidity data, automatically identifies key historical data points through the attention mechanism, improves the prediction accuracy of temperature and humidity changes in complex environments, enables the control system to have foresight, and can anticipate and respond to temperature and humidity anomalies in advance. The control decision-making system is constructed by using multi-agent deep reinforcement learning. The temperature control and humidity control are decomposed into sub-tasks that work together. Through the cooperation mechanism between agents and the design of the global reward function, the coupling problem in temperature and humidity control is solved, and the coordinated optimization control of temperature and humidity is realized. The control decision is divided into multiple operation modes and the control strategy is executed based on the state evaluation results and temperature and humidity change trends, improving the adaptability of the control system to different working conditions, being able to adopt the optimal control strategy for different demand scenarios such as normal, rapid temperature adjustment, rapid humidity adjustment, energy saving, and emergency, improving the control accuracy and response speed, and reducing energy consumption at the same time. The mechanism for adaptively updating the control decision-making system enables the system to have the ability of continuous learning and self-optimization, can mine the optimization space from historical operation data, and continuously adjust and improve the control strategy according to the execution effect, improving the robustness and long-term performance of the system. The present invention combines artificial intelligence algorithms with the cold chain physical model for the specific application field of on-vehicle cold chain temperature and humidity control, gives full play to the advantages of deep reinforcement learning in complex decision-making problems, the expertise of the long short-term memory network in processing time-series data, and the ability of the attention mechanism in identifying key information. At the same time, by introducing physical constraints and a multi-mode control architecture, it ensures that the algorithm output conforms to the actual physical laws and control requirements, realizes the substantial contribution of algorithm features to the solution, solves the complex scenario control problem that is difficult to handle by traditional methods, and significantly improves the accuracy, stability, foresight, and energy utilization efficiency of on-vehicle cold chain temperature and humidity control.
[0037] In a specific embodiment, the process of executing step S101 may specifically include the following steps:
[0038] Four sensor nodes are set in the top area of the carriage, six sensor nodes are set in the middle area, and four sensor nodes are set in the bottom area to form a three-dimensional monitoring network, collect the real-time temperature values and humidity values at each point in the carriage, and obtain the temperature and humidity data set in the carriage;
[0039] An environmental temperature and humidity sensor is set outside the carriage to collect the external environmental temperature, environmental humidity, vehicle speed, and door status, obtaining an external carriage dataset;
[0040] The local outlier factor algorithm is used to detect outliers in the in-carriage temperature and humidity dataset and the external carriage dataset, identifying data with too large a difference from adjacent time points to obtain outlier-labeled data;
[0041] Based on the valid data of adjacent time points, the time series interpolation algorithm is applied to the outlier-labeled data to complete the missing values, obtaining a complete dataset;
[0042] The moving average method is applied to the complete dataset for smoothing processing to obtain noise-reduced temperature and humidity data;
[0043] The noise-reduced temperature and humidity data are combined with the cargo type information, recording the physical and chemical properties of the cargo and the optimal preservation temperature and humidity range to obtain preprocessed temperature and humidity data for temperature and humidity status evaluation.
[0044] Specifically, four temperature and humidity sensor nodes are set in the top area of the carriage. These four nodes are located at the four corners of the carriage top, forming a rectangular distribution; six sensor nodes are set in the middle area, of which four are located at the central positions of the four sides in the middle of the carriage, and the other two are located in the center of the middle space of the carriage, one is slightly forward and the other is slightly backward; four sensor nodes are set in the bottom area, distributed at the four corners of the carriage bottom, corresponding to the top nodes. Such a distribution method forms a three-dimensional monitoring network to ensure comprehensive monitoring of the temperature and humidity conditions at different positions in the carriage. Each sensor node uses a high-precision temperature and humidity sensor, with a temperature measurement accuracy of ±0.1 °C and a humidity measurement accuracy of ±2%RH. These sensors transmit the collected data to the vehicle-mounted central processing unit through a bus or wireless communication method, and the data collection frequency is set to once every 30 seconds. The data collected by the sensors in the carriage include the real-time temperature values and humidity values at each point.
[0045] At the same time, an environmental temperature and humidity sensor is set outside the carriage to collect the external environmental temperature and environmental humidity. In addition, vehicle speed information is obtained through the vehicle's CAN bus, and the door switch status is obtained through a door magnetic switch sensor. These data together constitute the external carriage dataset, providing important environmental parameter information for the intelligent control system. The in-carriage temperature and humidity dataset and the external carriage dataset together form the original dataset, which needs to be further processed before being used for subsequent intelligent control decisions.
[0046] For the collected raw data, the Local Outlier Factor (LOF) algorithm is first used for outlier detection. The LOF algorithm is a density-based outlier detection method, and its core idea is to identify outliers by comparing the local density differences between data points and their neighborhoods. The specific steps include: first calculating the distances between data points, then determining the neighbors of each data point, calculating the local reachability density of each data point, and finally calculating the local outlier factor. For time series data, the data of consecutive multiple time points form a sliding window, and the distance and density relationships between each data point in the window and its neighboring points are calculated. When the local outlier factor value of a certain data point is significantly higher than the threshold, it is marked as abnormal data. For example, when the temperature of a certain sensor suddenly changes by more than 3°C or the humidity changes by more than 10% within a short period of time, and this change is not observed in other sensors, this data point will be marked as abnormal.
[0047] For the detected abnormal data and the missing data points, a time series interpolation algorithm is used for completion. The time series interpolation algorithm performs interpolation based on the valid data of adjacent time points. Common methods include linear interpolation, spline interpolation, and polynomial interpolation. In this method, the cubic spline interpolation method is mainly used. This method not only considers the continuity of the data but also the smoothness of the data. The specific operation is to use several valid data points before and after the abnormally marked data as reference points to construct a cubic spline function, and then calculate the interpolation result corresponding to the abnormal point. This method can effectively retain the trend and smoothness of the temperature and humidity changes, and avoid the zigzag changes that may be brought by simple linear interpolation.
[0048] After completing the filling of missing values, the moving average method is applied to the complete data set for smoothing processing to reduce the interference of random noise on the system judgment. The moving average method is to take the average of the data of the current time point and a certain number of time points before and after it as the new value of the current time point. In this method, the size of the sliding window is set to 5 sampling points, that is, the data of the current point and two points before and after it are used for averaging. Through the moving average processing, the data fluctuations within a short period of time can be effectively eliminated, and a more stable temperature and humidity change trend can be obtained.
[0049] Finally, the noise-reduced temperature and humidity data are combined with the cargo type information to record the physical and chemical properties of the cargo and the optimal storage temperature and humidity range. A parameter library of different types of goods is preset in the system, including the optimal storage temperature range and humidity range of different goods such as frozen meat, fresh fruits and vegetables, vaccines, etc. According to the type of goods being transported currently, the system automatically selects the corresponding parameter settings. These information, together with the noise-reduced temperature and humidity data, constitute the preprocessed temperature and humidity data, providing a basis for subsequent status evaluation.
[0050] Taking the transportation of chilled fresh food as an example, a cold chain vehicle transports a batch of fresh food that needs to be kept in an environment of 2°C to 8°C and a relative humidity of 75% to 85%. 14 sensors in the carriage are installed according to the above distribution method, and data is collected every 30 seconds. During a certain operation, sensor 11 suddenly reported a temperature of 15°C at a certain moment, while the other sensors still showed 4 - 6°C. The local outlier factor algorithm calculates that the outlier factor value of this point is much higher than the set threshold, so this point is marked as abnormal data. Subsequently, using the time series interpolation algorithm, the valid data points within two minutes before and after this sensor are taken, and the temperature at this moment is calculated to be 5.4°C through cubic spline interpolation. After processing all abnormal points, the sliding average method is applied to the entire data set for smoothing to eliminate random fluctuations within a short period. Finally, the system combines the processed temperature and humidity data with the parameters of the currently transported fresh food, records the optimal preservation temperature range and humidity range, and provides an accurate data basis for subsequent temperature and humidity state evaluation and control decision-making.
[0051] In a specific embodiment, the process of executing step S102 may specifically include the following steps:
[0052] According to the optimal preservation conditions of the transported goods, set the temperature threshold and humidity threshold to obtain the temperature and humidity target range;
[0053] Analyze the preprocessed temperature and humidity data. When the temperature of all points in the carriage is within the temperature threshold range and the humidity is within the humidity threshold range, divide the system state into the positive domain to obtain an ideal state identifier;
[0054] Analyze the preprocessed temperature and humidity data. When there are points in the carriage where the temperature or humidity exceeds the temperature threshold range or the humidity threshold range by a certain value, divide the system state into the negative domain to obtain an abnormal state identifier;
[0055] Analyze the preprocessed temperature and humidity data, and divide the states that do not belong to the positive domain and the negative domain into the boundary domain to obtain a critical state identifier;
[0056] Calculate the spatial uniformity index of the temperature and humidity in the carriage, including the standard deviation of the temperature and humidity at each measurement point in the carriage, to obtain the temperature and humidity distribution uniformity evaluation value;
[0057] Combine the ideal state identifier, abnormal state identifier, or critical state identifier with the temperature and humidity distribution uniformity evaluation value to generate a comprehensive state evaluation result as the input for control decision-making.
[0058] Specifically, the temperature threshold and humidity threshold are set according to the optimal storage conditions of the transported goods to obtain the temperature and humidity target ranges. Different types of cold-chain goods have different optimal storage temperature and humidity conditions, which are usually determined by food safety standards or drug storage specifications. In practical applications, the corresponding temperature and humidity parameters are obtained by querying the cold-chain goods database. For example, the optimal storage temperature range for frozen meat is from -18°C to -15°C, and the humidity range is from 85% to 90%; the optimal storage temperature range for fresh fruits and vegetables is from 2°C to 8°C, and the humidity range is from 90% to 95%; the optimal storage temperature range for vaccines is from 2°C to 5°C, and the humidity range is from 45% to 65%. These parameters are set as the control target ranges and used as the criteria for state evaluation.
[0059] After obtaining the temperature and humidity target ranges, a three-way decision model is used to divide the temperature and humidity states inside the carriage. The three-way decision model is derived from the three-way decision theory and divides the decision space into three parts: the positive domain, the negative domain, and the boundary domain. In this method, the positive domain represents the ideal state, the negative domain represents the abnormal state, and the boundary domain represents the critical state. First, the preprocessed temperature and humidity data are analyzed to determine whether the temperature and humidity at each point inside the carriage are within the target ranges. When the temperature at all measurement points inside the carriage is within the set temperature threshold range and the humidity is within the set humidity threshold range, the system state is divided into the positive domain and an ideal state identifier is generated. The specific judgment process is to traverse the temperature and humidity data of all sensor nodes inside the carriage and check whether the temperature of each node is within the temperature threshold range and the humidity is within the humidity threshold range. Only when the temperature and humidity data of all nodes meet the conditions, an ideal state identifier is generated. Then, the preprocessed temperature and humidity data are continuously analyzed. When the temperature or humidity of a point inside the carriage exceeds the set threshold range by a certain value, the system state is divided into the negative domain and an abnormal state identifier is generated. Here, "exceeding a certain value" means that the degree of exceeding the threshold range reaches the preset fault tolerance value. For example, if the set temperature fault tolerance value is 0.5°C and the humidity fault tolerance value is 3%, when there is a measurement point where the temperature is lower than the lowest temperature threshold minus 0.5°C, or higher than the highest temperature threshold plus 0.5°C, or the humidity is lower than the lowest humidity threshold minus 3%, or higher than the highest humidity threshold plus 3%, it is determined as an abnormal state. This design takes into account the actual requirements of cold-chain temperature and humidity control and avoids frequent state switching caused by minor fluctuations.
[0060] For the state that neither belongs to the positive domain nor the negative domain, that is, the temperature and humidity of a measurement point exceed the target range but do not reach the abnormal state determination standard, it is divided into the boundary domain and a critical state identifier is generated. The critical state is a state that needs attention but has not reached the abnormal level and is also a state that the control system needs to adjust. The specific judgment process is to first exclude the situations that meet the conditions of the positive domain and the negative domain, and the remaining states are all divided into the boundary domain.
[0061] In addition to determining whether the absolute values of temperature and humidity are within the target range, it is also necessary to calculate the spatial uniformity index of the temperature and humidity in the carriage to evaluate the degree of uniformity of the temperature and humidity distribution. The spatial uniformity index is represented by calculating the standard deviation of the temperature and humidity at each measurement point in the carriage. The smaller the standard deviation, the more uniform the temperature and humidity distribution; the larger the standard deviation, the more uneven the temperature and humidity distribution. The process of calculating the standard deviation is to first find the average value of the temperatures at all measurement points, then calculate the square of the difference between each measurement point temperature and the average value, sum them up and divide by the number of measurement points, and finally take the square root. The calculation method of the humidity standard deviation is similar.
[0062] Combine the ideal state identifier, abnormal state identifier, or critical state identifier obtained previously with the evaluation value of the temperature and humidity distribution uniformity to generate a comprehensive state evaluation result. The combination method is to encode the temperature and humidity state classification and the uniformity evaluation result into a comprehensive state code. For example, a two-digit state code can be used. The first digit represents the temperature and humidity state classification (0 represents the positive domain, 1 represents the boundary domain, 2 represents the negative domain), and the second digit represents the uniformity level (0 represents uniform, 1 represents relatively uniform, 2 represents non-uniform). The comprehensive state evaluation result obtained in this way can comprehensively reflect the state of the temperature and humidity environment in the carriage and serve as an important input for subsequent control decisions.
[0063] Taking the cold chain transportation of vaccines as an example, the optimal storage temperature range for a batch of vaccines is 2°C to 5°C, and the humidity range is 45% to 65%. The temperature tolerance value is set to 0.5°C, and the humidity tolerance value is set to 3%. During a transportation process, the preprocessed temperature and humidity data collected by 14 sensors show that the temperatures of 12 sensors are between 2.3°C and 4.7°C, and the humidities are between 48% and 62%. However, the temperatures of 2 sensors are 5.3°C and 5.4°C, and the humidities are 58% and 60%. First, judge whether the positive domain conditions are met. Since there are measurement points with temperatures exceeding 5°C, it does not belong to the positive domain. Then judge whether the negative domain conditions are met. The highest temperature of 5.4°C exceeds the upper limit of 5°C but does not exceed the upper limit plus the tolerance value of 5.5°C, so it does not belong to the negative domain either. Therefore, the current state is classified as the boundary domain, and a critical state identifier is generated. Then calculate the temperature uniformity index. The average value of the temperatures at all measurement points is 3.8°C, and the standard deviation is 0.9°C; the average humidity is 55%, and the standard deviation is 4.2%. According to the set uniformity evaluation criteria, when the temperature standard deviation is less than 1.0°C and the humidity standard deviation is less than 5%, it is determined that the temperature and humidity distribution is uniform. Therefore, the uniformity evaluation value is "uniform". Finally, combine the critical state identifier and the uniformity evaluation value to generate a comprehensive state evaluation result, encoded as "10", indicating "critical state and uniform". This comprehensive evaluation result will be used as an input to the control decision system to guide the system to perform appropriate temperature and humidity regulation.
[0064] In a specific embodiment, the process of executing step S103 may specifically include the following steps:
[0065] The pre - processed temperature and humidity data are divided according to time windows. Each time window contains multiple consecutive sampling points, forming an input sequence to obtain time - series feature data;
[0066] Perform feature normalization on the time - series feature data, convert temperature, humidity, and environmental parameters to a unified dimension to obtain normalized feature data;
[0067] Input the normalized feature data into a three - layer long short - term memory network for processing. The first layer contains 128 neurons, the second layer contains 64 neurons, and the third layer contains 32 neurons. A dropout layer is added between each layer to obtain a time - series feature representation;
[0068] Apply the attention mechanism to weight the time - series feature representation, sort the importance of historical data points by calculating weight coefficients to obtain a weighted feature representation;
[0069] Map the weighted feature representation to the output space through a fully - connected layer to obtain predicted values of the temperature and humidity change trends at multiple future time points;
[0070] Add a car body heat transfer model as a physical constraint to the predicted values of the temperature and humidity change trends, and optimize the model parameters through a loss function composed of the mean square error loss and the physical model constraint regularization term to obtain the temperature and humidity change trends that conform to physical laws.
[0071] Specifically, based on the pre - processed temperature and humidity data and the state evaluation results, a temperature and humidity trend analysis model needs to be constructed to accurately predict the future temperature and humidity change trends. First, the pre - processed temperature and humidity data are divided according to time windows. Each time window contains multiple consecutive sampling points, forming an input sequence to obtain time - series feature data. A time window refers to the range of data samples selected for analysis within a continuous period of time in time - series analysis. In this method, the length of the time window is set to L sampling points, and there can be overlap or no overlap between adjacent windows. For example, if the sampling frequency is once every 30 seconds and the window length L is set to 20, then each window contains 10 minutes of continuous data. For the data within each time window, the extracted features include the temperature vector inside the car, the humidity vector, the environmental temperature, the environmental humidity, the vehicle speed, the door state, etc. These data together form the input sequence for subsequent deep - learning model training.
[0072] After the extraction of time series feature data is completed, it is necessary to perform feature normalization on it, convert temperature, humidity, and environmental parameters to a unified dimension, and obtain normalized feature data. Normalization is the process of converting features with different dimensions and different numerical ranges to the same scale, which helps to improve the stability and convergence speed of model training. There are usually two normalization methods: min-max normalization and Z-score standardization. Min-max normalization scales the feature values to the range [0, 1], while Z-score standardization converts the features to a distribution with a mean of 0 and a standard deviation of 1. In this method, min-max normalization is used for temperature data, linear scaling is performed on humidity data (which is already in the range of 0 - 100%), and Z-score standardization is used for other features such as vehicle speed. After such processing, the data in each dimension has a similar numerical range, which is beneficial to the training of neural networks.
[0073] After the preparation of the normalized feature data, it is input into a three-layer long short-term memory network for processing. The long short-term memory network (LSTM) is a special type of recurrent neural network that can effectively process sequential data and capture long-term dependencies. LSTM solves the problem of gradient vanishing in traditional recurrent neural networks by introducing memory cells and gating mechanisms. In this method, a three-layer LSTM network is constructed. The first layer contains 128 neurons, the second layer contains 64 neurons, and the third layer contains 32 neurons. A dropout layer is added between each layer, and the dropout rate is set to 0.2 to prevent overfitting. The dropout layer is a regularization technique that randomly sets the outputs of a part of the neurons to zero during the training process, forcing the network to learn more robust feature representations. The output after processing by the LSTM network is the time series feature representation, which captures the time dependence and variation law of the temperature and humidity data.
[0074] For the time series feature representation extracted by LSTM, the attention mechanism is applied for weighting. By calculating the weight coefficients, the importance of historical data points is ranked, and a weighted feature representation is obtained. The attention mechanism is a technique that enables the model to focus on the important parts of the input sequence and is particularly effective in time series prediction. In this method, the dot product attention mechanism is adopted. The dot product of the query vector and the key vector is calculated, and the weight coefficients are obtained through the softmax function, and then the value vector is weighted and summed. Specifically, the hidden state of LSTM is used as the key vector and the value vector, and the query vector can be the hidden state of the last time step or obtained through additional parameter learning. Through the attention mechanism, the model can automatically identify the historical data points that have a greater impact on the prediction result and assign higher weights, thereby improving the prediction accuracy.
[0075] The weighted feature representation is mapped to the output space through a fully connected layer to obtain the predicted values of the temperature and humidity change trends at multiple future time points. The fully connected layer is a basic neural network layer that maps the input features to the output space through a linear transformation and a non-linear activation function. In this method, the input of the fully connected layer is the weighted feature representation, and the output is the predicted values of the temperature and humidity inside the carriage at the next K time points. The number of neurons in the output layer is K×N, where K is the number of predicted time steps and N is the number of sensor nodes inside the carriage. The activation function of the fully connected layer is selected as ReLU to introduce non-linearity and avoid the vanishing gradient problem.
[0076] Finally, the carriage heat transfer model is added to the predicted values of the temperature and humidity change trends as a physical constraint, and the model parameters are optimized through a loss function composed of the mean square error loss and the physical model constraint regularization term to obtain the temperature and humidity change trends that conform to physical laws. The loss function is defined as follows:
[0077]
[0078] Among them, represents the total loss function, represents the model parameters, M represents the number of training samples, represents the predicted value of the model for the i-th sample, represents the true value of the i-th sample, represents the square of the Euclidean distance, is the weight coefficient, is the physical model constraint regularization term. The physical model constraint regularization term is based on the thermodynamic equation and considers factors such as heat conduction through the carriage wall, convective heat transfer of air flow, and heat exchange caused by door opening and closing, and is defined as follows:
[0079]
[0080] Among them, and respectively represent the temperature and humidity of the j-th node predicted by the model at time t, and respectively represent the external environmental temperature and humidity at time t, represents the door state at time t, represents the vehicle speed at time t, represents the carriage heat transfer model function based on the thermodynamic equation. represents the physical model constraint regularization term, which is used to constrain the prediction results to conform to physical laws. K represents the number of predicted time steps, that is, how many time points in the future the temperature and humidity values are predicted, and N represents the number of sensor nodes inside the carriage. By adding this regularization term, while learning the data pattern, the model will also respect physical laws and avoid generating prediction results that violate the principles of thermodynamics.
[0081] Taking the transportation of refrigerated drugs as an example, the drugs loaded in a refrigerated truck need to be stored in an environment of 2-8°C. The temperature and humidity data are continuously collected by 14 sensors in the carriage every 30 seconds, and one day's historical data is accumulated. First, these data are divided into groups of 20 sampling points. Each group of data contains 10 minutes of continuous records, and there is an overlap of 10 points between adjacent groups. For each group of data, features such as the temperature and humidity of the 14 sensors in the carriage, ambient temperature, ambient humidity, vehicle speed, and door status are extracted. Then, these features are normalized. The temperature data is transformed to the [0,1] interval through min-max normalization, the humidity data is linearly scaled, and the vehicle speed is standardized by Z-score. The normalized data is input into a three-layer LSTM network and processed layer by layer through 128, 64, and 32 neurons. A dropout layer with dropout = 0.2 is added between each layer to prevent overfitting. The temporal features output by the LSTM are weighted through an attention mechanism to identify key historical data points. For example, the data points after the change in the door switch state will obtain higher weights because the data at these moments is more important for predicting future temperature and humidity changes. The weighted features are mapped to the output space through a fully connected layer to predict the temperature and humidity changes at each sensor point within the next 20 minutes. Finally, by adding a physical constraint regularization term, it is ensured that the prediction results conform to the laws of thermodynamics. For example, when the prediction shows that the temperature in a certain area rises abnormally, the heat transfer model will check whether this conforms to the influence laws of external factors such as the opening of the car door or the change in ambient temperature. If not, the model will correct the prediction results.
[0082] In a specific embodiment, the process of executing step S104 may specifically include the following steps:
[0083] Decompose the control task into two subtasks of temperature control and humidity control, respectively construct a temperature control agent and a humidity control agent, and form a multi-agent control framework;
[0084] Define the control system state space, including the current temperature and humidity distribution in the carriage, ambient temperature and humidity, vehicle speed, door status, state evaluation results, and temperature and humidity change trends, to obtain the agent state representation;
[0085] Define the action space of the temperature control agent to include the compressor power, fan speed, and circulation air damper opening, and the action space of the humidity control agent to include the desiccant wheel speed and regeneration heater power, to obtain the set of control actions that the agent can execute;
[0086] Construct the reward function of the temperature control agent to include a temperature deviation penalty term, a temperature non-uniformity penalty term, and an energy consumption penalty term, and the reward function of the humidity control agent to include a humidity deviation penalty term, a humidity non-uniformity penalty term, and an energy consumption penalty term, to obtain the agent optimization objective;
[0087] Design a global reward function, which combines the rewards of the temperature control agent, the humidity control agent, and the cooperation reward term, and is trained by the proximal policy optimization algorithm to obtain the parameters of the agent's policy network;
[0088] Input the state evaluation result and the temperature and humidity change trend into the trained agent policy network to generate an optimal control action sequence and obtain the temperature and humidity control decision.
[0089] Specifically, decompose the control task into two subtasks of temperature control and humidity control, and construct a temperature control agent and a humidity control agent respectively to form a multi-agent control framework. Multi-agent reinforcement learning is a distributed decision-making method, where each agent is responsible for making decisions on specific subtasks and achieves overall optimization through a coordination mechanism. In the temperature and humidity control of on-vehicle cold chain, temperature and humidity are two interrelated and relatively independent control objects, which are adjusted by different actuators respectively. The temperature control agent is mainly responsible for adjusting the operating parameters of the refrigeration system, while the humidity control agent is responsible for adjusting the operating parameters of the dehumidification or humidification equipment. This decomposition design conforms to the physical composition of the cold chain system and is also convenient for flexible adjustment according to the special needs of different cargo types. Define the control system state space, which includes the current temperature and humidity distribution in the carriage, the ambient temperature and humidity, the vehicle speed, the door state, the state evaluation result, and the temperature and humidity change trend, to obtain the agent state representation. The state space is the basis for the agent to perceive the environment and contains all the key information required for decision-making. The current temperature and humidity distribution in the carriage is composed of the real-time temperature and humidity data collected by 14 sensor nodes, reflecting the temperature and humidity conditions in different areas of the carriage; the ambient temperature and humidity data come from the external sensors of the carriage, providing important information about external interference; the vehicle speed information is obtained through the vehicle CAN bus, reflecting the vehicle operating state; the door state is obtained through the door magnetic switch sensor, recording the door opening and closing state, which has an important impact on the temperature and humidity fluctuations; the state evaluation result is the classification of the carriage temperature and humidity state generated by the three-way decision-making model, including the identification of the positive domain, negative domain, and boundary domain, as well as the uniformity evaluation value; the temperature and humidity change trend is the future temperature and humidity change situation predicted by the model constructed by the long short-term memory network and the attention mechanism. These information together constitute the complete state representation and provide a basis for the agent's decision-making.
[0090] Define the action space of the temperature control agent to include compressor power, fan speed, and the opening degree of the circulation air damper, and the action space of the humidity control agent to include the rotation speed of the desiccant wheel and the power of the regeneration heater, to obtain the set of control actions executable by the agent. The action space defines all possible actions that the agent can take. Compressor power controls the refrigeration capacity of the compartment, with a value range from 0 to the maximum power, usually expressed as a percentage; fan speed controls the cold air circulation speed, affecting temperature uniformity and adjustment rate, with a value range from 0 to the maximum speed; the opening degree of the circulation air damper controls the cold air flow direction, used to adjust the temperature distribution in different areas, with a value range from 0° to 90°. The rotation speed of the desiccant wheel controls the rate of humidity reduction, with a value range from 0 to the maximum speed; the power of the regeneration heater controls the regeneration efficiency of the desiccant, with a value range from 0 to the maximum power. These parameters together constitute the set of control actions of the agent. By adjusting these parameters, the agent can achieve precise control of the temperature and humidity environment in the compartment.
[0091] Construct the reward function of the temperature control agent to include a temperature deviation penalty term, a temperature non-uniformity penalty term, and an energy consumption penalty term, and the reward function of the humidity control agent to include a humidity deviation penalty term, a humidity non-uniformity penalty term, and an energy consumption penalty term, to obtain the optimization objective of the agent. The reward function is the core of reinforcement learning, defining the criteria for evaluating the quality of the agent's behavior. The temperature deviation penalty term calculates the gap between the temperature at each point in the compartment and the target temperature. The smaller the gap, the smaller the penalty; the temperature non-uniformity penalty term calculates the standard deviation of the temperature distribution in the compartment. The smaller the standard deviation, the smaller the penalty; the energy consumption penalty term calculates the energy consumption caused by the control action. The smaller the energy consumption, the smaller the penalty. The reward function design of the humidity control agent is similar to that of the temperature control agent, including the corresponding three penalty terms. Through the combination of these penalty terms, a comprehensive optimization objective that balances temperature and humidity control accuracy, uniformity, and energy efficiency is formed. Design a global reward function, combine the rewards of the temperature control agent, the humidity control agent, and the cooperation reward term, and train through the proximal policy optimization algorithm to obtain the parameters of the agent's policy network. The global reward function is used to promote cooperation between agents, including three parts: temperature control reward, humidity control reward, and cooperation reward. The cooperation reward term mainly considers the mutual influence between temperature and humidity control. For example, the cooling process will cause the relative humidity to increase, and the dehumidification process will cause the temperature to increase. The cooperation reward measures the cooperation effect of the two agents by evaluating the difference between the actual temperature and humidity and the predicted temperature and humidity. The proximal policy optimization algorithm is a classic policy gradient reinforcement learning algorithm, characterized by high sample efficiency and good stability. This algorithm avoids instability caused by excessive policy updates by restricting the difference between the new and old policies. During the training process, an experience replay buffer is used to store transition samples, and each sample contains information about the state, action, reward, and next state. Through multiple iterative trainings, the parameters of the agent's policy network are continuously optimized, and finally the optimal control strategy is learned.
[0092] Input the state evaluation result and the temperature and humidity change trend into the trained intelligent agent policy network to generate an optimal control action sequence and obtain the temperature and humidity control decision. In the actual control process, first obtain information such as the current temperature and humidity distribution, ambient temperature and humidity, vehicle speed, and door state inside the carriage, and combine the state evaluation result and the temperature and humidity change trend obtained in the previous steps to construct the current state representation. Then input the state representation into the policy networks of the trained temperature control intelligent agent and humidity control intelligent agent, and output their respective control actions. The temperature control intelligent agent outputs the control values of the compressor power, fan speed, and circulation air door opening degree, and the humidity control intelligent agent outputs the control values of the desiccant wheel speed and the regeneration heater power. These control values together constitute the temperature and humidity control decision, which is used to guide the subsequent actuator control.
[0093] Taking the cold chain transportation of vaccines as an example, a cold chain vehicle transports a batch of vaccines and needs to maintain a temperature range of 2 - 5°C and a humidity range of 45 - 65%. When the vehicle enters a mountain road section, the external temperature suddenly drops and the humidity increases simultaneously. 14 sensors inside the carriage monitor the temperature and humidity data in real time. Through the three-way decision-making model, it is evaluated that the current state is in the boundary domain, the temperature distribution is uniform but the temperature in some local areas is close to the lower limit. The temperature and humidity trend analysis model predicts that if no measures are taken, the temperature in some areas inside the carriage will be lower than 2°C after 20 minutes, and the humidity will rise above 70%. At this time, input these information into the multi-intelligent agent control decision system. The temperature control intelligent agent calculates according to the state information that the compressor power needs to be reduced to 30%, the fan speed is adjusted to 40%, and the circulation air door opening degree is set to 65° to reduce the cooling capacity and improve the cold air distribution. At the same time, the humidity control intelligent agent calculates that the desiccant wheel speed needs to be increased to 60% and the regeneration heater power is set to 50% to cope with the rising humidity. These control decisions consider factors such as the temperature and humidity change trend, energy consumption optimization, and coordinated operation of equipment, and can effectively respond to environmental changes, keep the temperature and humidity inside the carriage within the ideal range, and ensure the safe transportation of vaccines.
[0094] In a specific embodiment, the process of executing step S105 may specifically include the following steps:
[0095] According to the state evaluation result, divide the control strategy into five operation modes: normal mode, rapid temperature adjustment mode, rapid humidity adjustment mode, energy-saving mode, and emergency mode, and obtain the operation mode set;
[0096] For each operation mode in the operation mode set, design the corresponding actuator control strategy matrix, map the temperature and humidity control decision into specific actuator control parameters, and obtain the actuator control parameter set;
[0097] Adopt a hierarchical control architecture, generate control objectives for each subsystem based on the operation mode set and temperature and humidity control decisions, and obtain the subsystem control objective set;
[0098] Apply the proportional-integral control algorithm to the subsystem control objective set to generate actuator control signals, ensure that each actuator accurately tracks the control objective, and obtain control execution instructions;
[0099] Adjust the compressor power, fan speed, circulation damper opening, desiccant wheel speed, and regeneration heater power through the control execution instructions, perform temperature and humidity adjustment operations, and obtain real-time temperature and humidity change data;
[0100] Compare the real-time temperature and humidity change data with the temperature and humidity change trend, calculate the temperature and humidity deviation value, and determine the execution effect of the control strategy as the basis for adaptive update.
[0101] Specifically, according to the state evaluation results, the control strategy is divided into five operation modes: normal mode, rapid temperature adjustment mode, rapid humidity adjustment mode, energy-saving mode, and emergency mode, to obtain the operation mode set. The normal mode is applicable to the situation where the temperature and humidity state in the carriage belongs to the positive domain. At this time, the temperature and humidity of all points in the carriage are within the target range, and the control objective is to maintain the current temperature and humidity state and minimize energy consumption. The rapid temperature adjustment mode is applicable to the situation where the temperature in the carriage deviates from the target range while the humidity is within the target range, and the control objective is to adjust the temperature to the target range at the fastest speed. The rapid humidity adjustment mode is applicable to the situation where the humidity in the carriage deviates from the target range while the temperature is within the target range, and the control objective is to adjust the humidity to the target range at the fastest speed. The energy-saving mode is applicable to the situation of vehicle parking or good temperature and humidity state in the carriage and no significant change is predicted in a short time, and the control objective is to minimize the system energy consumption. The emergency mode is applicable to the situation where the temperature and humidity in the carriage seriously deviate from the target range or are predicted to seriously deviate in a short time, and the control objective is to adjust the temperature and humidity to the safe range in the shortest time. The mode selection process is a comprehensive judgment based on the state evaluation results of the three-way decision model, the temperature and humidity uniformity evaluation value, and the temperature and humidity change trend, and the most suitable operation mode is automatically selected using rule matching or decision tree algorithm.
[0102] For each operating mode in the set of operating modes, design the corresponding actuator control strategy matrix, map the temperature and humidity control decisions into specific actuator control parameters, and obtain the actuator control parameter set. The control strategy matrix is a mapping relation table that converts high-level control decisions into low-level actuator control parameters, and different mapping rules are designed for different operating modes. For example, for the normal mode, the control strategy matrix will prioritize minimizing energy consumption, appropriately reduce the compressor power and fan speed, and maintain a low desiccant wheel speed; for the rapid temperature adjustment mode, the control strategy matrix will prioritize adjusting the compressor power and fan speed, and at the same time adjust the circulation damper opening according to needs to reach the target temperature at the fastest speed; for the rapid humidity adjustment mode, the control strategy matrix will prioritize adjusting the desiccant wheel speed and the regeneration heater power to quickly adjust the humidity; for the energy-saving mode, the control strategy matrix will adjust the compressor power curve in advance according to the predicted temperature and humidity change trend, avoid frequent start-stop, and reduce peak power consumption; for the emergency mode, the control strategy matrix will maximize the temperature and humidity adjustment capabilities at the same time, regardless of energy consumption factors, and adjust all actuators to the optimal working state. These mapping relations are based on the control decisions obtained by multi-agent reinforcement learning, but further consider the specific requirements of different operating modes. Adopt a hierarchical control architecture to generate the control objectives of each subsystem based on the set of operating modes and the temperature and humidity control decisions, and obtain the subsystem control objective set. The hierarchical control architecture is a control method that decomposes complex control problems into different levels, with the high level responsible for decision-making and the low level responsible for execution. In vehicle-mounted cold chain temperature and humidity control, the high-level controller generates the control objectives of each subsystem according to the selected operating mode and the control actions output by the agent; the low-level controller is responsible for ensuring that each actuator accurately tracks the control objective. The subsystems include the refrigeration system, the fan system, the damper system, the drying system, etc., and each subsystem is responsible for controlling the corresponding actuator. The control objective of the refrigeration system is the set value of the compressor power; the control objective of the fan system is the set value of the fan speed; the control objective of the damper system is the set value of the circulation damper opening; the control objective of the drying system is the set values of the desiccant wheel speed and the regeneration heater power. These control objectives constitute the subsystem control objective set and serve as the input to the low-level controller.
[0103] Apply the proportional-integral control algorithm to the subsystem control target set to generate actuator control signals, ensuring that each actuator accurately tracks the control target and obtaining control execution instructions. The proportional-integral control algorithm is a classic feedback control algorithm that achieves precise tracking of the target value through the combination of the proportional term and the integral term. The role of the proportional term is to give the corresponding control output according to the current error magnitude, and the role of the integral term is to accumulate historical errors and eliminate static errors. For compressor speed control, the control law of the PI controller is to calculate the deviation between the current speed and the set speed, then multiply the deviation by the proportional coefficient and add the integral value of the deviation multiplied by the integral coefficient to obtain the control output. Similarly, for fan speed, recycle damper opening, desiccant wheel speed, and regeneration heater power, corresponding PI controllers are used for precise control. The parameters of the PI controller are obtained through offline optimization or online adaptive adjustment to meet the control requirements under different working conditions. In this way, the high-level control decisions are accurately converted into specific control signals for each actuator, forming control execution instructions.
[0104] By controlling the execution of instructions, the power of the compressor, the speed of the fan, the opening degree of the circulation air damper, the speed of the desiccant wheel, and the power of the regeneration heater are adjusted to perform temperature and humidity regulation operations, and real-time temperature and humidity change data are obtained. The compressor is the core component of the refrigeration system, and the amount of refrigeration can be controlled by adjusting its power; the fan is responsible for the circulation and distribution of cold air, and the temperature uniformity and adjustment rate are affected by adjusting the speed; the circulation air damper controls the flow direction of the cold air, and the temperature distribution in different areas is affected by adjusting the opening degree; the desiccant wheel adjusts the humidity through the absorption or desorption effect, and the humidity adjustment rate is affected by the speed; the regeneration heater is responsible for the regeneration of the desiccant, and the dehumidification ability of the desiccant is affected by the power. These actuators work together to achieve precise control of the temperature and humidity environment in the compartment. During the execution of the control operation, the sensor network in the compartment continuously collects real-time temperature and humidity data, reflecting the execution effect of the control operation. After preprocessing these real-time data, temperature and humidity change data are formed for subsequent evaluation of the control effect. The real-time temperature and humidity change data are compared with the temperature and humidity change trend, the temperature and humidity deviation value is calculated, and the execution effect of the control strategy is determined as the basis for adaptive update. The temperature and humidity deviation value is the difference between the actual temperature and humidity and the predicted temperature and humidity, reflecting the accuracy and effectiveness of the execution of the control strategy. By calculating the temperature deviation and humidity deviation at each sensor point, the overall deviation distribution is obtained. The temperature deviation calculation method is to subtract the predicted temperature at the corresponding moment from the actual measured temperature; the calculation method of the humidity deviation is similar. These deviation data are statistically analyzed, and indicators such as the mean and variance are calculated to comprehensively evaluate the execution effect of the control strategy. If a significant deviation is detected between the actual temperature and humidity change and the expectation (for example, the temperature deviation exceeds ±0.5°C or the humidity deviation exceeds ±5%RH), it indicates that the execution effect of the control strategy is not ideal and needs to be adjusted. These deviation data and their statistical analysis results serve as an important basis for subsequent adaptive update, providing data support for the optimization of the control decision-making system.
[0105] Taking the cold chain transportation of fresh food as an example, a cold chain vehicle transports a batch of fresh food that needs to be kept in an environment of 0 - 4°C and a humidity of 85 - 95%. The temperature and humidity data obtained through the temperature and humidity monitoring network inside the carriage are evaluated by the three-way decision-making model. It is found that the temperature inside the carriage is within the normal range, but the humidity is on the low side. The humidity at multiple sensor points has dropped to 80 - 83%, and the state evaluation result is the boundary region with a relatively uniform temperature and humidity distribution. The temperature and humidity trend analysis predicts that if no measures are taken, the humidity will continue to drop. Based on this information, the control system automatically selects the rapid humidity adjustment mode and inputs this decision into the actuator control strategy matrix. For the rapid humidity adjustment mode, the parameters generated by the control strategy matrix are set as follows: reduce the rotational speed of the desiccant wheel to the lowest value of 20%, turn off the regeneration heater (power 0%), appropriately reduce the compressor power to 60%, keep the fan rotational speed at 70%, and adjust the damper opening to 45°. These parameters are passed to each low-level controller as the subsystem control target. Through the PI control algorithm, each actuator accurately tracks these target values. For example, the PI controller of the desiccant wheel motor takes the deviation between the current rotational speed and the target rotational speed as the input and generates the corresponding control voltage. After 30 minutes of execution operation, the real-time temperature and humidity data show that the humidity has risen to 86 - 91%, and the temperature remains within the range of 1 - 3.5°C, which is basically consistent with the expected trend and the deviation is within the allowable range. This indicates that the execution effect of the control strategy is good, and there is no need for major adjustment, only a fine adjustment of the rotational speed of the desiccant wheel is required to maintain the current humidity level.
[0106] In a specific embodiment, the process of executing step S106 may specifically include the following steps:
[0107] Deeply mine the collected historical operation data, and use the hierarchical clustering algorithm to divide the historical operation data into multiple typical operation modes according to the similarity to obtain a set of typical operation modes;
[0108] Based on the set of typical operation modes, calculate the key performance indicators, including the temperature adjustment rate, humidity adjustment rate, and energy utilization efficiency, to obtain the mode performance evaluation data;
[0109] According to the mode performance evaluation data, identify the working condition points with significant performance differences, extract the operation rules and potential optimization space of the control system, and obtain the optimization direction;
[0110] Adopt an incremental learning strategy, add the data in the execution effect to the training set, and update the parameters of the temperature and humidity trend analysis model to obtain an updated temperature and humidity trend analysis model;
[0111] Maintain an experience buffer pool during actual operation, store the state-action-reward sequence during the operation process, and when the data volume in the buffer pool reaches the threshold, use the policy distillation technology to update the parameters of the control decision-making system to obtain an optimized control strategy;
[0112] An anomaly detection method based on the Isolation Forest algorithm is adopted to monitor the operating parameters of the vehicle-mounted cold chain system in real time. When an abnormal pattern is detected, the diagnostic process is initiated, and the control strategy is automatically adjusted to obtain an optimized temperature and humidity control system.
[0113] Specifically, the historical operation data collected is deeply mined. The hierarchical clustering algorithm is used to divide the historical operation data into multiple typical operation patterns according to similarity, and a set of typical operation patterns is obtained. The hierarchical clustering algorithm is a method of constructing a clustering hierarchy based on the distance or similarity between data points, which is divided into bottom-up agglomerative clustering and top-down divisive clustering. In the temperature and humidity control of vehicle-mounted cold chains, agglomerative hierarchical clustering is mainly adopted, and its processing steps include: first, each data point is regarded as a separate category, and the distance between categories is calculated; then, the two closest categories are merged, and the distance between the merged category and other categories is recalculated; this process is repeated until the termination condition is met. The distance metric uses the Euclidean distance, and the inter-class distance uses the average linkage method, that is, the distance between two categories is the average of the pairwise distances between all points in the categories. The clustering features include the temperature distribution, humidity distribution, ambient temperature, ambient humidity, vehicle speed, load condition, door opening and closing frequency, etc. in the carriage. Through clustering analysis, the historical operation data is divided into different typical operation patterns, such as "fully loaded operation in a high-temperature environment", "partially loaded in a low-temperature environment", "frequent door opening operation", etc. Each pattern represents a specific combination of operating conditions. Based on the identified set of typical operation patterns, the key performance indicators under each pattern are calculated, including the temperature adjustment rate, humidity adjustment rate, and energy utilization efficiency, to obtain the pattern performance evaluation data. The temperature adjustment rate is defined as the change in the carriage temperature per unit time, and the calculation method is the temperature change divided by the adjustment time; the definition and calculation method of the humidity adjustment rate are similar. The energy utilization efficiency is defined as the refrigeration / dehumidification amount generated per unit energy consumption, and the calculation method is the refrigeration / dehumidification amount divided by the product of the system power and the operation time. For each typical operation pattern, the corresponding data segment is extracted, these key performance indicators are calculated, and statistical analysis is performed to obtain statistics such as the average value, standard deviation, maximum value, and minimum value of each indicator. These statistics together constitute the pattern performance evaluation data, reflecting the control effect and energy utilization of the cold chain system under different operation patterns.
[0114] According to the pattern performance evaluation data, identify the operating points with significant performance differences, extract the operating rules and potential optimization space of the control system, and obtain the optimization direction. The operating points with significant performance differences refer to the working state points where the system performance indicators have large differences from the average level under the same conditions. The identification methods include box plot analysis, Z-score calculation, and outlier detection, etc. By comparing the performance indicators under different operating modes, find the operating modes with the best and worst performance, analyze the differences in conditions and control strategies between them, and extract the operating rules of the control system from them. For example, analyze under what conditions the temperature regulation rate is the fastest or the energy utilization efficiency is the highest, and these rules can guide the optimization of the control strategy. At the same time, compare the performance differences at different time periods under the same operating mode to find the potential optimization space. For example, in some working conditions, the energy consumption is too high but the temperature and humidity control effect is not ideal, indicating that there is room for improvement in the control strategy. Through this analysis, a clear optimization direction is obtained, providing guidance for subsequent model updates and strategy optimization.
[0115] Adopt an incremental learning strategy, add the data in the execution effect to the training set, update the parameters of the temperature and humidity trend analysis model, and obtain the updated temperature and humidity trend analysis model. Incremental learning is a learning method that does not require retraining the entire model but updates the existing model based on new data. In this method, at regular intervals or when the prediction error exceeds a threshold, the newly collected data is added to the training set, and only some parameters of the model are updated, thereby reducing the computational overhead. The processing steps of incremental learning include: first, select the latest execution effect data and perform the same preprocessing as the original training data; then, combine the processed data with a part of the historical data to form a new training batch; finally, use this batch of data to update the model parameters, usually only updating the weights of the last few layers. The update frequency is automatically adjusted according to the prediction error: when the prediction error is continuously monitored to exceed the preset threshold (for example, the temperature error is greater than 0.5°C or the humidity error is greater than 5%RH), the model update is triggered; when the prediction error continues to be lower than the threshold, the update frequency is reduced to reduce the computational burden. Through incremental learning, the temperature and humidity trend analysis model can continuously adapt to new operating data and improve prediction accuracy. Maintain an experience buffer pool during actual operation to store the state-action-reward sequences during the operation. When the data volume in the buffer pool reaches the threshold, use the policy distillation technique to update the parameters of the control decision-making system and obtain an optimized control strategy. The experience buffer pool is a data structure that stores the interaction experiences of the agent. Each experience includes the current state, the executed action, the obtained reward, and the next state. In this method, for each execution of the control decision, the corresponding state-action-reward sequence is stored in the experience buffer pool. When the data volume in the buffer pool reaches the set threshold (for example, 10,000 records), the policy update process is triggered. Policy distillation is a technique for transferring knowledge from one model (teacher model) to another model (student model). In this method, the current policy is combined with the historical optimal policy to generate an improved policy that comprehensively considers historical experience and the characteristics of the new environment. The specific processing steps include: first, extract data from the experience buffer pool and calculate the action distributions of the current policy and the historical optimal policy; then, update the parameters of the current policy with the goal of minimizing the KL divergence between the two distributions; finally, apply the updated policy to actual control. In this way, the control decision-making system can be continuously optimized and gradually approach the optimal control strategy.
[0116] An anomaly detection method based on the Isolation Forest algorithm is adopted to monitor the operating parameters of the vehicle-mounted cold chain system in real time. When an abnormal pattern is detected, the diagnostic process is initiated, and the control strategy is automatically adjusted to obtain an optimized temperature and humidity control system. The Isolation Forest algorithm is an unsupervised anomaly detection method based on a tree structure. Its basic idea is that abnormal data points tend to be isolated more quickly in randomly constructed decision trees. The algorithm processing steps include: constructing multiple isolation trees, and each tree divides data points into different subspaces by randomly selecting features and split points; calculating the average path length required for each data point to be isolated, and the shorter the path, the more likely the data point is an abnormal point. In this method, the operating parameters such as compressor current, evaporator temperature, condenser temperature, and system pressure are monitored in real time, and these parameters are input into the Isolation Forest model to calculate the anomaly score. When the anomaly score exceeds the threshold, the anomaly diagnosis process is triggered. The diagnosis process is based on a preset fault decision tree or an inference system based on a knowledge graph to locate the cause of the anomaly and automatically adjust the control strategy according to the diagnosis result. For example, when a refrigerant leak is detected, the system will reduce the compressor operating frequency and increase the fan speed to minimize the impact on the cold chain quality; at the same time, send a warning message to the maintenance personnel, providing fault diagnosis and handling suggestions.
[0117] Taking the cold-chain transportation of fresh milk as an example, a cold-chain vehicle has been transporting fresh milk for a long time, and in-depth mining is carried out by collecting operation data for half a year. First, the hierarchical clustering algorithm is applied to analyze the historical data, and the data is divided into typical operation modes such as full load in summer, full load in winter, partial load in summer, partial load in winter, and frequent door-opening operations. For each mode, key performance indicators are calculated: for example, in the full-load mode in summer, the average temperature regulation rate is 1.2 °C per hour, the average humidity regulation rate is 3.5% per hour, and the energy utilization efficiency is 2.8 kWh / ton. By comparing the performance data under different modes, it is found that the energy utilization efficiency in the full-load mode in summer is significantly lower than that in other modes, which indicates that there is room for optimization of the control strategy under high-temperature environments and full-load conditions. Further analysis reveals that under such operating conditions, the compressor power often reaches its peak but the duration is short, and the frequent start-stop leads to a reduction in energy efficiency. Based on this finding, the optimization direction is determined: to improve the compressor control strategy under full-load conditions in summer and reduce frequent start-stop. Subsequently, the recent operation data is applied to the incremental learning of the temperature and humidity trend analysis model to update the model parameters and improve the prediction accuracy of temperature and humidity changes under high-temperature conditions in summer. At the same time, an experience buffer pool is maintained during operation. When enough state-action-reward sequences under full-load conditions in summer are collected, the policy distillation technology is applied to update the control decision-making system to generate a new control strategy: pre-adjust the fan speed to delay the temperature rise in the carriage and reduce the start-up frequency of the compressor. While implementing the new strategy, the system operation parameters are continuously monitored through the isolation forest algorithm to ensure that the control strategy can be adjusted in a timely manner under abnormal conditions. For example, when abnormal fluctuations in the compressor current are detected, the operation frequency is immediately reduced to avoid damaging the equipment.
[0118] The method for intelligently controlling the temperature and humidity of a vehicle-mounted cold chain described in the embodiments of the present application has been described above. Next, the system for intelligently controlling the temperature and humidity of a vehicle-mounted cold chain in the embodiments of the present application will be described. Please refer to Figure 2 , an embodiment of the system for intelligently controlling the temperature and humidity of a vehicle-mounted cold chain in the embodiments of the present application includes:
[0119] A processing module, configured to collect the temperature and humidity data inside and outside the carriage through multiple temperature and humidity sensors inside the carriage, and process the temperature and humidity data to obtain preprocessed temperature and humidity data;
[0120] A partitioning module, configured to partition the temperature and humidity state inside the carriage based on the preprocessed temperature and humidity data by using a three-way decision model to obtain a state evaluation result of the temperature and humidity inside the carriage;
[0121] An analysis module, configured to use the preprocessed temperature and humidity data and the state evaluation result to construct a temperature and humidity trend analysis model based on a long short-term memory network and an attention mechanism to obtain the temperature and humidity change trend;
[0122] A building block for constructing a control decision system using multi-agent deep reinforcement learning based on the state evaluation result and the temperature and humidity change trend to obtain a temperature and humidity control decision;
[0123] A control module for classifying the control decision into multiple operation modes based on the temperature and humidity control decision, and executing a control strategy in combination with the state evaluation result and the temperature and humidity change trend to obtain an execution effect;
[0124] An update module for adaptively updating the control decision system according to the execution effect to obtain an optimized temperature and humidity control system.
[0125] Through the collaborative cooperation of the above-mentioned various components, the present invention collects temperature and humidity data through multiple temperature and humidity sensors arranged in the carriage and preprocesses them, constructing a comprehensive and accurate data foundation, avoiding the one-sidedness of data caused by single-point sampling, and ensuring the reliability of subsequent analysis and decision-making. The three-way decision model is used to divide the temperature and humidity states in the carriage, accurately classifying the states into the positive domain, negative domain, and boundary domain, overcoming the limitations of traditional binary judgment, and being able to more carefully identify the critical situations of temperature and humidity states, providing a clear state representation for precise control. The temperature and humidity trend analysis model constructed based on the long short-term memory network and attention mechanism effectively captures the temporal characteristics and long-term dependence relationships of temperature and humidity data, automatically identifies key historical data points through the attention mechanism, improves the prediction accuracy of temperature and humidity changes in complex environments, enables the control system to have foresight, and can anticipate and respond to temperature and humidity anomalies in advance. The control decision system is constructed using multi-agent deep reinforcement learning, decomposing temperature control and humidity control into sub-tasks that work together. Through the cooperation mechanism between agents and the design of the global reward function, the coupling problem in temperature and humidity control is solved, and the collaborative optimization control of temperature and humidity is realized. The control decision is divided into multiple operation modes and the control strategy is executed based on the state evaluation results and temperature and humidity change trends, improving the adaptability of the control system to different working conditions. It can adopt the optimal control strategy for different demand scenarios such as normal, rapid temperature adjustment, rapid humidity adjustment, energy saving, and emergency, improving the control accuracy and response speed while reducing energy consumption. The mechanism for adaptively updating the control decision system enables the system to have the ability of continuous learning and self-optimization. It can mine the optimization space from historical operation data and continuously adjust and improve the control strategy according to the execution effect, improving the robustness and long-term performance of the system. The present invention combines artificial intelligence algorithms with the cold chain physical model for the specific application field of on-vehicle cold chain temperature and humidity control, giving full play to the advantages of deep reinforcement learning in complex decision-making problems, the expertise of long short-term memory networks in temporal data processing, and the ability of attention mechanisms in key information recognition. At the same time, by introducing physical constraints and a multi-mode control architecture, it ensures that the algorithm output conforms to actual physical laws and control requirements, realizes the substantial contribution of algorithm features to the solution, solves the complex scenario control problems difficult to handle by traditional methods, and significantly improves the accuracy, stability, foresight, and energy utilization efficiency of on-vehicle cold chain temperature and humidity control.
[0126] Above Figure 2 The system for intelligent control of on-vehicle cold chain temperature and humidity in the embodiments of the present invention is described in detail from the perspective of modular functional entities. Next, the device for intelligent control of on-vehicle cold chain temperature and humidity in the embodiments of the present invention is described in detail from the perspective of hardware processing.
[0127] Figure 3FIG. 0 is a schematic structural diagram of a device for intelligently controlling the temperature and humidity of a vehicle-mounted cold chain provided by an embodiment of the present invention. The device 300 for intelligently controlling the temperature and humidity of a vehicle-mounted cold chain may vary greatly due to different configurations or performances, and may include one or more processors (central processing units, CPU) 310 (for example, one or more processors) and a memory 320, and one or more storage media 330 for storing application programs 333 or data 332 (for example, one or more mass storage device terminals). Among them, the memory 320 and the storage media 330 may be transient storage or persistent storage. The program stored in the storage media 330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the device 300 for intelligently controlling the temperature and humidity of a vehicle-mounted cold chain. Further, the processor 310 may be configured to communicate with the storage media 330 and execute a series of instruction operations in the storage media 330 on the device 300 for intelligently controlling the temperature and humidity of a vehicle-mounted cold chain to implement the steps of the above method for intelligently controlling the temperature and humidity of a vehicle-mounted cold chain.
[0128] The device 300 for intelligently controlling the temperature and humidity of a vehicle-mounted cold chain may further include one or more power supplies 340, one or more wired or wireless network interfaces 350, one or more input / output interfaces 360, and / or one or more operating systems 331, such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, and so on. Those skilled in the art can understand that Figure 3 The shown structure of the device for intelligently controlling the temperature and humidity of a vehicle-mounted cold chain does not constitute a limitation on the device for intelligently controlling the temperature and humidity of a vehicle-mounted cold chain provided by the present invention, and may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0129] The present invention also provides a computer-readable storage medium. The computer-readable storage medium may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions run on a computer, the computer is caused to execute the steps of the method for intelligently controlling the temperature and humidity of a vehicle-mounted cold chain.
[0130] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described system, system, and unit can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0131] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a device for intelligent control of vehicle cold chain temperature and humidity (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0132] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for intelligent control of vehicle cold chain temperature and humidity, characterized in that, The method includes: Collecting the in - vehicle temperature and humidity data and the out - vehicle temperature and humidity data through multiple temperature and humidity sensors in the vehicle compartment, and processing the temperature and humidity data to obtain pre - processed temperature and humidity data; Based on the pre - processed temperature and humidity data, using a three - way decision - making model to divide the in - vehicle temperature and humidity state, and obtaining the state evaluation result of the in - vehicle temperature and humidity; Using the pre - processed temperature and humidity data and the state evaluation result, constructing a temperature and humidity trend analysis model based on a long short - term memory network and an attention mechanism to obtain the temperature and humidity change trend, including: dividing the pre - processed temperature and humidity data according to a time window, where each time window contains a continuous number of sampling points to form an input sequence and obtain time - series feature data; performing feature normalization on the time - series feature data to convert temperature, humidity, and environmental parameters to a unified dimension to obtain normalized feature data; inputting the normalized feature data into a three - layer long short - term memory network for processing, where the first layer contains 128 neurons, the second layer contains 64 neurons, the third layer contains 32 neurons, and a dropout layer is added between each layer to obtain a time - series feature representation; applying the attention mechanism to weight the time - series feature representation, sorting the importance of historical data points by calculating the weight coefficient to obtain a weighted feature representation; mapping the weighted feature representation to the output space through a fully - connected layer to obtain the predicted values of the temperature and humidity change trends at multiple future time points; adding a vehicle heat transfer model as a physical constraint to the predicted values of the temperature and humidity change trends, and optimizing the model parameters through a loss function composed of the mean - square error loss and a physical model constraint regularization term to obtain a temperature and humidity change trend that conforms to physical laws. The loss function is: ; Among them, represents the total loss function, represents the model parameters, M represents the number of training samples, represents the predicted value of the model for the i-th sample, represents the true value of the i-th sample, representing the square of the Euclidean distance, is the weight coefficient, is the physical model constraint regularization term. The physical model constraint regularization term is based on the thermodynamic equation and considers the factors of heat conduction of the carriage wall, convective heat transfer of the air flow, and heat exchange caused by the opening and closing of the door. It is defined as follows: ; Among them, and respectively represent the temperature and humidity of the j-th node predicted by the model at time t, and respectively represent the external environmental temperature and humidity at time t, represents the door state at time t, represents the vehicle speed at time t, represents the car body heat transfer model function based on the thermodynamic equation, represents the physical model constraint regularization term, K represents the number of predicted time steps, and N represents the number of sensor nodes in the car body; According to the state evaluation result and the temperature and humidity change trend, using multi - agent deep reinforcement learning to construct a control decision - making system to obtain a temperature and humidity control decision; Based on the temperature and humidity control decision, dividing the control decision into multiple operation modes, where the multiple operation modes include five operation modes: normal mode, rapid temperature adjustment mode, rapid humidity adjustment mode, energy - saving mode, and emergency mode, to obtain an operation mode set, and combining the state evaluation result and the temperature and humidity change trend to execute the control strategy to obtain the execution effect; According to the execution effect, adaptively updating the control decision - making system to obtain an optimized temperature and humidity control system.
2. The method for intelligently controlling the temperature and humidity of an in-vehicle cold chain according to claim 1, wherein, The step of collecting the in - vehicle temperature and humidity data and the out - vehicle temperature and humidity data through multiple temperature and humidity sensors in the vehicle compartment, and processing the temperature and humidity data to obtain pre - processed temperature and humidity data includes: Setting four sensor nodes in the top area of the vehicle compartment, six sensor nodes in the middle area, and four sensor nodes in the bottom area to form a three - dimensional monitoring network, collecting the real - time temperature values and humidity values at each point in the vehicle compartment to obtain the in - vehicle temperature and humidity data set; Setting an environmental temperature and humidity sensor outside the vehicle compartment to collect the external environmental temperature, environmental humidity, vehicle speed, and door status to obtain the out - vehicle data set; The local outlier factor algorithm is used to detect outliers in the in-carriage temperature and humidity data set and the external data set of the carriage, identify data with too large a difference from adjacent time points, and obtain anomaly-labeled data; Based on the valid data of adjacent time points, the time series interpolation algorithm is applied to the anomaly-labeled data to complete the missing values, and a complete data set is obtained; The sliding average method is applied to the complete data set for smoothing processing to obtain the temperature and humidity data after noise reduction; The temperature and humidity data after noise reduction are combined with the cargo type information, and the physical and chemical properties of the cargo and the optimal storage temperature and humidity range are recorded to obtain the preprocessed temperature and humidity data for the evaluation of the temperature and humidity state.
3. The method for intelligently controlling the temperature and humidity of an in-vehicle cold chain according to claim 1, characterized in that, Based on the preprocessed temperature and humidity data, a three-way decision-making model is used to divide the in-carriage temperature and humidity state, and the state evaluation result of the in-carriage temperature and humidity is obtained, including: According to the optimal storage conditions of the transported cargo, temperature thresholds and humidity thresholds are set to obtain the temperature and humidity target range; The preprocessed temperature and humidity data are analyzed. When the temperature of all points in the carriage is within the temperature threshold range and the humidity is within the humidity threshold range, the system state is divided into the positive domain to obtain an ideal state identifier; The preprocessed temperature and humidity data are analyzed. When the temperature or humidity of some points in the carriage exceeds the temperature threshold range or the humidity threshold range by a certain value, the system state is divided into the negative domain to obtain an abnormal state identifier; The preprocessed temperature and humidity data are analyzed, and the state that does not belong to the positive domain and the negative domain is divided into the boundary domain to obtain a critical state identifier; Calculate the spatial uniformity index of the in-carriage temperature and humidity, including the standard deviation of the temperature and humidity at each measurement point in the carriage, to obtain the evaluation value of the temperature and humidity distribution uniformity; The ideal state identifier or the abnormal state identifier or the critical state identifier is combined with the temperature and humidity distribution uniformity evaluation value to generate a comprehensive state evaluation result as the input of the control decision.
4. The method for intelligently controlling the temperature and humidity of an in-vehicle cold chain according to claim 1, characterized in that According to the state evaluation result and the temperature and humidity change trend, a multi-agent deep reinforcement learning is used to construct a control decision system to obtain the temperature and humidity control decision, including: The control task is decomposed into two sub-tasks of temperature control and humidity control, and a temperature control agent and a humidity control agent are constructed respectively to form a multi-agent control framework; Define the control system state space, including the current in-carriage temperature and humidity distribution, ambient temperature and humidity, vehicle speed, door state, the state evaluation result and the temperature and humidity change trend, to obtain the agent state representation; Define the action space of the temperature control agent, including the compressor power, the fan speed and the opening degree of the circulation air damper, and the action space of the humidity control agent, including the rotation speed of the desiccant wheel and the power of the regeneration heater, to obtain the set of control actions that the agent can execute; Construct the reward function of the temperature control agent, including the temperature deviation penalty term, the temperature non-uniformity penalty term and the energy consumption penalty term, and the reward function of the humidity control agent, including the humidity deviation penalty term, the humidity non-uniformity penalty term and the energy consumption penalty term, to obtain the agent optimization objective; Design a global reward function, combine the rewards of the temperature control agent, the humidity control agent, and the collaboration reward term, and train it through the proximal policy optimization algorithm to obtain the parameters of the agent policy network; Input the state evaluation result and the temperature and humidity change trend into the trained agent policy network to generate an optimal control action sequence and obtain the temperature and humidity control decision.
5. The method for intelligently controlling the temperature and humidity of an in-vehicle cold chain according to claim 1, characterized in that, Based on the temperature and humidity control decision, divide the control decision into multiple operation modes. Among them, the multiple operation modes include five operation modes: normal mode, rapid temperature adjustment mode, rapid humidity adjustment mode, energy-saving mode, and emergency mode, to obtain an operation mode set, and combine the state evaluation result and the temperature and humidity change trend to execute the control strategy to obtain the execution effect, including: Divide the control strategy into five operation modes: normal mode, rapid temperature adjustment mode, rapid humidity adjustment mode, energy-saving mode, and emergency mode according to the state evaluation result to obtain an operation mode set; For each operation mode in the operation mode set, design a corresponding actuator control strategy matrix to map the temperature and humidity control decision to specific actuator control parameters to obtain an actuator control parameter set; Adopt a hierarchical control architecture to generate the control objectives of each subsystem based on the operation mode set and the temperature and humidity control decision to obtain a subsystem control objective set; Apply the proportional-integral control algorithm to the subsystem control objective set to generate an actuator control signal to ensure that each actuator accurately tracks the control objective to obtain a control execution instruction; Adjust the compressor power, fan speed, circulating air damper opening, desiccant wheel speed, and regeneration heater power through the control execution instruction to perform the temperature and humidity adjustment operation to obtain the real-time temperature and humidity change data; Compare the real-time temperature and humidity change data with the temperature and humidity change trend, calculate the temperature and humidity deviation value, and determine the execution effect of the control strategy as the basis for adaptive update.
6. The method for intelligently controlling the temperature and humidity of a vehicle-mounted cold chain according to claim 1, characterized in that, According to the execution effect, perform adaptive update on the control decision system to obtain an optimized temperature and humidity control system, including: Deeply mine the collected historical operation data, and use the hierarchical clustering algorithm to divide the historical operation data into multiple typical operation modes according to similarity to obtain a typical operation mode set; Based on the typical operation mode set, calculate the key performance indicators, including the temperature adjustment rate, humidity adjustment rate, and energy utilization efficiency, to obtain the mode performance evaluation data; According to the mode performance evaluation data, identify the working condition points with significant performance differences, extract the operation rules and potential optimization space of the control system to obtain the optimization direction; Adopt an incremental learning strategy to add the data in the execution effect to the training set and update the parameters of the temperature and humidity trend analysis model to obtain an updated temperature and humidity trend analysis model; Maintain an experience buffer pool during actual operation to store the state-action-reward sequence during the operation process. When the data volume in the buffer pool reaches the threshold, use the policy distillation technology to update the parameters of the control decision system to obtain an optimized control strategy; An anomaly detection method based on the Isolation Forest algorithm is adopted to monitor the operating parameters of the vehicle-mounted cold chain system in real time. When an abnormal pattern is detected, the diagnostic process is started, and the control strategy is automatically adjusted to obtain an optimized temperature and humidity control system.
7. A system for intelligent control of vehicle cold chain temperature and humidity, characterized in that, For implementing the method for intelligently controlling the temperature and humidity of a vehicle-mounted cold chain as described in any one of claims 1-6, the system for intelligently controlling the temperature and humidity of a vehicle-mounted cold chain includes: A processing module for collecting the temperature and humidity data inside and outside the carriage through multiple temperature and humidity sensors inside the carriage and processing the temperature and humidity data to obtain preprocessed temperature and humidity data; A partitioning module for partitioning the temperature and humidity state inside the carriage based on the preprocessed temperature and humidity data by using a three-way decision-making model to obtain a state evaluation result of the temperature and humidity inside the carriage; An analysis module for using the preprocessed temperature and humidity data and the state evaluation result to construct a temperature and humidity trend analysis model based on a long short-term memory network and an attention mechanism to obtain the temperature and humidity change trend; A construction module for constructing a control decision-making system by using multi-agent deep reinforcement learning according to the state evaluation result and the temperature and humidity change trend to obtain a temperature and humidity control decision; A control module for dividing the control decision into multiple operation modes based on the temperature and humidity control decision, where the multiple operation modes include five operation modes: normal mode, rapid temperature adjustment mode, rapid humidity adjustment mode, energy-saving mode, and emergency mode, to obtain an operation mode set, and combining the state evaluation result and the temperature and humidity change trend to execute a control strategy to obtain an execution effect; An update module for adaptively updating the control decision-making system according to the execution effect to obtain an optimized temperature and humidity control system.
8. An apparatus for intelligently controlling the temperature and humidity of a vehicle-mounted cold chain, characterized in that, It includes a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the computer program, it implements the method for intelligently controlling the temperature and humidity of a vehicle-mounted cold chain as described in any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is run by the processor, the processor is caused to execute the method for intelligently controlling the temperature and humidity of a vehicle-mounted cold chain as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Power supply service early warning method based on three-way decision theory and LSTM neural network
CN115271993A
Mushroom house temperature prediction model training method, mushroom house temperature prediction method and mushroom house temperature prediction device
CN117972433A
Temperature regulation and control method and system for hydrogen power cold chain transport vehicle based on Beidou
CN118700787A
Charging pile cable monitoring method and system and charging pile cable
CN119005006A
Crop planting management method and device, electronic equipment and storage medium
CN119784527A