A method for regulating irradiation warehouse environment based on multi-modal data fusion
By using multimodal data fusion and deep reinforcement learning networks, the system predicts the state of the storage environment and optimizes control parameters, solving the problem of insufficient precision in storage environment regulation. This enables personalized and adaptive environmental management, improving regulation accuracy and stability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HUNAN INST OF INFORMATION TECH
- Filing Date
- 2026-01-16
- Publication Date
- 2026-04-21
AI Technical Summary
Existing methods for controlling the warehousing environment rely on fixed thresholds, resulting in insufficient control precision, making it difficult to meet the refined needs of modern processing enterprises, and posing a risk of environmental fluctuations.
By fusing multimodal data, voice commands and image data are used to determine warehouse environment parameters. Combined with Kalman filtering algorithm and deep reinforcement learning network, the future environmental state is predicted and control parameters are optimized to achieve personalized and adaptive environmental management.
It improves the precision and stability of warehousing environment control, avoids environmental fluctuations, achieves precise management of irradiated items, and meets personalized needs.
Smart Images

Figure CN121544179B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of warehouse control technology, and in particular relates to a control method for irradiated warehouse environment based on multimodal data fusion. Background Technology
[0002] In the irradiation processing industry (such as food, pharmaceuticals, and medical devices), items awaiting irradiation typically need to be temporarily stored in warehouses while awaiting irradiation treatment. Due to factors such as the operation and scheduling of irradiation processing equipment and production plans, the waiting time for irradiated items can be lengthy and unpredictable. During this period, improper control of the storage environment (such as temperature and humidity) may lead to spoilage, decay, or performance degradation of the items, affecting the quality and safety of the final product. Currently, warehouse environment control mainly relies on fixed threshold adjustments, which suffers from insufficient precision and fails to meet the refined warehousing needs of modern processing enterprises. Summary of the Invention
[0003] This application provides a method for controlling the irradiated storage environment based on multimodal data fusion, which can solve the problem of insufficient control accuracy in the storage environment.
[0004] This application provides a method for controlling the irradiated storage environment based on multimodal data fusion, including:
[0005] The system acquires user-input voice commands and image data of the items to be irradiated; the voice commands are used to describe the required storage environment information for the items to be irradiated.
[0006] Based on voice commands and image data, determine the storage environment parameter values of the items to be irradiated;
[0007] By fusing and estimating the environmental data of the storage area where the irradiated items are located, the current environmental state vector, the current equipment state vector, the historical environmental state vector, and the historical equipment state vector of the storage area are obtained.
[0008] Based on the storage environment parameter values, historical environmental state vectors, and historical equipment state vectors, environmental state prediction is performed to obtain the predicted environmental state value of the storage area in the future period, as well as the comprehensive risk assessment value of the predicted environmental state value.
[0009] The current environmental state vector, the current equipment state vector, the predicted environmental state value, the comprehensive risk assessment value, and the warehouse environment parameter value are used as the current state space. The current state space is then input into the trained deep reinforcement learning network for processing to obtain the current action. The current action includes the control parameters of the environmental control equipment in the warehouse area.
[0010] Control the environmental control equipment in the storage area based on the current action.
[0011] Optionally, based on voice commands and image data, the storage environment parameters of the items to be irradiated are determined, including:
[0012] The storage environment parameter values for the items to be irradiated are calculated using the following formula. :
[0013] ;
[0014] in, This indicates the preset voice confidence level. This represents the environmental parameter values identified from the voice commands. This indicates the preset image confidence level. This indicates the environmental parameter values determined from the knowledge base of environmental requirements for stored goods based on image data.
[0015] Optionally, the environmental data of the storage area where the irradiated item is located is fused and estimated to obtain the current environmental state vector, current equipment state vector, historical environmental state vector, and historical equipment state vector of the storage area, including:
[0016] The Kalman filter algorithm is used to fuse and estimate the environmental data of the storage area where the irradiated items are located.
[0017] Based on the environmental estimates obtained by the Kalman filter algorithm fusion estimation, the current environmental state vector, current equipment state vector, historical environmental state vector, and historical equipment state vector of the storage area are obtained.
[0018] Optionally, the current environmental state vector includes: the current temperature estimate of the storage area, the current humidity estimate of the storage area, and the current light intensity estimate of the storage area;
[0019] The current equipment state vector includes: the current estimated power of the air conditioners in the storage area, the current estimated power of the dehumidifiers in the storage area, and the current estimated speed of the exhaust fans in the storage area;
[0020] The historical environmental state vector includes: temperature estimates, humidity estimates, and light intensity estimates for the storage area at multiple historical times.
[0021] The historical equipment state vector includes: the estimated power of air conditioners in the storage area at multiple historical moments, the estimated power of dehumidifiers in the storage area at multiple historical moments, and the estimated speed of exhaust fans in the storage area at multiple historical moments.
[0022] Optionally, environmental status prediction is performed based on warehousing environment parameter values, historical environmental state vectors, and historical equipment state vectors to obtain predicted environmental status values for the warehousing area in future periods, as well as a comprehensive risk assessment value based on the predicted environmental status values, including:
[0023] The storage environment parameter values, historical environment state vectors, and historical equipment state vectors are input into the evaluation model for processing to obtain the predicted environmental state values of the storage area in future time periods. The evaluation model consists of an LSTM encoder, an attention module, an LSTM decoder, and an output layer connected in sequence.
[0024] Based on the hidden state of the LSTM decoder when generating the last prediction point, a comprehensive risk assessment value for the predicted environmental state is obtained.
[0025] Optionally, based on the hidden state of the LSTM decoder when generating the last prediction point, a comprehensive risk assessment value for the predicted environmental state is obtained, including:
[0026] The hidden state of the LSTM decoder when generating the last prediction point is calculated using a fully connected layer, and the calculation result is obtained.
[0027] The calculation results are normalized using the Sigmoid activation function to obtain the comprehensive risk assessment value of the environmental state prediction.
[0028] Optionally, the deep reinforcement learning network is obtained through deep reinforcement learning training. The state space of deep reinforcement learning includes environmental state vector, device state vector, environmental state prediction value, comprehensive risk assessment value, and warehouse environment parameter value.
[0029] The actions learned through deep reinforcement learning are the control parameters of the environmental equipment in the storage area where the items to be irradiated are located.
[0030] Optional environmental control equipment includes air conditioners, dehumidifiers, and exhaust fans.
[0031] The above-mentioned solution in this application has the following beneficial effects:
[0032] In the embodiments of this application, by analyzing the historical environmental state and equipment state of the storage area where the irradiated item is located, as well as the storage environment parameter values of the irradiated item, the evolution trend of the environmental state and its comprehensive risk assessment value in the future time period are predicted, thereby effectively solving the lag problem of traditional control. Based on this, the current environmental state, current equipment state, environmental state evolution trend, and comprehensive risk assessment value of the storage area are used as the current state space. This current state space integrates the current state and future knowledge, enabling the processing of complex and uncertain information using a deep reinforcement learning network. This ensures that the output control parameters of the environmental control equipment are stable and forward-looking, effectively avoiding environmental fluctuations.
[0033] Meanwhile, since the control method of this application is based on the storage environment information required by the irradiated items, the environment of the storage area and the status of the equipment, this application can achieve personalized and adaptive precise management according to the needs of the irradiated items, effectively avoid the shortcomings of fixed threshold control, and significantly improve the control accuracy of the storage environment.
[0034] Other beneficial effects of this application will be described in detail in the following detailed description section. Attached Figure Description
[0035] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0036] Figure 1 A flowchart illustrating a method for controlling an irradiated storage environment based on multimodal data fusion, provided as an embodiment of this application. Detailed Implementation
[0037] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0038] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0039] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0040] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."
[0041] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0042] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0043] To address the issue of insufficient precision in current warehousing environment control, this application provides a method for controlling irradiated warehousing environments based on multimodal data fusion. This method analyzes the historical environmental state, historical equipment state, and environmental parameter values of the irradiated item's storage area to predict the future evolution trend of the environmental state and its comprehensive risk assessment value, effectively solving the lag problem of traditional control methods. Furthermore, the current environmental state, current equipment state, environmental state evolution trend, and comprehensive risk assessment value of the storage area are used as the current state space. This current state space integrates the current state with future knowledge, enabling the processing of complex and uncertain information using a deep reinforcement learning network. This results in stable and forward-looking control parameters for the output environmental control equipment, effectively avoiding environmental fluctuations.
[0044] Meanwhile, since the control method of this application is based on the storage environment information required by the irradiated items, the environment of the storage area and the status of the equipment, this application can achieve personalized and adaptive precise management according to the needs of the irradiated items, effectively avoid the shortcomings of fixed threshold control, and significantly improve the control accuracy of the storage environment.
[0045] The following describes the method for controlling the irradiated storage environment based on multimodal data fusion provided in this application by way of specific embodiments.
[0046] like Figure 1 As shown in the embodiments of this application, the method for controlling the irradiated storage environment based on multimodal data fusion includes the following steps:
[0047] Step 11: Obtain the voice command input by the user and obtain the image data of the item to be irradiated. The voice command is used to describe the storage environment information required for the item to be irradiated.
[0048] The items to be irradiated mentioned above are items that require irradiation treatment, such as food, medicine, and medical devices.
[0049] In some embodiments of this application, when items to be irradiated need to be stored in a warehouse awaiting irradiation, operators can record voice commands via the microphone of a smart terminal, which then transmits the voice commands to the system used to execute the control method of this application. The voice commands include the required storage environment information for the items to be irradiated, such as temperature, humidity, and light intensity. For example, the items to be irradiated need to be stored at 20±2℃, in a dry environment, and away from direct sunlight.
[0050] Furthermore, when items to be irradiated need to be stored in a warehouse awaiting irradiation, operators can acquire image data of the items using an image acquisition device, which then transmits the image data to the system used to execute the control method of this application. This image data provides visual supplementary information for the system to obtain user requirements, especially correcting incomplete or ambiguous verbal descriptions.
[0051] Step 12: Determine the storage environment parameter values of the items to be irradiated based on voice commands and image data.
[0052] In some embodiments of this application, after receiving the voice command, the system performs the following processing:
[0053] (1) Speech Recognition Processing: An end-to-end Transformer-based learning model is used to convert speech signals into text. This model is trained on a large amount of Chinese speech data and can effectively overcome noise interference in the warehouse environment. Its core is to solve for the most probable text sequence T:
[0054] ;
[0055] in, It is an acoustic model. It is the input speech feature sequence. It is a sequence of candidate texts, that is, a given text. , generate speech The probability is given by this formula, which is used to find the maximum value parameter.
[0056] (2) Semantic understanding and information extraction: Natural language processing is performed on the identified text. First, entity recognition is performed using a speech model to extract key entities from the text. Then, entity types are defined, including item categories, temperature requirements, humidity requirements, lighting requirements, and special requirements. Finally, relation extraction is performed: a dependency syntax tree is constructed to determine the modification relationships between the extracted entities, such as binding "temperature" with "20±2℃".
[0057] In some embodiments of this application, after acquiring the image data, the system performs the following processing:
[0058] (1) The first step is item recognition and classification. A pre-trained convolutional neural network (CNN) is used to identify and classify items in the image. The model is fine-tuned on a dataset containing a large number of warehouse items such as medicines and food, and outputs the category labels of the irradiated items;
[0059] (2) Next, the item attribute analysis is carried out. In order to further refine the requirements, a deeper attribute analysis is performed on the image.
[0060] (3) Finally, the packaging materials are identified. Through convolutional neural networks, the outer packaging of the items (such as cardboard boxes, plastic seals, and metal cans) is analyzed. Different packaging materials have different sensitivities to humidity.
[0061] In some embodiments of this application, after extracting relevant key information (e.g., temperature, humidity, light intensity) from voice commands, if the value corresponding to the extracted key information is a fixed value, then the fixed value is directly used as the environmental parameter value identified from the voice command; if the value corresponding to the extracted key information is a range, then the range of values corresponding to the extracted key information can be used as the environmental parameter value identified from the voice command. After extracting relevant key information (e.g., item category) from image data, the corresponding environmental parameter value can be obtained based on the storage item environmental requirements knowledge base shown in Table 1 (this storage item environmental requirements knowledge base can be set by relevant operators based on experience, and it contains information such as temperature, humidity, light intensity, and remarks for various items that need to be irradiated; Table 1 here only provides illustrative information for medicine A and cooked food B).
[0062] Specifically, if the environmental conditions (temperature, humidity, light intensity) of a certain type of item recorded in the knowledge base of storage item environmental requirements are a fixed value, then this fixed value is directly used as the environmental parameter value determined from the knowledge base of storage item environmental requirements based on the image data. If the environmental conditions (temperature, humidity, light intensity) of a certain type of item recorded in the knowledge base of storage item environmental requirements are a range, then this range can be used as the environmental parameter value determined from the knowledge base of storage item environmental requirements based on the image data.
[0063] In Table 1, ℃ stands for degrees Celsius, which is a unit of temperature, and %RH stands for relative humidity, which is a unit of relative humidity that measures the water vapor content in the air.
[0064] Table 1 Knowledge Base of Environmental Requirements for Stored Goods
[0065]
[0066] In some embodiments of this application, the storage environment parameter values of the items to be irradiated can be calculated using the following formula. :
[0067] ;
[0068] in, This indicates the preset voice confidence level. This represents the environmental parameter values identified from the voice commands. This indicates the preset image confidence level. This refers to the environmental parameter values determined from the knowledge base of storage item environmental requirements based on image data. These environmental parameter values can be temperature, humidity, and light intensity. When the environmental parameter value is temperature, the storage temperature parameter value is calculated using the above formula; when the environmental parameter value is humidity, the storage humidity parameter value is calculated using the above formula; and when the environmental parameter value is light intensity, the storage light intensity parameter value is calculated using the above formula. In other words, in this embodiment, the storage environmental parameter values ultimately obtained based on voice commands and image data... This includes storage temperature parameters, storage humidity parameters, and storage light intensity parameters.
[0069] It should be noted that in practical applications, the corresponding storage environment parameters can be calculated separately for the three environmental conditions: temperature, humidity, and light intensity. Specifically, when calculating the storage temperature parameter, all values in the above formula correspond to the actual temperature; when calculating the storage humidity parameter, all values in the above formula correspond to the actual humidity; and when calculating the storage light intensity parameter, all values in the above formula correspond to the actual light intensity.
[0070] Regarding light intensity, if the voice or image indicates strict light avoidance, the corresponding environmental parameter value (i.e., light intensity value) is 0 to 50 lux; if the voice or image indicates light avoidance, the corresponding environmental parameter value (i.e., light intensity value) is 0 to 100 lux; if the voice or image indicates weak light, the corresponding environmental parameter value (i.e., light intensity value) is 100 to 300 lux; if the voice or image indicates normal lighting, the corresponding environmental parameter value (i.e., light intensity value) is 300 to 500 lux; and if the voice or image indicates no requirements, the corresponding environmental parameter value (i.e., light intensity value) is 0 to 500 lux.
[0071] It is understood that the above environmental parameter values (such as temperature, humidity, and light intensity) can all be within a range, and the corresponding calculated storage environment parameter values... It can also be a range.
[0072] It depends on speech clarity; the higher the speech clarity, the better. The larger the value, for example, when the speech clarity is high. The value is 0.9; It depends on the confidence level of the image classification and the reliability rating of the record in the knowledge base. For example, when the confidence level of the image classification is average and the reliability rating is medium, The value is 0.6. This is understandable. and The value can also be preset.
[0073] For example, when an operator puts a box of items into the warehouse and enters a voice message: "This is medicine, it needs to be refrigerated, 2 to 8 degrees Celsius," the system's camera simultaneously captures an image of the item's packaging box.
[0074] At this point, through speech recognition processing, the words "medicine" and the temperature range [2, 8] were successfully extracted. The median value of 5 can then be used as... Due to its clear speech and high recognition confidence, it provides comprehensive... =0.9. Image processing identified the "medical" label on the outer packaging, but failed to accurately identify the specific drug, classifying it as "general medicine" with moderate confidence. A recommended temperature was obtained by searching the "general medicine" knowledge base for warehousing environmental requirements. =5, and this record is the default value, so its reliability is moderate. Overall, it is given... =0.6.
[0075] Storage temperature parameter values The calculation is as follows: In this case, the two information sources give the same value. Regardless of the confidence level, the final result is still 5. The fusion plays a role in verification and reinforcement.
[0076] For humidity and light intensity, refer to the formulas above. If not mentioned in the voice, the system will directly adopt the data from the knowledge base of environmental requirements for stored goods.
[0077] Final warehouse environment parameter values =[Storage temperature parameter value] Storage temperature parameter values Warehouse light intensity parameter values ].
[0078] Step 13: Perform fusion estimation on the environmental data of the storage area where the irradiated item is located to obtain the current environmental state vector, current equipment state vector, historical environmental state vector, and historical equipment state vector of the storage area.
[0079] In some embodiments of this application, the Kalman filter algorithm can be used to fuse and estimate the environmental data of the storage area where the irradiated item is located; then, based on the environmental estimate obtained by the Kalman filter algorithm fusion estimation, the current environmental state vector, the current equipment state vector, the historical environmental state vector, and the historical equipment state vector of the storage area are obtained.
[0080] Specifically, various sensors within the storage area (such as temperature sensors, humidity sensors, light intensity sensors, and speed sensors) can collect data over a preset time period (usually including the current time and the time before the current time). A historic moment, The values can be set according to actual conditions. The system collects the temperature, humidity, light intensity, and exhaust fan speed of the storage area, and obtains the power of the air conditioner and dehumidifier within a preset time period through the control system of the air conditioner and dehumidifier within the storage area. In practical applications, the device sensing layer can collect the temperature, humidity, light intensity, and exhaust fan speed of the storage area within a preset time period, as well as the power of the air conditioner and dehumidifier within a preset time period.
[0081] Based on this environmental data, the Kalman filter algorithm can be used to fuse and estimate the temperature, humidity, light intensity, air conditioning power, dehumidifier power, and exhaust fan speed in the storage area within a preset time period, yielding environmental estimates. These estimates include the estimated temperature, humidity, and light intensity of the storage area within the preset time period, as well as the estimated power of the air conditioning, dehumidifier, and exhaust fan speeds within the storage area within the preset time period.
[0082] Specifically, the aforementioned current environmental state vector includes: the current temperature estimate of the storage area (i.e., the temperature estimate of the storage area at the current moment), the current humidity estimate of the storage area (i.e., the humidity estimate of the storage area at the current moment), and the current light intensity estimate of the storage area (i.e., the light intensity estimate of the storage area at the current moment).
[0083] The aforementioned current equipment state vector includes: the current estimated power of the air conditioner in the storage area (i.e., the estimated power of the air conditioner in the storage area at the current moment), the current estimated power of the dehumidifier in the storage area (i.e., the estimated power of the dehumidifier in the storage area at the current moment), and the current estimated speed of the exhaust fan in the storage area (i.e., the estimated speed of the exhaust fan in the storage area at the current moment).
[0084] The historical environment state vector includes: the storage area at multiple historical moments (i.e., the aforementioned Temperature estimates at multiple historical moments, and the storage area at multiple historical moments (i.e., the aforementioned) Humidity estimates at multiple historical moments and storage area at multiple historical moments (i.e., the aforementioned) The estimated light intensity at each historical moment.
[0085] Historical equipment state vectors include: air conditioners within the warehouse area at multiple historical moments (i.e., the aforementioned...). The estimated power of dehumidifiers in the storage area at multiple historical moments (i.e., the aforementioned) The estimated power of the exhaust fans in the storage area at multiple historical moments (i.e., the aforementioned) The estimated rotational speed (at a historical moment).
[0086] It should be noted that the Kalman filter algorithm is a commonly used data estimation method, therefore, its principle will not be elaborated on here, but only briefly explained as follows.
[0087] Before fusion estimation, since real-time acquired data typically contains noise, missing values, and asynchronous features, after the device perception layer transmits the acquired data to the edge nodes through the transport layer (wired or wireless transmission), the edge nodes perform data preprocessing for each type of data, as follows:
[0088] (1) Outlier detection: A statistical method based on standard scores (Z-score) is used to quickly identify and remove obviously abnormal readings.
[0089] ;
[0090] In the above formula, These are the standardized values. For the observed values, For sensors in recent times.
[0091] The mean of the data, This represents the standard deviation of recent sensor data.
[0092] (2) Missing value handling: For short-term missing values, time series linear interpolation is used to fill in the missing values; for long-term missing values, the sensor is marked as faulty and the redundant sensor switching mechanism is activated.
[0093] (3) Spatiotemporal alignment: Since each type of sensor is deployed in different locations, its data represents the state of different spatial points. A three-dimensional spatial model is established for the warehouse, and each sensor is bound to its spatial coordinates. Through spatial interpolation algorithms, the measurements of discrete points are fused into a continuous field distribution for the entire space.
[0094] ;
[0095] In the above formula, Represents any point in space The estimated value, Indicates sensor At point The measured value, It is the Euclidean distance between two points. It is the attenuation coefficient. It refers to the number of sensors.
[0096] It should be noted that the above data preprocessing is based on the data output by the sensor. The data preprocessing for air conditioner power and dehumidifier power also needs to be performed in accordance with the above process. To avoid too much repetition, the preprocessing process for these two will not be described in detail here.
[0097] After the above data preprocessing, the data still needs to be deeply fused to generate a more reliable environmental state estimate. The specific process is as follows:
[0098] (1) Fusion Model: Multi-sensor data (including temperature, humidity, light intensity, exhaust fan speed, air conditioner power, and dehumidifier power) are fused using a Kalman filter, combining "model-based internal estimation" and "sensor-based observations" to obtain a more accurate and reliable optimal estimate. A mathematical model is defined based on the storage environment, including state vectors and observation vectors.
[0099] (a) State vector It includes the internal state of the storage environment, namely static parameters and dynamic rate of change.
[0100] ;
[0101] in: , , These are the temperature estimate, humidity estimate, and light intensity estimate for a certain area of the warehouse at the previous moment (e.g., current temperature estimate = previous temperature + rate of change). , , These are the rates of change for temperature (°C / min), humidity (%RH / min), and light intensity, respectively. Because environmental changes are continuous, rates of change need to be added to improve prediction accuracy. For example, if the temperature was 10°C at the previous moment and was rising at a rate of 0.1°C / min, without a rate of change in the state variables, the model would predict the next temperature to be 20°C. Conversely, without a rate of change, the model could predict 20.1°C.
[0102] The state space can be represented as: Describe how the state changes from the previous moment. Natural evolution to the present moment .in, For the previous moment state, For a moment state, It is the state transition matrix. It is process noise, which represents the inaccuracy of the model.
[0103] (b) Observation vector , including at time The actual readings of all sensors. and These represent temperature sensor 1 and humidity sensor 1, respectively. and These represent temperature sensor 2 and humidity sensor 2, respectively. and Representing light intensity sensor 1 and light intensity sensor 2 respectively, and so on, we obtain the multidimensional observation vector:
[0104] ;
[0105] The observation space can be represented as: The purpose is to establish the internal state of the system. and external observation The connection between them. The observation matrix represents how the state vector is mapped to the observation vector. This is observation noise, representing the measurement error of the sensor.
[0106] (2) The Kalman filter is used for fusion, as follows:
[0107] (a) Predicting state and predicting uncertainty, i.e., predicting the state at the current moment based on the optimal estimate of the previous moment.
[0108] The predicted state is represented as: ,in, For prior state estimation, for ( for The optimal state estimate (posterior estimate) of the state variables at time 1. Let be the state transition matrix.
[0109] Prediction uncertainty is expressed as: ,in, It is the error covariance of the prior estimate, representing the uncertainty of the prediction. Let be the error covariance matrix of the state estimate at the previous time step, representing the uncertainty of the estimate. Represents the state transition matrix Mapping the uncertainty of the previous moment to the current moment. Let be the process noise covariance matrix.
[0110] (b) Calculate the Kalman gain: The Kalman gain value determines the intelligent trade-off between "model-based predictions" and "actual sensor observations":
[0111] ;
[0112] Kalman gain It's a trade-off factor that determines whether to trust the predictive model output or rely more on sensor observations. If sensor noise... Even with very small uncertainties (where the observations are reliable), the gain will increase, meaning there is greater confidence in the observations. If the uncertainty in the model predictions... If the gain is small (the model is very reliable), it means there is more confidence in the prediction.
[0113] (c) Data fusion processing: This step fuses the actual sensor observations (real-time data) with the predicted values to obtain the optimal estimate. .
[0114] ;
[0115] in, It is a predicted value (based on the state and model prediction of the previous time step). These are observations (sensor readings). This maps the predicted values to values in the observation space. This represents the difference between the observed value and the predicted expected value. Kalman gain. By incorporating a portion of this difference into the final state estimate, the fusion of multi-sensor data is achieved.
[0116] (d) Update uncertainty:
[0117] ;
[0118] The observation matrix maps the state space to the observation space. For Kalman gain, It is the error covariance of the prior estimate. Given an identity matrix with 1s on the diagonal and 0s elsewhere, after the update, the intelligent control module obtains this value, and the uncertainty of the state estimation is... It will decrease, if A sudden increase in value can indicate a potential problem in the sensor network, triggering an early warning. The Kalman filter, through a "prediction-update" loop, continuously corrects the state estimate, making it closer to the true value and effectively smoothing out random fluctuations.
[0119] (3) Data Fusion Output: After data fusion, the edge nodes will output unified and high-quality environmental state vectors and device state vectors. The environmental state vector here includes the current environmental state vector. Current device state vector , A historical environment state vector (denoted as) ), A historical device state vector (denoted as) ).
[0120] Among them, time Let be the current moment, and be the historical environment state vector. for The environmental state vector at time (the time before the current time, also known as the first historical time), also known as the first historical environmental state vector, includes the storage area in... The estimated values for temperature, humidity, and light intensity at any given time. And so on. for The environment state vector at time t, also known as the first time step, is the state vector of the environment at time t. A historical environmental state vector, which includes the storage area in The estimated values for temperature, humidity, and light intensity at any given time.
[0121] Historical device state vector for The device state vector at time (the time before the current time, also known as the first historical time), also known as the first historical device state vector, includes the air conditioning in the storage area. Estimated power at any time, dehumidifiers in the storage area Estimated power at any time and exhaust fans in the storage area The estimated rotational speed at any given time. And so on. for The device state vector at time t, also known as the t-th time vector, is... A historical device state vector, which includes the air conditioners in the storage area... Estimated power at any time, dehumidifiers in the storage area Estimated power at any time and exhaust fans in the storage area Estimated rotational speed at any given time.
[0122] Step 14: Based on the storage environment parameter values, historical environment state vectors, and historical equipment state vectors, perform environmental state prediction to obtain the predicted environmental state value of the storage area in the future period, as well as the comprehensive risk assessment value of the predicted environmental state value.
[0123] In some embodiments of this application, the cloud server can predict the environmental state based on the historical environmental state vector and historical device state vector transmitted by the edge node, as well as the calculated warehouse environment parameter values.
[0124] In some embodiments of this application, step 14 is specifically implemented as follows: steps 14.1 and 14.2:
[0125] Step 14.1: Input the storage environment parameter values, historical environmental state vectors and historical equipment state vectors into the evaluation model for processing to obtain the predicted environmental state values of the storage area in the future period. The predicted environmental state values include the temperature, humidity and light intensity values of the storage area in the future period.
[0126] The evaluation model described above includes a Long Short-Term Memory (LSTM) encoder, an attention module, an LSTM decoder, and an output layer connected in sequence. For example, the LSTM encoder can be a single-layer LSTM encoder, the LSTM decoder can be a single-layer LSTM decoder, and the output layer can be a fully connected layer.
[0127] It should be noted that the LSTM encoder, attention module, and LSTM decoder are respectively traditional LSTM encoders, traditional attention modules, and traditional LSTM decoders. Therefore, their working principles will not be elaborated upon here, but only illustrated by the following examples.
[0128] In this LSTM encoder, the first few LSTM layers are responsible for capturing short-term local dependencies and subtle patterns in the sequence, including short-term noise patterns in sensor readings. The later LSTM layers receive and process the features passed from the earlier layers, fusing and upscaling them to form high-level, global situational features. In this way, the encoder ultimately encodes the entire input sequence into an information-rich hidden state sequence. Each of them These all represent the deep feature representation of the input sequence at the corresponding time step. The length of the input sequence.
[0129] To address the problem of forgetting information in long sequences, an attention mechanism is introduced between the encoder and decoder. Its role is to handle the hidden states of the encoder output at all time steps. Dynamic importance assessment and selection are performed. When the decoder generates each prediction point, it assigns a hidden state to the most relevant historical moment in the input sequence. Higher weight Adaptive feature enhancement is performed to address the shortcomings of traditional models that treat all historical data equally.
[0130] ;
[0131] ;
[0132] in, Represents the context vector. It is the hidden state of the decoder at the previous time step. It is an alignment model. This allows the model to focus on historical information that is most relevant to the current prediction. The encoder is reading the first input sequence. The hidden state generated at each time step It is the first Hidden state The corresponding attention weights.
[0133] The initial state of the LSTM decoder is initialized by the final state of the LSTM encoder and the context vector. At each step... The decoder receives the predicted value from the previous step. In addition to the current context vector and condition vector (i.e., the storage environment parameter values mentioned earlier), output the current hidden state. .
[0134] Output layer: Set the weight matrix of the output layer to... The bias vector is .
[0135] ;
[0136] The above formula represents the first time... The hidden state of a time decoder Mapped to predicted values through a fully connected layer. It includes predicted data such as temperature, humidity, and light sensitivity. Based on this formula, the final output environmental state prediction sequence of the model within the future time range of M is obtained. =[ ,..., ], for Predicted environmental conditions at any given time, including the warehouse area. The temperature, humidity, and light intensity values at any given time, and so on. for Predicted environmental conditions at any given time, including the warehouse area. Temperature, humidity, and light intensity values at any given time.
[0137] Understandably, in practical applications, the above evaluation model can also predict the equipment status and obtain the predicted equipment status value for subsequent decision-making.
[0138] It should be noted that before using the above evaluation model for prediction, it can be trained using common model training methods. For example, the Adam optimizer can be used for backpropagation to iteratively optimize the model parameters and complete the training. During training, supervised learning can be performed using historically collected real datasets, with mean squared error loss as the loss function.
[0139] Step 14.2: Based on the hidden state of the LSTM decoder when generating the last prediction point, obtain the comprehensive risk assessment value of the environmental state prediction value.
[0140] In some embodiments of this application, a fully connected layer can be used to calculate the hidden state of the LSTM decoder when generating the last prediction point, obtaining the calculation result; then, the Sigmoid activation function is used to normalize the calculation result to obtain a comprehensive risk assessment value of the environmental state prediction. This comprehensive risk assessment value comprehensively reflects the probability and severity of the environmental state prediction deviating from the item storage requirements in the future time period. The Sigmoid activation function is a basic and commonly used activation function in artificial neural networks.
[0141] Specifically, the comprehensive risk assessment value can be calculated using the following formula. :
[0142] ;
[0143] in, The LSTM decoder generates the last (the...) The hidden state when predicting 1) points, and These are the weights and biases of the fully connected layer used to calculate the overall risk assessment value. and These are learnable parameters. It is the Sigmoid activation function.
[0144] Understandably, to improve the accuracy of regulation, the fully connected layer used to calculate the comprehensive risk assessment value can be trained together with the aforementioned assessment model.
[0145] Step 15: Use the current environmental state vector, current equipment state vector, environmental state prediction value, comprehensive risk assessment value, and warehouse environment parameter value as the current state space, and input the current state space into the trained deep reinforcement learning network for processing to obtain the current action; the current action includes the control parameters of the environmental control equipment in the warehouse area.
[0146] It is understandable that the current actions obtained based on the current state space include the control parameters of the environmental control equipment in the storage area for future periods.
[0147] In some embodiments of this application, the deep reinforcement learning network described above is obtained through deep reinforcement learning training, and the deep reinforcement learning network can be understood as an intelligent agent.
[0148] The state space of the aforementioned deep reinforcement learning includes environmental state vectors, equipment state vectors, predicted environmental state values, comprehensive risk assessment values, and warehouse environmental parameter values. The actions of deep reinforcement learning are the control parameters of the environmental equipment in the warehouse area where the irradiated items are located. This environmental control equipment includes air conditioners, dehumidifiers, and exhaust fans in the warehouse area. The corresponding control parameters include the control status of the air conditioners (off, cooling, heating, etc.) and a reasonable fine-tuning range (e.g., temperature fluctuation range), the power of the dehumidifiers, and the rotation speed of the exhaust fans (described by wind speed levels).
[0149] The reward function for the deep reinforcement learning described above is:
[0150] ;
[0151] in, This represents the value of the reward function. , , , All represent preset weighting coefficients. , , , The importance of balancing stability, energy consumption, and security. This represents the squared Euclidean distance between the current environmental state vector and the warehouse environmental parameter values, penalizing environmental fluctuations. Represents the current environment state vector. This represents the storage environment parameter values. This includes storage temperature parameters, storage humidity parameters, and storage light intensity parameters. As a stability penalty, express The total energy consumption of environmental control equipment within the storage area at any given time (usually in kilowatts, obtained from equipment operation). As an energy penalty, (.) represents a conditional function. express The corresponding environmental parameter range, This indicates that the environment has exceeded the limits. This represents the closest distance between the current environmental estimate and the safety boundary (i.e., The difference between the maximum (or minimum) value in the corresponding environmental parameter range and the current environmental estimate in the current environmental state vector. This indicates a preset safety threshold. In order to transcend punishment, As a safety reward. (.) is a conditional function that imposes a large penalty or reward when the environment exceeds the limits or is in a safe state.
[0152] It should be noted that the above reward function uses the same quantity for calculation, and each quantity is calculated using only numerical values (i.e., units are not introduced). For example, when the current environmental state vector is the current temperature estimate, the warehouse environment parameter value... Environmental parameter range, current environmental estimate, and safety range boundary (i.e.) The maximum (or minimum) value in the corresponding environmental parameter range should be the relevant value corresponding to the temperature. The corresponding environmental parameter range refers to the calculated range. The corresponding numerical ranges. The same applies to humidity and light intensity, which will not be elaborated upon here.
[0153] It should be noted that the aforementioned states include the current environmental and equipment states, short-term forecast data, and storage control conditions (i.e., warehouse environment parameter values). Problems involving high dimensionality and complex relationships require deep neural networks to automatically learn effective feature representations from high-dimensional states, making them suitable for handling such high-dimensional nonlinear problems. The network structure takes states as input and outputs an estimate of the expected value Q (i.e., the expected long-term cumulative reward) for each selectable action. The training process of a deep reinforcement learning network is as follows:
[0154] (1) Experience replay: The model stores the interaction experience of each step in a fixed-size replay buffer. During training, small batches of data are randomly sampled from the buffer to break the correlation between data.
[0155] (2) Target network: Use a target network with the same structure but slower parameter updates to calculate the target value for stable training.
[0156] (3) Minimize the loss function by gradient descent and iteratively update Q.
[0157] Deep reinforcement learning algorithms construct the current state, select the optimal action according to the policy, and complete a relatively simple calculation process. Once trained, this lightweight DRL network can run efficiently on edge devices (i.e., edge nodes).
[0158] It is understandable that after obtaining a deep reinforcement learning network based on deep reinforcement learning training, the current state space is input into the deep reinforcement learning network for processing, and the current actions of the environmental control equipment in the storage area can be obtained in a future time period (generally a period of time after the current time).
[0159] Specifically, the current action includes the control status (off, cooling or heating) of the air conditioning in the storage area in the future time period and a reasonable fine-tuning range (e.g., temperature fluctuation range), the power of the dehumidifier in the storage area, and the speed of the exhaust fan in the storage area (which can also be described by wind speed level).
[0160] Step 16: Control the environmental control equipment in the storage area based on the current action.
[0161] That is, after obtaining the current action, it is necessary to control the environmental control equipment according to the control parameters of each environmental control device in the current action in order to achieve environmental regulation of the irradiation warehouse.
[0162] Understandably, upon receiving the current action, the edge node can transmit these control parameters through the transport layer to the control terminal of the corresponding environmental control device, enabling the device to control the environment and achieve environmental regulation. Alternatively, for environmental control devices lacking a control terminal, the control parameters can be transmitted to the terminal devices of relevant operators, allowing them to adjust the equipment.
[0163] It should be noted that by constructing a real-time data acquisition and fusion processing architecture (which consists of a device perception layer, a transmission layer, edge nodes, and a cloud server, the functions of which have been explained in detail above), edge nodes (such as terminal devices near the storage area) combine real-time status data, cloud predictions, and storage control conditions (i.e., storage environment parameter values) to make millisecond-level optimal control decisions to control environmental control equipment. This requires collaborative processing between the "edge" and the cloud. Edge processing continuously collects real-time data, performs preprocessing and data fusion, and outputs status data (environmental and equipment status data). On the one hand, it periodically uploads a small segment of historical status data to the cloud; on the other hand, it inputs this data into the state builder of the control agent to achieve millisecond-level rapid response control. Cloud processing collects status data from various storage edge nodes (different regions may correspond to different edge nodes), runs the environmental assessment model at a relatively low frequency, generates short-term future environmental predictions and risk values, and distributes them to the edge nodes. The edge nodes obtain the optimal action based on the current real-time status and issue instructions through the action executor, forming online decisions. Through continuous learning, the intelligent agent at this edge node eventually achieves dynamic, precise, and adaptive intelligent control.
[0164] In summary, the control method provided in this application has the following advantages:
[0165] (1) Designing a "perception-prediction-decision" paradigm with forward-looking regulation:
[0166] Based on the perception of multimodal data such as voice and object images, a key "prediction-decision" link is introduced to construct a predictive intelligent system. Its core is a deep learning-based environmental assessment model. This model proactively predicts the evolution trend of the environmental state and its comprehensive risk data over a future period by analyzing environmental state data, equipment state data, and the storage and control conditions of objects. This completely solves the lag problem of traditional control methods, providing an extremely stable storage environment for irradiated objects, enabling forward-looking control, and avoiding environmental fluctuations.
[0167] (2) Multi-module perception + evaluation model with dynamic adaptation and personalized strategies:
[0168] Personalized "item storage control conditions" are generated through multimodal perception (speech, image). These conditions are introduced into the cloud-based environmental assessment model prediction and the edge-based control agent decision-making. This means that when the stored items change, the system can automatically switch to the corresponding storage requirements and make predictions and controls that meet the new needs, thus achieving personalized, adaptive, and precise management.
[0169] (3) Constructing a deep reinforcement learning agent. This agent is not an engine that executes fixed rules, but a device controller that can autonomously learn the optimal strategy from data through continuous interaction with the environment. The agent's state space integrates the environmental state and device state, the short-term future predictions and risk assessments output by the environmental assessment model, and the storage control conditions of the items, comprehensively considering "current state + future cognition", which makes the decision control robust and able to handle complex information containing uncertainty.
[0170] Based on the above advantages, this application can achieve personalized and adaptive precise management for the needs of items to be irradiated, effectively avoid the shortcomings of fixed threshold control, significantly improve the control accuracy of the storage environment, and effectively avoid environmental fluctuations.
[0171] The above description is the preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principles described in this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A method for controlling the irradiated storage environment based on multimodal data fusion, characterized in that, include: The system acquires voice commands input by the user and obtains image data of the item to be irradiated; the voice commands are used to describe the required storage environment information for the item to be irradiated. Based on the voice command and the image data, the storage environment parameter values of the item to be irradiated are determined; The environmental data of the storage area where the irradiated item is located is fused and estimated to obtain the current environmental state vector, current equipment state vector, historical environmental state vector, and historical equipment state vector of the storage area. Based on the storage environment parameter values, the historical environment state vector, and the historical equipment state vector, environmental state prediction is performed to obtain the predicted environmental state value of the storage area in the future period, as well as the comprehensive risk assessment value of the predicted environmental state value. The current environmental state vector, the current equipment state vector, the predicted environmental state value, the comprehensive risk assessment value, and the warehouse environment parameter value are used as the current state space. The current state space is then input into a trained deep reinforcement learning network for processing to obtain the current action. The current action includes the control parameters of the environmental control equipment in the warehouse area. The environmental control equipment in the warehouse area is controlled based on the current action. The environmental state prediction based on the warehouse environment parameter values, the historical environmental state vector, and the historical equipment state vector yields a predicted environmental state value for the warehouse area in a future time period, and a comprehensive risk assessment value based on the predicted environmental state value, including: The warehouse environment parameter values, the historical environment state vector, and the historical equipment state vector are input into the evaluation model for processing to obtain the predicted environmental state value of the warehouse area in the future time period; the evaluation model includes an LSTM encoder, an attention module, an LSTM decoder, and an output layer connected in sequence. Based on the hidden state of the LSTM decoder when generating the last prediction point, the comprehensive risk assessment value of the environmental state prediction is obtained. The comprehensive risk assessment value of the environmental state prediction obtained based on the hidden state of the LSTM decoder when generating the last prediction point includes: The hidden state of the LSTM decoder when generating the last prediction point is calculated using a fully connected layer to obtain the calculation result; The calculation results are normalized using the Sigmoid activation function to obtain a comprehensive risk assessment value for the predicted environmental state. This comprehensive risk assessment value reflects the probability and severity of the predicted environmental state deviating from the storage requirements of the items in the future.
2. The control method according to claim 1, characterized in that, The step of determining the storage environment parameter values of the item to be irradiated based on the voice command and the image data includes: The storage environment parameter values of the items to be irradiated are calculated using the following formula. : ; in, This indicates the preset voice confidence level. This represents the environmental parameter values identified from the voice commands. This indicates the preset image confidence level. This indicates the environmental parameter values determined from the knowledge base of warehouse item environmental requirements based on the image data.
3. The control method according to claim 1, characterized in that, The step of fusing and estimating the environmental data of the storage area where the irradiated item is located to obtain the current environmental state vector, current equipment state vector, historical environmental state vector, and historical equipment state vector of the storage area includes: The Kalman filter algorithm is used to fuse and estimate the environmental data of the storage area where the irradiated items are located; Based on the environmental estimates obtained by fusion estimation using the Kalman filter algorithm, the current environmental state vector, current equipment state vector, historical environmental state vector, and historical equipment state vector of the storage area are obtained.
4. The control method according to claim 3, characterized in that, The current environmental state vector includes: the estimated current temperature of the storage area, the estimated current humidity of the storage area, and the estimated current light intensity of the storage area; The current device state vector includes: the current estimated power of the air conditioner in the storage area, the current estimated power of the dehumidifier in the storage area, and the current estimated speed of the exhaust fan in the storage area; The historical environmental state vector includes: temperature estimates of the storage area at multiple historical moments, humidity estimates of the storage area at multiple historical moments, and light intensity estimates of the storage area at multiple historical moments; The historical equipment state vector includes: the estimated power of the air conditioner in the storage area at multiple historical moments, the estimated power of the dehumidifier in the storage area at multiple historical moments, and the estimated rotational speed of the exhaust fan in the storage area at multiple historical moments.
5. The control method according to claim 1, characterized in that, The deep reinforcement learning network is obtained through deep reinforcement learning training. The state space of deep reinforcement learning includes environmental state vector, equipment state vector, environmental state prediction value, comprehensive risk assessment value, and warehouse environment parameter value. The actions learned through deep reinforcement learning are the control parameters of the environmental equipment in the storage area where the items to be irradiated are located.
6. The control method according to claim 1, characterized in that, The environmental control equipment includes air conditioners, dehumidifiers, and exhaust fans.
Citation Information
Patent Citations
Cold storage energy-saving control method based on deep reinforcement learning
CN118912801A
Intelligent granary ventilation and energy consumption optimization decision-making method based on reinforcement learning
CN120725247A