Board clamping detection method and device, computer equipment and storage medium
By collecting multi-source data from multiple sensors in real time, pre-processing and data fusion, and using reinforcement learning algorithms to detect card abnormalities, the problem of difficulty in accurately detecting card plate phenomena in the vertical furnace in the existing technology is solved, and accurate monitoring of equipment status and improvement of production efficiency is achieved.
Patent Information
- Application Number
- CN202510277465.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art is difficult to accurately detect the phenomenon of the snail in the vertical furnace, especially when it is difficult to install a camera or image sensor inside the equipment.
By obtaining multi-source data that monitors the device status in real time, pre-processing and data fusion are performed, board abnormality detection is performed using reinforcement learning algorithms, and equipment parameters are adjusted when abnormalities are detected.
Accurate identification and real-time adjustment of the cardboard phenomenon is achieved, avoiding equipment shutdown or damage, and improving production efficiency.
Smart Images

Figure CN120217149A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a pallet detection method, and more specifically to a pallet detection method, device, computer device and storage medium. Background Art
[0002] The phenomenon of pallet jamming in a vertical furnace refers to the situation where the sheet material gets stuck, blocked or fails to move smoothly when passing through devices such as conveyor belts and guide rails inside the furnace or in the conveying system. The reasons for the occurrence of pallet jamming in a vertical furnace can be analyzed from aspects such as equipment design, operation problems, material problems, insufficient maintenance and environmental factors. In terms of equipment design, if the conveyor belt or guide rail is not reasonably designed, it may cause friction or jamming; improper operation, incorrect program setting or sensor configuration problems may also lead to pallet jamming. Material problems such as inconsistent sheet size, deformation or defects will also affect the smooth transmission. Lack of maintenance such as conveyor belt wear, motor failure, etc., will exacerbate the pallet jamming phenomenon. Environmental factors, such as high temperature or dust accumulation, may also cause abnormal operation of the equipment, thus triggering the pallet jamming problem.
[0003] Currently, it is a relatively common method to detect whether pallet jamming occurs through software algorithms. These algorithms can monitor the operating state of the equipment in real time and determine whether there is a pallet jamming phenomenon. Specifically, traditional vision detection methods mainly rely on cameras or image sensors installed on the equipment to capture image data on the surface of the equipment or a specific area. By processing and analyzing these images, the system attempts to identify the pallet jamming phenomenon or other abnormal situations. However, since pallet jamming often occurs inside the furnace, it is actually very difficult to install cameras and image sensors inside the furnace, and traditional vision detection is difficult to accurately judge whether pallet jamming occurs.
[0004] Therefore, it is necessary to design a new method to effectively identify the pallet jamming phenomenon and make adjustments or give alarms in a timely manner. This can avoid equipment shutdown or damage and improve production efficiency. Summary of the Invention
[0005] The purpose of the present invention is to overcome the defects of the prior art and provide a pallet detection method, device, computer device and storage medium.
[0006] To achieve the above purpose, the present invention adopts the following technical solutions: A pallet detection method, comprising:
[0007] Obtaining multi-source data collected by different sensors for real-time monitoring of the equipment state;
[0008] Performing preprocessing on the multi-source data to obtain a preprocessing result;
[0009] Performing data fusion on the preprocessing result to obtain a fusion result;
[0010] Perform pallet anomaly detection using a reinforcement learning algorithm based on the fusion result to obtain a detection result;
[0011] When the detection result indicates a pallet anomaly state, adjust the device parameters to real-time adjust the device operating state.
[0012] A further technical solution thereof is: The sensors include vibration sensors, photoelectric sensors, pressure sensors, temperature sensors, and encoders.
[0013] A further technical solution thereof is: The preprocessing of the multi-source data to obtain a preprocessing result includes:
[0014] Perform denoising filtering on the multi-source data to obtain a filtering result;
[0015] Perform timestamp alignment on the filtering result to obtain an alignment result;
[0016] Extract features from the alignment result to obtain a preprocessing result.
[0017] A further technical solution thereof is: The preprocessing result includes time domain features, frequency domain features, features obtained by time-frequency analysis, and statistical features.
[0018] A further technical solution thereof is: The data fusion of the preprocessing result to obtain a fusion result includes:
[0019] Use Kalman filtering and covariance matrix weighting techniques to perform data fusion on the preprocessing result to obtain a fusion result.
[0020] A further technical solution thereof is: The performing pallet anomaly detection using a reinforcement learning algorithm based on the fusion result to obtain a detection result includes:
[0021] Input the fusion result into an anomaly detection model to perform pallet anomaly detection to obtain a detection result;
[0022] Among them, the anomaly detection model is obtained by training a Q-Learning strategy by defining states, device actions, and reward functions.
[0023] A further technical solution thereof is: The anomaly detection model is obtained by training a Q-Learning strategy by defining states, device actions, and reward functions, including:
[0024] The anomaly detection model is formed by defining states, device actions, and reward functions, using interactive learning, where the agent executes actions in the environment and collects experience, using an ε-greedy strategy to balance exploration and exploitation, updating the strategy by minimizing the temporal difference error, and optimizing the reward function to maximize the cumulative reward.
[0025] The present invention also provides a pallet detection device, including:
[0026] An acquisition unit, configured to acquire multi-source data collected by different sensors for real-time monitoring of the device status;
[0027] A preprocessing unit, configured to preprocess the multi-source data to obtain a preprocessing result;
[0028] A data fusion unit, configured to perform data fusion on the preprocessing result to obtain a fusion result;
[0029] An anomaly detection unit, configured to use a reinforcement learning algorithm to perform pallet anomaly detection based on the fusion result to obtain a detection result;
[0030] An adjustment unit, configured to adjust the device parameters when the detection result indicates a pallet anomaly state, so as to adjust the device operation state in real time.
[0031] The present invention also provides a computer device, which includes a memory and a processor. A computer program is stored on the memory, and when the processor executes the computer program, the above method is implemented.
[0032] The present invention also provides a storage medium, which stores a computer program, and when the computer program is executed by a processor, the above method is implemented.
[0033] The beneficial effects of the present invention compared with the prior art are as follows: By collecting multi-source data from multiple sensors in real time, preprocessing and performing data fusion, the present invention can accurately identify the operation state of the device; based on the fused data, a reinforcement learning algorithm is used to perform pallet anomaly detection to evaluate in real time whether the device is in an abnormal state; once a pallet anomaly is detected, the system will automatically adjust the device parameters, such as reducing the operating speed or starting a cooling device, or sending an alarm signal, so as to timely respond to the abnormal situation, thereby effectively avoiding device downtime or damage, ensuring the stable operation of the device, and improving production efficiency.
[0034] The following further describes the present invention with reference to the accompanying drawings and specific embodiments. Description of the Drawings
[0035] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0036] Figure 1Schematic flowchart of the card board detection method provided by an embodiment of the present invention;
[0037] Figure 2 Schematic sub - flowchart of the card board detection method provided by an embodiment of the present invention;
[0038] Figure 3 Schematic block diagram of the card board detection device provided by an embodiment of the present invention;
[0039] Figure 4 Schematic block diagram of the pre - processing unit of the card board detection device provided by an embodiment of the present invention;
[0040] Figure 5 Schematic block diagram of the computer device provided by an embodiment of the present invention. Detailed implementation manners
[0041] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0042] It should be understood that when used in this specification and the appended claims, the terms "include" and "comprise" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0043] It should also be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in this specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.
[0044] It should be further understood that the term " / and / or" used in this specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0045] Please refer to Figure 1 , Figure 1Schematic flowchart of the pallet detection method provided by the embodiments of the present invention. The pallet detection method is applied to a server. The server interacts with multiple different sensors and the controller of the device, monitors the device status in real time through multiple sensors, collects multi-source data and performs preprocessing, including denoising, timestamp alignment, and feature extraction. Then, data fusion is performed using Kalman filtering and weighted covariance matrix techniques to obtain a fusion result. Based on the fusion result, a reinforcement learning algorithm (such as Q-Learning) is used for pallet anomaly detection to identify abnormal states in a timely manner. If a pallet anomaly is detected, the system will automatically adjust the device parameters or issue an alarm to prevent the device from stopping or being damaged, thereby improving production efficiency.
[0046] Figure 2 is the schematic flowchart of the pallet detection method provided by the embodiments of the present invention. As Figure 2 shown, the method includes the following steps S110 to S150.
[0047] S110. Obtain multi-source data collected by different sensors for real-time monitoring of the device status.
[0048] In this embodiment, the multi-source data includes monitoring data from multiple sensors, and these sensors are respectively used to collect various physical parameters of the device in different operating states. Specifically, the multi-source data includes the following categories:
[0049] Vibration data: Collected by vibration sensors, mainly monitoring the mechanical vibration of the device to identify whether there is abnormal wear or mechanical failure of the device.
[0050] Position and motion state data: Collected by optoelectronic sensors, real-time monitoring the position and motion state of the sheet on the conveyor belt to ensure that the sheet moves correctly and avoid misalignment or blockage.
[0051] Pressure data: Collected by pressure sensors, monitoring the contact pressure between the sheet and the guide rail to detect whether there is excessive pressure or uneven pressure, and timely discover jams or abnormal phenomena.
[0052] Temperature data: Collected by temperature sensors, monitoring the temperature of key components of the device to help prevent overheating or potential device failures.
[0053] Speed data: Collected by encoders, feeding back the movement speed of the conveyor belt to ensure that the materials flow at a predetermined speed and avoid misalignment or blockage.
[0054] These data are sampled in real time through a data acquisition system and can be transmitted to a monitoring terminal or the cloud through a communication protocol for further processing and analysis, so as to achieve comprehensive monitoring and anomaly detection of the device status.
[0055] Specifically, the sensors include vibration sensors, photoelectric sensors, pressure sensors, temperature sensors, and encoders.
[0056] The multi-source data includes monitoring data from multiple sensors, which are respectively used to collect various physical parameters of the equipment under different operating states to comprehensively reflect the state of the equipment.
[0057] The vibration sensors are installed at easily worn positions such as the bearings and guide rails of the conveyor belt to monitor the mechanical vibration of the equipment. A commonly used model is the IEPE acceleration sensor, which can cover vibration signals with a frequency range of 0.5 - 10 kHz. By monitoring the changes in vibration, it is possible to identify whether there is abnormal wear or mechanical failure in the equipment.
[0058] The photoelectric sensors are deployed at the feeding port, discharging port, and key corners to detect the position and motion state of the sheet in real time. For example, using Omron E3Z series photoelectric sensors, position detection can be carried out by reflecting or passing through the sheet to ensure the correct position of the sheet on the conveyor belt and avoid problems such as misalignment or blockage.
[0059] The pressure sensors are integrated in the guide wheels or transmission mechanisms to monitor the contact pressure between the sheet and the guide rail. The common measurement range is 0 - 500 N, and the accuracy is ±1%. By measuring the contact pressure, it is possible to detect whether the sheet is under excessive or uneven pressure during the conveying process and timely discover possible jams or abnormal phenomena.
[0060] The temperature sensors are mainly installed at positions such as motors, bearings, and high-temperature areas of the furnace body. Common temperature sensors include PT100 thermal resistors or infrared temperature measurement modules, which can accurately measure the temperature of key components of the equipment. Excessive temperature may mean that the equipment is overheating or there is a risk of failure. Timely obtaining temperature data helps prevent equipment damage.
[0061] The encoders work in cooperation with the servo motors to provide real-time feedback on the speed of the conveyor belt. Incremental encoders usually have a high resolution (such as ≥1000 pulses / revolution), which can accurately monitor the motion state of the conveyor belt to ensure that the materials flow at a predetermined speed and avoid misalignment or blockage during transportation.
[0062] These sensors are connected to a data acquisition system (such as an industrial PLC or an edge computing device) to form a complete monitoring system. Through synchronous sampling, the data acquisition system can obtain sensor data in real time and transmit the data to the monitoring terminal or the cloud for processing and analysis through a suitable communication protocol (such as Profinet, EtherCAT, or MQTT protocol), thereby achieving comprehensive monitoring and anomaly detection of the equipment state.
[0063] S120. Preprocess the multi-source data to obtain a preprocessing result.
[0064] In this embodiment, the preprocessing result includes time-domain features, frequency-domain features, features obtained from time-frequency analysis, and statistical features.
[0065] In one embodiment, referring to Figure 2 , the above step S120 may include steps S121 to S123.
[0066] S121. Denoise and filter the multi-source data to obtain a filtering result.
[0067] In this embodiment, the filtering result refers to the filtered signal.
[0068] Specifically, multi-source data usually consists of various types of data such as vibration signals, pressure signals, and temperature signals, and these data may be disturbed by various noises. The purpose of denoising and filtering is to make the data purer by filtering out the noise signals, so as to provide a more reliable input for subsequent analysis. The specific filtering methods are as follows:
[0069] Denoising of vibration signals: To remove high-frequency noise in vibration signals, a low-pass filter (usually with a cut-off frequency of 500 Hz) can be used to remove the high-frequency components in the signal. In addition, wavelet denoising technology (for example, using Daubechies wavelet basis) can also be used to effectively remove noise through wavelet transform while retaining the main features of the signal.
[0070] Denoising of pressure signals: Pressure signals are often affected by instantaneous fluctuations. To reduce the impact of such fluctuations, moving average filtering (for example, with a window length set to 50 ms) can be used. This method smooths the signal and removes short-term irregular changes, making the signal more stable.
[0071] The final result of denoising and filtering is the filtered signal, which removes the noise and can more accurately reflect the actual state of the sensor.
[0072] S122. Align the timestamps of the filtering result to obtain an alignment result.
[0073] In this embodiment, the alignment result refers to the signal obtained after aligning the timestamps of the filtered signal.
[0074] In a multi-sensor system, each sensor may have a different sampling frequency, resulting in inconsistent timestamps of the data they collect. Therefore, it is necessary to align the timestamps of these data to ensure that the data from different sensors can be compared and fused at the same moment. Common alignment methods include:
[0075] Linear interpolation: For sensors with a relatively low sampling frequency (such as optoelectronic sensors), their data can be interpolated onto the timestamps of sensors with a higher sampling frequency (such as encoders). Through interpolation techniques, low-frequency data can be aligned to high-frequency timestamps.
[0076] Moving average: For high-frequency sensor data (such as encoder speed data), its data can be downsampled to match the time resolution of low-frequency sensors (such as optoelectronic sensor position data). This method smooths the data to ensure that data of different frequencies can be synchronized.
[0077] After timestamp alignment, the data of all sensors are adjusted to the same time base, eliminating the impact of time differences on subsequent analysis.
[0078] S123. Extract features from the alignment result to obtain a preprocessing result.
[0079] In this embodiment, the aligned multi-source data is already data with consistent time series. Next, useful features need to be extracted from it to support subsequent fault diagnosis or prediction. Feature extraction can be classified into the following categories:
[0080] Time-domain features: Extract some statistical features from the original time-domain data of the signal, such as the mean, variance, peak factor, and kurtosis of the signal, etc. These features reflect the overall trend, volatility, and impact of the signal. For example, kurtosis can be used to reflect the impact vibration in the signal, which is very important in mechanical fault diagnosis.
[0081] Frequency-domain features: Analyze the spectrum of the signal by performing a fast Fourier transform (FFT) on the signal. Frequency-domain features help to identify the frequency characteristics of specific faults such as bearing faults (such as BPFO and BPFI). These frequency components can reveal whether there are damages or faults in mechanical components.
[0082] Time-frequency analysis: Use methods such as wavelet packet decomposition to perform time-frequency analysis on the signal and extract the energy proportion of different frequency bands. This analysis method is applicable to non-stationary signals and can better capture the dynamic change characteristics of the signal.
[0083] Statistical features: For example, the moving standard deviation can be used to monitor the fluctuation trend of the pressure signal. Through statistical features, the stability and health status of the system can be evaluated.
[0084] Since the sampling frequencies of the encoder (speed sensor) and the optoelectronic sensor (position sensor) may be different, time synchronization needs to be performed through methods such as interpolation (e.g., linear interpolation) or moving average. Unit conversion is performed on different types of data. For example, the speed data of the encoder (unit: mm / s) is converted into a position increment, while the optoelectronic sensor directly outputs the position coordinates (unit: mm).
[0085] The time-domain, frequency-domain, time-frequency, and statistical features extracted through the above steps constitute the final result of data preprocessing. These features will provide the necessary basic data for subsequent tasks such as health monitoring and fault diagnosis.
[0086] The multi-source data preprocessing process in step S120 converts the original data into a clearer, more accurate, and usable preprocessing result through methods such as denoising filtering, timestamp alignment, and feature extraction. These preprocessing results include different types of features (time-domain features, frequency-domain features, time-frequency analysis, and statistical features), providing reliable inputs for subsequent data analysis, modeling, and decision-making.
[0087] S130. Perform data fusion on the preprocessing result to obtain a fusion result.
[0088] In this embodiment, the fusion result refers to the signal after the preprocessed signals are fused together.
[0089] Specifically, the Kalman filter and covariance matrix weighting technology are used to perform data fusion on the preprocessing result to obtain a fusion result.
[0090] The Kalman filter algorithm will be applied to the preprocessed data for data fusion. The basic principle of the Kalman filter is to dynamically adjust the weights of prediction and observation through the weighted combination of state estimation and observation data to obtain the optimal fusion result.
[0091] State definition: The state vector X = [x, v]T, where x represents the position and v represents the speed.
[0092] Observation vector Z = [x photo , v encoder T, where x photo is the measured position of the optoelectronic sensor, and v encoder is the measured speed of the encoder.
[0093] Prediction stage:
[0094] The state transition matrix F assumes that the system is in a uniform motion state, that is, the position is predicted based on the speed. For example, X k = F k-1 X k-1 + B k-1 μ k-1 + ωk-1 ;
[0095] The process noise covariance Q is usually set according to the dynamic characteristics of the system and is used to describe the uncertainty in the process. For example, Q = diag(0.1, 0.5), which represents the noise characteristics in the position and velocity directions.
[0096] Update phase: P k|k = (I - K k H k )P k|k-1 ;
[0097] The observation matrix H describes the relationship between the sensor measurement results and the system state, usually a direct linear relationship.
[0098] The observation noise covariance R is set according to the performance and calibration data of different sensors. For example, the noise of a photoelectric sensor may be higher, while the noise of an encoder is lower. Therefore, R11 and R22 can be set to 1.0 and 0.5 respectively.
[0099] Output the fusion result: Finally, through the update step of the Kalman filter, the fused state estimate value is obtained, which includes the optimal position estimate value (x fused ) and the velocity estimate value (v fused ).
[0100] Specifically, the final position estimate value xfused = xfused[0] = Xk[0], and the velocity estimate value vfused = vfused[1] = Xk[1].
[0101] Specifically, the state vector X = [x, v], where x refers to the real-time position of the sheet on the conveyor belt (unit: mm), reflecting whether the sheet deviates from the center line. v is the real-time speed of the conveyor belt (unit: mm / s), which affects the stability of sheet transportation.
[0102] State transition matrix Describes the variation law of the state over time, assuming the sheet moves at a constant speed. t is the sampling period.
[0103] The process noise covariance Q represents the uncertainty in state prediction, and the sources include conveyor belt slippage or motor control error (velocity noise), and random position offset caused by mechanical vibration. For example, Q = diag(0.1, 0.5), (position noise 0.1mm 2 , velocity noise 0.5 (mm / s) 2 );
[0104] The observation matrix H represents that the encoder directly measures speed and the optoelectronic sensor directly measures position. The observation noise matrix R represents the encoder speed measurement error and the optoelectronic sensor position error.
[0105] In addition to Kalman filtering, covariance matrix weighting technology is also adopted to fuse multi-modal data. This method is used to weight data according to the characteristics and reliability of different sensors. The specific steps include:
[0106] Standardize various sensor data (such as vibration, pressure, temperature, encoder speed, etc.) to ensure consistent data dimensions. Z-score standardization or Min-Max normalization can be used.
[0107] Allocate weights based on expert experience. For example, according to Failure Mode and Effects Analysis (FMEA), determine the weight of each sensor. For instance, the vibration sensor may have a greater association with faults, so its weight is 0.4, while the temperature sensor has a lower weight of only 0.2. In a high-temperature environment, the weight of the temperature sensor can be increased, as shown in Table 1.
[0108] Table 1. Results of weight allocation based on expert experience
[0109]
[0110]
[0111] Alternatively, use principal component analysis to reduce the dimension of the standardized data and allocate weights according to the variance contribution rate of the principal components. λi is the variance of the i-th principal component, reflecting the contribution of this component to the data variation. ω i It means that the weight of the vibration sensor is 60% and vibration anomalies need to be monitored preferentially. The weight of the temperature sensor is 15%, which is a secondary factor, but its weight may need to be dynamically increased under high-temperature working conditions.
[0112] For example, if the eigenvalue of the vibration sensor is the largest, it indicates that the vibration data is the main influencing factor of the pallet. For instance, PCA analysis shows that the variance contribution rate of the vibration data is at most 60%, so its weight is 0.6, the weight of pressure is 0.25, and the weight of temperature is 0.15.
[0113] Finally, based on Kalman filtering and covariance weighting technology, fuse the preprocessed signals of multiple sensors to obtain a more accurate and reliable result of the sheet metal trajectory tracking. The output of the fusion result includes:
[0114] Position estimate (x fused ): Represents the optimized position data.
[0115] Velocity estimate (vfused ) represents the optimized speed data.
[0116] Combining the advantages of Kalman filtering and covariance matrix weighting, through multi-sensor data fusion, the accuracy and robustness of the system can be effectively improved, especially in dynamic and complex environments.
[0117] S140. Use a reinforcement learning algorithm to perform pallet abnormality detection based on the fusion result to obtain a detection result.
[0118] In this embodiment, the detection result refers to the determination result of whether there is a pallet abnormality.
[0119] Specifically, input the fusion result into an abnormality detection model to perform pallet abnormality detection to obtain a detection result;
[0120] Among them, the abnormality detection model is obtained by training the Q-Learning strategy by defining states, device actions, and reward functions.
[0121] The abnormality detection model is a model formed by using interactive learning by defining states, device actions, and reward functions. The agent executes actions in the environment and collects experiences, uses the ε-greedy strategy to balance exploration and exploitation, updates the strategy by minimizing the temporal difference error, and optimizes the reward function to maximize the cumulative reward.
[0122] For the training of the model, in the application scenario of pallet abnormality detection, the core goal of the problem is to learn an optimal control strategy through the interaction between the agent and the environment, so as to detect and prevent the occurrence of pallets in real time during the operation of the device. This process uses the Q-Learning algorithm to train an effective abnormality detection model through the definition of states, actions, and reward functions.
[0123] If the Q-learning or DQN model is used for abnormality detection, the following methods can be adopted:
[0124] First, it is necessary to define what kind of behaviors or states can be regarded as "abnormal". For example:
[0125] Sensor data such as vibration amplitude and pressure value exceeds the normal operation range;
[0126] When the model selects certain actions (such as reducing the speed by 10% or triggering cleaning), and fails to bring the expected reward, it may be considered abnormal.
[0127] Next, during the model training process, Q-learning will learn which state-action pairs are optimal and select the most effective strategy by accumulating rewards. If, in the new test data, the strategy executed by the model is significantly different from the optimal strategy obtained during training, it may indicate that some anomaly has occurred.
[0128] Finally, when the agent takes certain actions, if the reward value is significantly lower than the reward in the normal state, it indicates that the current state may be abnormal. If the decision of the model deviates significantly from the optimal decision learned in historical training, this may indicate that the current state is abnormal. For example, if the agent frequently selects the deceleration action in certain specific states without obtaining the expected positive reward, it may be that there is a problem with the device.
[0129] In a normally operating environment, train the Q-learning or DQN model through the normal state-action-reward sequence. Input the new, potentially abnormal state data into the model and observe whether the actions it takes are consistent with the expected optimal actions. By comparing the rewards returned by the model, the deviation of action selection, and whether abnormal states occur, determine whether there is abnormal behavior.
[0130] In addition to the reinforcement learning model, traditional anomaly detection methods (such as statistics-based anomaly detection or machine learning-based classification methods) can also be combined to assist in identifying anomalies. For example, the historical distribution of sensor data can be used to determine whether the current data deviates from the normal range.
[0131] For model training, the definition of the state space is the key to anomaly detection and represents the state of the device at a certain moment. According to the sensor data, the state space can be divided into two forms:
[0132] Continuous state: The original data provided by the sensor, which is also the fusion result of this embodiment, such as vibration amplitude, pressure, temperature, speed, etc.
[0133] Discrete state: By segmenting the sensor data, the continuous data is converted into discrete states. For example:
[0134] Vibration amplitude: High, medium, low.
[0135] Pressure value: High, normal, low.
[0136] Speed value: Fast, normal, slow.
[0137] Temperature value: High, normal, low.
[0138] In practical applications, these continuous sensor data are usually normalized or discretized so that the Q-Learning model can process them better.
[0139] The action space defines the various behaviors that the agent can take. In this embodiment, the actions are divided into two categories: discrete actions and continuous actions; among them, the discrete actions include:
[0140] Reduce speed by 10%: Reduce the speed of the device by 10%.
[0141] Maintain speed: Keep the current speed unchanged.
[0142] Increase speed by 10%: Increase the device speed by 10%.
[0143] Trigger cleaning: Start the cleaning operation to prevent jamming.
[0144] The continuous actions include:
[0145] Speed adjustment amount (Δv ∈ [-0.1, 0.1] m / s): The agent selects a suitable speed adjustment value according to the state of the current environment, making the operation of the device smoother and preventing abnormalities from occurring.
[0146] The design of the reward function is the most important part of the reinforcement learning model, which determines the optimization direction of the agent's behavior. In the scenario of jamming anomaly detection, the reward function includes two aspects: immediate reward and energy-saving reward, comprehensively evaluating the agent's behavior:
[0147] Normal operation: If the device operates normally in the current state, the agent gets a positive reward (+0.1).
[0148] Jamming occurs: If jamming occurs, the agent will get a large negative reward (-10), and this round will terminate. This is a strong punishment for the occurrence of jamming anomalies.
[0149] Energy-saving reward: When the device achieves energy saving by reducing speed, an additional reward (+0.05) is given to encourage the agent to perform energy-saving control without causing jamming.
[0150] The core of the anomaly detection model is trained through a reinforcement learning algorithm (such as Q-Learning). The agent executes actions in the environment and gradually learns how to select the optimal action according to different states, so as to maximize the cumulative reward and avoid the occurrence of jamming anomalies.
[0151] At each time step, the agent selects an action according to the current state (described by sensor data).
[0152] After executing the action, the environment feedbacks the new state and calculates the reward according to the state change.
[0153] The agent updates its policy through these interaction data (i.e., trajectories) to optimize the decision-making process.
[0154] To ensure the learning effect of the model, the agent needs to find a balance between exploration (trying different actions) and exploitation (selecting the optimal action based on existing experience). The ε-greedy strategy is adopted to control this process. Initially, the agent explores more (ε = 1). As the training progresses, the exploration rate is gradually reduced, and the utilization rate is increased (ε gradually decays to 0.1).
[0155] Q-Learning updates the policy by minimizing the temporal difference error (TD Error). The update formula is as follows: L(θ) = E[(r + γa max Q(s′, a′; θ-) - Q(s, a; θ))2]
[0156] where r is the immediate reward, γ is the discount factor, Q(s, a; θ)) is the value estimate of the current state-action pair, and θ is the parameter of the policy network.
[0157] The target network (θ-) is used to stabilize the training, and the parameters of the policy network are synchronously updated to the target network regularly.
[0158] DQN (Deep Q-Network): To handle the high-dimensional state space, the tabular update method of Q-Learning may become inapplicable. Therefore, a deep neural network is adopted to approximate the Q function. DQN approximates the state-action value through a convolutional neural network (CNN) or a fully connected network (FCN).
[0159] Through the anomaly detection model trained by the above Q-Learning strategy, the agent can make reasonable decisions based on the current state of the device and predict whether a pallet jamming anomaly occurs.
[0160] The fusion results (sensor data such as vibration, pressure, speed, etc.) are input into the anomaly detection model trained by Q-Learning. The model judges the device state based on the current sensor data and selects appropriate actions (such as reducing speed, triggering cleaning, etc.) according to the learned strategy to avoid pallet jamming. The model judges whether the current device is in normal operation based on the actions of the agent and outputs the result of whether there is a pallet jamming anomaly. If pallet jamming occurs, the agent will receive a negative reward and terminate the current episode.
[0161] The training effect of the Q-Learning model depends largely on the setting of hyperparameters. The following are some key hyperparameter tuning strategies:
[0162] Learning rate (α): Usually selected between 1e-4 and 1e-3 to control the update speed of the Q value.
[0163] Discount factor (γ): Select a moderate discount factor (such as 0.9 - 0.99) to balance the weights of short-term and long-term rewards.
[0164] Exploration rate (ε): Adopt a linear decay strategy, starting from 1.0 and gradually decreasing to 0.1 to promote the agent to explore and utilize the learned knowledge.
[0165] Batch Size: Usually select a batch size of 32 to 256. A larger batch size can improve the stability and convergence speed of training.
[0166] By defining an appropriate state space, action space, and reward function, the Q-Learning strategy can effectively train an agent model for pallet abnormality detection. During the interaction between the agent and the environment, by balancing exploration and exploitation, the agent gradually optimizes its strategy and finally realizes the prevention and detection of pallet abnormalities.
[0167] S150. When the detection result is the pallet abnormal state, adjust the device parameters to adjust the device operation state in real time.
[0168] In the pallet abnormality detection system, ensuring the normal operation of the device and avoiding pallet occurrence is crucial. By combining the mechanisms of Action Masking and Safety Layer, it can quickly respond when detecting pallet abnormalities in real time, adjust the operation parameters of the device, and ensure the stability and safety of the device.
[0169] The core idea of action masking is to mask some unsafe or inappropriate actions when the agent (usually the agent in reinforcement learning) makes a decision. In the scenario of pallet abnormality detection, the operation state of the device needs to be strictly controlled, and over-speed operation may cause the device to enter a dangerous area and increase the risk of pallet occurrence. Therefore, the action masking mechanism can be used to mask these "dangerous" actions during the agent's decision-making process.
[0170] Specifically, the operation parameters of the device (such as vibration, temperature, pressure, etc.) are monitored in real time through sensors to determine whether the current device is in an abnormal state. If the possibility of pallet abnormality is detected or the device is about to have a pallet problem, the system will enter the warning state.
[0171] The application of action masking includes:
[0172] Disable over-speed actions: If a pallet abnormality is detected, the system will automatically disable "acceleration" - type actions to prevent the agent from choosing to run at an accelerated speed, thus avoiding the device entering a more dangerous state.
[0173] Prohibit unsafe adjustments: For example, if the device is already close to its limit operating conditions (such as temperature and pressure approaching the upper limit), then prohibit the agent from taking further radical adjustments, such as increasing speed or reducing cooling, etc.
[0174] With such an action mask, the system can ensure that in the abnormal state of the pallet, the agent can only select safe operations (such as decelerating or stopping), thereby reducing the probability of equipment damage and failures.
[0175] The safety layer mechanism is a real-time monitoring and alarm system. Its function is to detect abnormal situations in real time during the operation of the equipment and take corresponding emergency measures according to the detection results. In the system for detecting pallet abnormalities, when it is detected that the equipment may have a pallet abnormality, the safety layer will respond in a timely manner, issue graded alarms, and take actions.
[0176] Specifically, the system continuously monitors various sensor data of the equipment (such as temperature, pressure, vibration, rotational speed, etc.) and performs real-time analysis on the data through set thresholds. If it is found that the sensor data is abnormal (such as too high vibration value, too high temperature, abnormal pressure, etc.), the system will determine it as a preliminary sign of "pallet abnormality".
[0177] According to the severity of the abnormality, the safety layer will issue early warnings of different levels:
[0178] Level 1 early warning (yellow): When a minor abnormality is detected (such as slightly higher vibration or temperature of the equipment, but not reaching the equipment safety threshold), the system will issue a Level 1 early warning to prompt the operator to check the equipment. At this time, the operating parameters of the equipment can be appropriately adjusted, the load can be reduced, or regular inspections can be carried out.
[0179] Level 2 early warning (red): If the equipment has a more serious abnormality (such as obvious temperature fluctuations, excessive vibration, too high pressure, etc.) in the equipment, the system will issue a Level 2 early warning and require immediate shutdown. At this time, the action mask of the agent will disable all actions that may exacerbate the abnormality (such as accelerating, heating up, etc.), and force the equipment to stop or perform other emergency operations (such as starting the cooling system, adjusting the equipment parameters, etc.).
[0180] Handling of minor abnormalities: For the situation of Level 1 early warning, the agent can choose to reduce the operating speed of the equipment, reduce the load, or even start the equipment self-check program, thereby reducing the risk of abnormalities.
[0181] Handling of serious abnormalities: When a serious abnormality is detected, the system will immediately stop the equipment operation to prevent the equipment from suffering more serious failures due to being in an unsafe state for a long time. At the same time, the system may start some protective measures, such as starting the emergency shutdown program, enabling the standby cooling system, or starting the emergency exhaust system, etc.
[0182] When the abnormal state of the pallet is detected, the system not only needs to give a real-time alarm, but also must adjust the operating parameters of the equipment according to the detection results to ensure that the pallet problem is repaired in a timely manner and further deterioration is avoided.
[0183] Detect abnormal status of the card board: Through real-time sensor data, the system identifies whether there are signs of abnormal card board in the current device (such as excessive vibration, unstable pressure, etc.). Once the card board problem is confirmed, the system will immediately make corresponding parameter adjustments.
[0184] To prevent the abnormal card board from worsening further, the intelligent agent will slow down the running speed of the device through an action mask. Slowing down can reduce the load and stress of the device and avoid device damage caused by excessive speed.
[0185] If the status of the device is on the verge of danger (such as the temperature approaching the limit), the intelligent agent may choose to maintain the current low speed state to prevent the device from entering a more dangerous state.
[0186] If the temperature of the device is too high, it may cause the card board to occur. Therefore, the intelligent agent can adjust the parameters of the cooling system (such as starting more cooling devices) to reduce the device temperature and reduce the risk of abnormal card board.
[0187] When an abnormal card board occurs, the intelligent agent can also start the device self-check program to check whether there are other potential faults in the device, such as insufficient lubricating oil, abnormal transmission system, etc. If problems are found, the intelligent agent will make decisions to stop or repair according to the situation.
[0188] When the abnormal card board is in a severe state that cannot be controlled, the intelligent agent will automatically start the emergency stop program to completely stop the operation of the device and prevent the device from causing more serious damage due to continuous operation.
[0189] By combining the action mask and the security layer mechanism, the device can quickly respond in the abnormal card board state and adjust the operating parameters in real time. These adjustments include controlling the device speed through the intelligent agent, starting the cooling system, performing device self-check, or executing an emergency stop, etc., to ensure that the device will not deteriorate further due to the abnormal card board. The action mask mechanism avoids unsafe operations, and the security layer mechanism helps the operator intervene in a timely manner through hierarchical alarms, thus ensuring the stability and safety of the device.
[0190] The above card board detection method can accurately identify the operating state of the device by collecting multi-source data from multiple sensors in real time and performing preprocessing and data fusion; based on the fused data, an enhanced learning algorithm is used for card board abnormal detection to evaluate in real time whether the device is in an abnormal state; once an abnormal card board is detected, the system will automatically adjust the device parameters, such as reducing the running speed or starting the cooling device, or sending out an alarm signal to timely respond to the abnormal situation, thereby effectively avoiding device shutdown or damage, ensuring the stable operation of the device, and improving production efficiency.
[0191] Figure 3It is a schematic block diagram of a pallet detection device 300 provided by an embodiment of the present invention. As Figure 3 shown, corresponding to the above pallet detection method, the present invention also provides a pallet detection device 300. The pallet detection device 300 includes units for executing the above pallet detection method, and the device can be configured in a server. Specifically, please refer to Figure 3 , the pallet detection device 300 includes an acquisition unit 301, a preprocessing unit 302, a data fusion unit 303, an anomaly detection unit 304, and an adjustment unit 305.
[0192] The acquisition unit 301 is used to acquire multi-source data collected by different sensors for real-time monitoring of the device status; the preprocessing unit 302 is used to preprocess the multi-source data to obtain a preprocessing result; the data fusion unit 303 is used to perform data fusion on the preprocessing result to obtain a fusion result; the anomaly detection unit 304 is used to perform pallet anomaly detection using a reinforcement learning algorithm based on the fusion result to obtain a detection result; the adjustment unit 305 is used to adjust the device parameters when the detection result is a pallet abnormal state to adjust the device operation state in real time.
[0193] In one embodiment, as Figure 4 shown, the preprocessing unit 302 includes a filtering subunit 3021, an alignment subunit 3022, and a feature extraction subunit 3023.
[0194] The filtering subunit 3021 is used to perform denoising filtering on the multi-source data to obtain a filtering result; the alignment subunit 3022 is used to perform timestamp alignment on the filtering result to obtain an alignment result; the feature extraction subunit 3023 is used to extract features from the alignment result to obtain a preprocessing result.
[0195] In one embodiment, the data fusion unit 303 is used to perform data fusion on the preprocessing result by using Kalman filtering and covariance matrix weighting technology to obtain a fusion result.
[0196] In one embodiment, the anomaly detection unit 304 is used to input the fusion result into an anomaly detection model for pallet anomaly detection to obtain a detection result; wherein, the anomaly detection model is obtained by training a Q-Learning strategy by defining states, device actions, and reward functions.
[0197] It should be noted that those skilled in the art can clearly understand that the specific implementation processes of the above pallet detection device 300 and each unit can refer to the corresponding descriptions in the foregoing method embodiments. For the sake of convenience and brevity of description, they will not be repeated here.
[0198] The above pallet detection device 300 can be implemented in the form of a computer program, which can run on a computer device as shown in Figure 5 .
[0199] Please refer to Figure 5 , Figure 5 which is a schematic block diagram of a computer device provided by an embodiment of the present application. The computer device 500 can be a server. Among them, the server can be an independent server or a server cluster composed of multiple servers.
[0200] Referring to Figure 5 , the computer device 500 includes a processor 502, a memory, and a network interface 505 connected through a system bus 501. Among them, the memory can include a non-volatile storage medium 503 and an internal memory 504.
[0201] The non-volatile storage medium 503 can store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions, which when executed, can cause the processor 502 to execute a pallet detection method.
[0202] The processor 502 is used to provide computing and control capabilities to support the operation of the entire computer device 500.
[0203] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can be caused to execute a pallet detection method.
[0204] The network interface 505 is used for network communication with other devices. Those skilled in the art can understand that Figure 5 the structure shown in
[0205] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device 500 to which the solution of the present application is applied. The specific computer device 500 may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0206] Obtain multi-source data collected by different sensors for real-time monitoring of the device status; preprocess the multi-source data to obtain a preprocessing result; perform data fusion on the preprocessing result to obtain a fusion result; use a reinforcement learning algorithm to perform pallet anomaly detection based on the fusion result to obtain a detection result; when the detection result is a pallet abnormal state, adjust the device parameters to real-time adjust the device operation state.
[0207] Among them, the sensors include vibration sensors, photoelectric sensors, pressure sensors, temperature sensors, and encoders.
[0208] The preprocessing results include time-domain features, frequency-domain features, features obtained from time-frequency analysis, and statistical features.
[0209] In one embodiment, when the processor 502 implements the step of preprocessing the multi-source data to obtain preprocessing results, the following steps are specifically implemented:
[0210] Perform denoising filtering on the multi-source data to obtain a filtering result; perform timestamp alignment on the filtering result to obtain an alignment result; extract features from the alignment result to obtain preprocessing results.
[0211] In one embodiment, when the processor 502 implements the step of performing data fusion on the preprocessing results to obtain fusion results, the following steps are specifically implemented:
[0212] Use Kalman filtering and covariance matrix weighting techniques to perform data fusion on the preprocessing results to obtain fusion results.
[0213] In one embodiment, when the processor 502 implements the step of using a reinforcement learning algorithm to perform pallet abnormality detection based on the fusion results to obtain detection results, the following steps are specifically implemented:
[0214] Input the fusion results into an abnormality detection model to perform pallet abnormality detection to obtain detection results; among them, the abnormality detection model is obtained by training a Q-Learning strategy by defining states, device actions, and reward functions.
[0215] In one embodiment, when the processor 502 implements the step that the abnormality detection model is obtained by training a Q-Learning strategy by defining states, device actions, and reward functions, the following steps are specifically implemented:
[0216] The abnormality detection model is a model formed by using interactive learning by defining states, device actions, and reward functions. The agent executes actions in the environment and collects experience, uses the ε-greedy strategy to balance exploration and exploitation, updates the strategy by minimizing the temporal difference error, and optimizes the reward function to maximize the cumulative reward.
[0217] It should be understood that in the embodiments of the present application, the processor 502 may be a central processing unit (CPU), and the processor 502 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0218] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program includes program instructions, and the computer program can be stored in a storage medium, and the storage medium is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.
[0219] Therefore, the present invention also provides a storage medium. The storage medium may be a computer-readable storage medium. The storage medium stores a computer program, and when the computer program is executed by a processor, the processor executes the following steps:
[0220] Obtain multi-source data collected by different sensors for real-time monitoring of the device status; preprocess the multi-source data to obtain a preprocessing result; perform data fusion on the preprocessing result to obtain a fusion result; use a reinforcement learning algorithm to perform pallet abnormality detection based on the fusion result to obtain a detection result; when the detection result is a pallet abnormal state, adjust the device parameters to adjust the device operation state in real time.
[0221] Among them, the sensors include vibration sensors, photoelectric sensors, pressure sensors, temperature sensors, and encoders.
[0222] The preprocessing result includes time-domain features, frequency-domain features, features obtained by time-frequency analysis, and statistical features.
[0223] In one embodiment, when the processor executes the computer program to implement the step of preprocessing the multi-source data to obtain a preprocessing result, the following steps are specifically implemented:
[0224] Denoise and filter the multi-source data to obtain a filtering result; align the timestamps of the filtering result to obtain an alignment result; extract features from the alignment result to obtain a preprocessing result.
[0225] In one embodiment, when the processor executes the computer program to implement the step of performing data fusion on the preprocessing result to obtain a fusion result, the specific implementation is as follows:
[0226] Use Kalman filtering and covariance matrix weighting techniques to perform data fusion on the preprocessing result to obtain a fusion result.
[0227] In one embodiment, when the processor executes the computer program to implement the step of using a reinforcement learning algorithm to perform pallet abnormality detection based on the fusion result to obtain a detection result, the specific implementation is as follows:
[0228] Input the fusion result into an abnormality detection model to perform pallet abnormality detection to obtain a detection result;
[0229] Among them, the abnormality detection model is obtained by training a Q-Learning strategy by defining states, device actions, and reward functions.
[0230] In one embodiment, when the processor executes the computer program to implement the step that the abnormality detection model is obtained by training a Q-Learning strategy by defining states, device actions, and reward functions, the specific implementation is as follows:
[0231] The abnormality detection model is a model formed by defining states, device actions, and reward functions, adopting interactive learning, where the agent executes actions in the environment and collects experience, using the ε-greedy strategy to balance exploration and exploitation, updating the strategy by minimizing the temporal difference error, and optimizing the reward function to maximize the cumulative reward.
[0232] The storage medium can be various computer-readable storage media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disc that can store program codes.
[0233] Those of ordinary skill in the art will realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of the examples have been generally described in terms of function in the above description. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0234] In several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of each unit is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0235] The steps in the method embodiments of the present invention can be adjusted, combined, and deleted according to actual needs. The units in the device embodiments of the present invention can be combined, divided, and deleted according to actual needs. In addition, the functional units in each embodiment of the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0236] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention.
[0237] As described above, the above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A card board detection method, characterized in that: include: Acquire multi-source data collected by different sensors to monitor the status of equipment in real time; Preprocessing the multi-source data to obtain a preprocessing result; Performing data fusion on the preprocessing results to obtain a fusion result; Using a reinforcement learning algorithm to perform card board abnormality detection according to the fusion result to obtain a detection result; When the detection result is that the card board is in an abnormal state, the device parameters are adjusted to adjust the device operation state in real time.
2. The card board detection method according to claim 1, characterized in that: The sensors include vibration sensors, photoelectric sensors, pressure sensors, temperature sensors and encoders.
3. The card board detection method according to claim 1, characterized in that: The preprocessing of the multi-source data to obtain a preprocessing result includes: Performing denoising filtering on the multi-source data to obtain a filtering result; Performing timestamp alignment on the filtering results to obtain an alignment result; Features are extracted from the alignment result to obtain a preprocessing result.
4. The card board detection method according to claim 1, characterized in that: The preprocessing results include time domain features, frequency domain features, features obtained from time-frequency analysis, and statistical features.
5. The card board detection method according to claim 1, characterized in that: The performing data fusion on the preprocessing results to obtain a fusion result includes: Kalman filtering and covariance matrix weighting technology are used to perform data fusion on the preprocessing results to obtain a fusion result.
6. The card board detection method according to claim 1, characterized in that: The method of using a reinforcement learning algorithm to perform card board abnormality detection according to the fusion result to obtain a detection result includes: Inputting the fusion result into an anomaly detection model to perform card board anomaly detection to obtain a detection result; The anomaly detection model is obtained by training the Q-Learning strategy by defining states, device actions and reward functions.
7. The card board detection method according to claim 6, characterized in that: The anomaly detection model is obtained by training the Q-Learning strategy by defining the state, device action and reward function, including: The anomaly detection model adopts interactive learning by defining states, device actions and reward functions. The agent performs actions in the environment and collects experience. The ε-greedy strategy is used to balance exploration and utilization. The strategy is updated by minimizing the temporal difference error and the reward function is optimized to maximize the cumulative reward.
8. A cardboard detection device, characterized in that: include: An acquisition unit, used to acquire multi-source data collected by different sensors to monitor the status of equipment in real time; A preprocessing unit, used for preprocessing the multi-source data to obtain a preprocessing result; A data fusion unit, used for performing data fusion on the preprocessing results to obtain a fusion result; An anomaly detection unit, used to perform card board anomaly detection using a reinforcement learning algorithm according to the fusion result to obtain a detection result; The adjustment unit is used to adjust the equipment parameters when the detection result is that the card board is in an abnormal state, so as to adjust the equipment operation state in real time.
9. A computer device, characterized in that: The computer device comprises a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
10. A storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.