Intelligent control method and system for shield machine main drive based on Internet of Things
Through the Internet of Things and deep reinforcement learning algorithm, the mapping model of the shield machine stratigraphic parameters and toolbar speed and torque is established, which solves the excavation efficiency and energy consumption problems of traditional shield machine under complex geological conditions, and realizes the intelligent control of shield machine and the extension of equipment life.
Patent Information
- Application Number
- CN202510805865.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2045-06-17
AI Technical Summary
The traditional shield machine main drive system lacks the ability to adapt to complex geological conditions and comprehensive analysis of multi-dimensional data, resulting in low excavation efficiency or excessive energy consumption.
A multi-dimensional shield machine stratigraphic boring parameter database based on the Internet of Things is adopted, combined with deep reinforcement learning algorithms and model prediction control algorithms, a mapping relationship model between stratigraphic parameters and toolbar speed and torque is established. Through state space construction, action space selection and reward function calculation, dynamic adjustment of toolbar speed and torque is realized, and the excavation process of shield machine is optimized.
It improves the excavation efficiency and energy consumption balance of the shield machine under complex geological conditions, extends the service life of the equipment, and improves construction safety and reliability.
Smart Images

Figure CN120312244B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of shield machines, and in particular to an intelligent control method and system for a main drive of a shield machine based on the Internet of Things. Background Art
[0002] During the traditional shield machine construction process, the control of the main drive system mainly relies on manual adjustments by operators based on experience, or a simple feedback control system is used to adjust basic parameters.
[0003] With the development of Internet of Things technology, various sensors can monitor the working status of shield machines in real time, providing a data basis for the intelligent control of the shield machine's main drive system. However, in practical applications, the existing technology has the following defects and shortcomings:
[0004] First, existing shield machine main drive system control methods lack adaptability to complex geological conditions. Shield machine excavation parameters vary significantly under different geological conditions, and traditional fixed parameter control cannot dynamically adjust to geological changes. This results in low excavation efficiency or excessive energy consumption in areas with complex geological conditions.
[0005] Secondly, traditional shield machine main drive systems lack the ability to comprehensively analyze multi-dimensional data. Existing systems typically only consider a single or a small number of parameters for simple feedback control, failing to fully utilize multimodal data spanning time, geological characteristics, and tunneling parameters, making it difficult to fully optimize the tunneling process. Summary of the Invention
[0006] The embodiments of the present invention provide an intelligent control method and system for the main drive of a shield machine based on the Internet of Things, which can solve the problems in the prior art.
[0007] A first aspect of an embodiment of the present invention provides an intelligent control method for a main drive of a shield machine based on the Internet of Things, comprising:
[0008] Acquire real-time monitoring data from the multimodal sensors of the shield machine's main drive system, and combine this with geological parameter information of the area where the shield machine is located to establish a multi-dimensional shield machine stratum excavation parameter database that includes time dimensions, geological feature dimensions, and excavation parameter dimensions.
[0009] According to the multi-dimensional shield machine stratum excavation parameter database, a deep reinforcement learning algorithm is used to establish a mapping relationship model between stratum parameters and cutterhead speed and cutterhead torque. The deep reinforcement learning algorithm includes a state space construction module, an action space selection module and a reward function calculation module, wherein:
[0010] The state space construction module constructs a current excavation state vector based on the real-time monitoring data; the action space selection module selects an optimal cutterhead speed within a preset speed range based on the current excavation state vector; the reward function calculation module calculates a reward value based on a ratio of excavation efficiency per unit time to energy consumption, and feeds the reward value back to the action space selection module for optimizing the selection strategy of the optimal cutterhead speed;
[0011] Based on the optimal cutterhead speed, the optimal torque output value is calculated based on the model predictive control algorithm to achieve dynamic adjustment of the cutterhead speed; based on the optimal torque output value, a frequency converter control instruction sequence is generated, and the frequency converter control instruction sequence is used to adjust the output voltage and current parameters of the frequency converter to achieve dynamic balance control of the torque of the shield machine main drive system.
[0012] The action space selection module selects the optimal cutterhead speed within a preset speed range based on the current tunneling state vector, including:
[0013] Constructing a speed candidate set within a preset speed range, wherein the upper limit of the preset speed range is determined based on the mechanical safety speed of the shield machine cutterhead, the lower limit of the preset speed range is determined based on the minimum tunneling efficiency of the shield machine, and the speed value interval in the speed candidate set is determined based on the control accuracy of the shield machine;
[0014] Based on the current excavation state vector, the excavation distance per unit time is calculated to obtain an excavation efficiency index and the energy consumption per unit excavation distance is calculated to obtain an energy consumption index;
[0015] The ratio of the excavation efficiency index to the standard excavation efficiency is recorded as a first score; the ratio of the standard energy consumption to the energy consumption index is recorded as a second score; and the weighted sum of the first score and the second score is recorded as the fitness score;
[0016] The fitness score of each speed value in the speed candidate set is obtained, and a speed value with the highest fitness score is randomly selected as the optimal cutterhead speed from the speed value set according to the ε greedy strategy.
[0017] The reward function calculation module calculates a reward value based on the ratio of excavation efficiency to energy consumption per unit time, and feeds the reward value back to the action space selection module for optimizing the selection strategy of the optimal cutterhead speed, including:
[0018] Calculating the tunneling efficiency per unit time based on the tunneling time and tunneling distance of the shield machine, and calculating the energy consumption per unit tunneling distance based on the power of the main motor of the shield machine; dividing the tunneling efficiency per unit time by the energy consumption per unit tunneling distance to obtain an efficiency ratio;
[0019] When the performance ratio is greater than a preset benchmark performance ratio, a positive incentive value is calculated based on the difference between the two; when the performance ratio is less than the preset benchmark performance ratio, a negative inhibition value is calculated based on the difference between the two; the positive incentive value or the negative inhibition value is used as a performance evaluation value;
[0020] Dynamically adjusting the cutterhead speed using an adaptive variable step size algorithm based on the performance evaluation value, calculating a speed adjustment amount, and using the speed adjustment amount as a reward value;
[0021] The reward value is fed back to the action space selection module to optimize the selection strategy of the optimal cutterhead speed.
[0022] Based on the performance evaluation value, the cutterhead speed is dynamically adjusted using an adaptive variable step size algorithm. The speed adjustment amount is calculated by:
[0023] Determine the rate of change of the performance evaluation value within a preset sampling period,
[0024] When the change rate and the positive excitation value have the same sign, the speed adjustment amount is the square root of the product of the change rate and the positive excitation value; when the change rate and the negative inhibition value have the same sign, the speed adjustment amount is the absolute value of the product of the change rate and the negative inhibition value;
[0025] When the rate of change has the same sign as the positive excitation value, the speed adjustment amount is added to the current cutterhead speed until a local optimal point is detected, wherein determining the local optimal point includes:
[0026] Continuously collect the efficiency ratios of three adjacent adjustment points; calculate the efficiency ratio difference of adjacent adjustment points; when the absolute value of the latter difference is less than the absolute value of the former difference, record the middle point as the peak point. When three peak points are obtained continuously, select the peak point with the largest efficiency ratio as the local optimal point.
[0027] Based on the optimal cutterhead speed, the optimal torque output value is calculated based on the model predictive control algorithm to achieve dynamic adjustment of the cutterhead speed, including:
[0028] Determining a predicted total duration based on the optimal cutterhead speed, the predicted total duration being the shortest time required for the optimal cutterhead speed to stably operate; dividing the predicted total duration into a plurality of control periods, the duration of each control period being proportional to the inverse of the optimal cutterhead speed;
[0029] setting a speed regulation acceleration upper limit value in each control period, wherein the speed regulation acceleration upper limit value is inversely proportional to the duration of the control period; and calculating a speed regulation amount in each control period according to the speed regulation acceleration upper limit value;
[0030] Calculating the deviation between the actual cutter head torque and the preset reference torque; constructing an optimization objective function based on a weighted sum of the deviation and the speed adjustment amount; determining a maximum torque limit based on the rated torque of the equipment and a torque change rate limit based on the motor characteristics; and solving the optimization objective function using a gradient projection algorithm under the constraints of the maximum torque limit and the torque change rate limit to obtain an optimal torque adjustment sequence;
[0031] The first value of the optimal torque adjustment sequence is extracted as the initial torque adjustment amount; a smoothing factor is determined based on the torque adjustment sensitivity, and the smoothing factor is used to suppress torque mutations; the product of the initial torque adjustment amount and the smoothing factor is used as the torque output increment; and the torque output increment is superimposed on the current torque to obtain the optimal torque output value.
[0032] Solving the optimization objective function by a gradient projection algorithm under the constraints of the maximum torque limit and the torque change rate limit to obtain an optimal torque adjustment sequence includes:
[0033] Setting a prediction time domain length with the current moment as the starting point, and dividing the prediction time domain length into a plurality of control periods at equal intervals; calculating an allowable torque adjustment range in each of the control periods, and using the allowable torque adjustment range as a local search space;
[0034] In each of the control periods, calculating a gradient vector of the optimization objective function; projecting the gradient vector to the boundary of the local search space;
[0035] Calculate the angle between the projected gradient direction and the boundary normal vector; adaptively adjust the search step size according to the angle; iteratively update the torque value along the projection direction; record the torque value of each iteration;
[0036] When the torque value change rate of two adjacent iterations is less than a preset change rate threshold, the current torque value is determined as the optimal torque value of the control period; the optimal torque values of each control period are combined into an optimal torque adjustment sequence.
[0037] A second aspect of the embodiments of the present invention provides an Internet of Things-based shield machine main drive intelligent control system, including:
[0038] The first unit is used to obtain real-time monitoring data from the multimodal sensors of the shield machine's main drive system, and to establish a multi-dimensional shield machine stratum excavation parameter database that includes time dimensions, geological feature dimensions, and excavation parameter dimensions, combined with geological parameter information of the area where the shield machine is located.
[0039] The second unit is used to establish a mapping relationship model between formation parameters and cutterhead speed and cutterhead torque using a deep reinforcement learning algorithm based on the multi-dimensional shield machine formation excavation parameter database. The deep reinforcement learning algorithm includes a state space construction module, an action space selection module and a reward function calculation module, wherein:
[0040] a third unit, configured for the state space construction module to construct a current tunneling state vector based on the real-time monitoring data, the action space selection module to select an optimal cutterhead speed within a preset speed range based on the current tunneling state vector, the reward function calculation module to calculate a reward value based on a ratio of tunneling efficiency to energy consumption per unit time, and to feed the reward value back to the action space selection module for optimizing a selection strategy for the optimal cutterhead speed;
[0041] The fourth unit is used to calculate the optimal torque output value based on the optimal cutterhead speed and the model predictive control algorithm to achieve dynamic adjustment of the cutterhead speed; according to the optimal torque output value, generate a frequency converter control instruction sequence, and the frequency converter control instruction sequence is used to adjust the output voltage and current parameters of the frequency converter to achieve dynamic balance control of the torque of the shield machine main drive system.
[0042] According to a third aspect of an embodiment of the present invention, an electronic device is provided, including:
[0043] processor;
[0044] a memory for storing processor-executable instructions;
[0045] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0046] According to a fourth aspect of an embodiment of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described above is implemented.
[0047] The beneficial effects of this application are as follows:
[0048] The present invention uses Internet of Things technology to collect multimodal sensor data and geological parameter information of the shield machine's main drive system, establishes a multidimensional database including time dimension, geological feature dimension and tunneling parameter dimension, realizes comprehensive perception and data support of the shield machine's working environment, and provides a solid foundation for intelligent detection.
[0049] The present invention adopts a deep reinforcement learning algorithm to establish a mapping relationship model between formation parameters and cutterhead speed and cutterhead torque. Through the coordinated work of three key modules: state space construction, action space selection and reward function calculation, it realizes the intelligent optimization selection of cutterhead speed under complex geological conditions, balances the relationship between excavation efficiency and energy consumption, and improves the operating efficiency of the shield machine.
[0050] The present invention calculates the optimal torque output value based on the model predictive control algorithm and generates a frequency converter control instruction sequence, thereby realizing dynamic balance control of the torque of the shield machine's main drive system, effectively reducing the equipment failure rate, extending the service life of the main drive system, and improving the adaptability and reliability of the shield machine under complex geological conditions. It is of great significance for improving the safety of urban underground engineering construction. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 Schematic diagram of the process of the intelligent control method of the shield machine main drive based on the Internet of Things according to an embodiment of the present invention;
[0052] Figure 2 The simulation diagram of the optimal cutterhead speed selection strategy;
[0053] Figure 3 This is a schematic diagram of the cutterhead speed adjustment performance curve of the adaptive variable step size algorithm;
[0054] Figure 4 Schematic diagram of the torque optimization performance comparison of the gradient projection algorithm. DETAILED DESCRIPTION
[0055] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0056] The following specific embodiments are used to describe the technical solution of the present invention in detail. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0057] Figure 1 FIG. 1 is a flow chart of an intelligent control method for a shield machine main drive based on the Internet of Things according to an embodiment of the present invention. Figure 1 As shown, the method includes:
[0058] Acquire real-time monitoring data from the multimodal sensors of the shield machine's main drive system, and combine this with geological parameter information of the area where the shield machine is located to establish a multi-dimensional shield machine stratum excavation parameter database that includes time dimensions, geological feature dimensions, and excavation parameter dimensions.
[0059] According to the multi-dimensional shield machine stratum excavation parameter database, a deep reinforcement learning algorithm is used to establish a mapping relationship model between stratum parameters and cutterhead speed and cutterhead torque. The deep reinforcement learning algorithm includes a state space construction module, an action space selection module and a reward function calculation module, wherein:
[0060] The state space construction module constructs a current excavation state vector based on the real-time monitoring data; the action space selection module selects an optimal cutterhead speed within a preset speed range based on the current excavation state vector; the reward function calculation module calculates a reward value based on a ratio of excavation efficiency per unit time to energy consumption, and feeds the reward value back to the action space selection module for optimizing the selection strategy of the optimal cutterhead speed;
[0061] Based on the optimal cutterhead speed, the optimal torque output value is calculated based on the model predictive control algorithm to achieve dynamic adjustment of the cutterhead speed; based on the optimal torque output value, a frequency converter control instruction sequence is generated, and the frequency converter control instruction sequence is used to adjust the output voltage and current parameters of the frequency converter to achieve dynamic balance control of the torque of the shield machine main drive system.
[0062] In an optional embodiment, the action space selection module selects the optimal cutterhead speed within a preset speed range based on the current tunneling state vector, including:
[0063] Constructing a speed candidate set within a preset speed range, wherein the upper limit of the preset speed range is determined based on the mechanical safety speed of the shield machine cutterhead, the lower limit of the preset speed range is determined based on the minimum tunneling efficiency of the shield machine, and the speed value interval in the speed candidate set is determined based on the control accuracy of the shield machine;
[0064] Based on the current excavation state vector, the excavation distance per unit time is calculated to obtain an excavation efficiency index and the energy consumption per unit excavation distance is calculated to obtain an energy consumption index;
[0065] The ratio of the excavation efficiency index to the standard excavation efficiency is recorded as a first score; the ratio of the standard energy consumption to the energy consumption index is recorded as a second score; and the weighted sum of the first score and the second score is recorded as the fitness score;
[0066] The fitness score of each speed value in the speed candidate set is obtained, and a speed value with the highest fitness score is randomly selected as the optimal cutterhead speed from the speed value set according to the ε greedy strategy.
[0067] In a specific technical solution for implementing the action space selection module to select the optimal cutterhead speed within a preset speed range based on the current tunneling state vector, it is first necessary to construct a speed candidate set within the preset speed range. The upper limit of the preset speed range is determined based on the mechanical safety speed of the shield machine cutterhead. For example, when the mechanical safety speed of the shield machine cutterhead is 3.5 rpm, the upper limit of the preset speed range can be set to 3.0 rpm to ensure a safety margin. The lower limit of the preset speed range is determined based on the minimum tunneling efficiency of the shield machine. For example, when the minimum tunneling efficiency requirement of the shield machine is 20 mm / min, the corresponding minimum cutterhead speed may be 0.5 rpm, so the lower limit of the preset speed range is set to 0.5 rpm. The speed value interval in the speed candidate set is determined based on the control accuracy of the shield machine. For example, when the control accuracy of the shield machine is 0.1 rpm, the speed value interval can be set to 0.1 rpm. In this way, within the speed range [0.5, 3.0], with an interval of 0.1 rpm, a speed candidate set {0.5, 0.6, 0.7, ..., 2.9, 3.0} is constructed, which contains a total of 26 candidate speed values.
[0068] After constructing the speed candidate set, the excavation efficiency index is calculated based on the current excavation state vector, as is the excavation distance per unit time. The energy consumption index is also calculated based on the energy consumption per unit excavation distance. The excavation state vector includes formation parameters, equipment parameters, and control parameters. For example, when the cutterhead speed is 2.0 rpm, the propulsion speed is 40 mm / min, the soil pressure is 200 kPa, and the cutterhead torque is 800 kN·m, the excavation efficiency index can be calculated as 40 mm / min using the excavation model. This excavation model can be a machine learning model trained based on historical data or a mechanical model based on physical mechanisms. The energy consumption index can be calculated by dividing the sum of the cutterhead motor power and the propulsion system power by the excavation speed. For example, when the cutterhead motor power is 450 kW, the propulsion system power is 120 kW, and the excavation speed is 40 mm / min, the energy consumption index is (450 + 120) ÷ 40 = 14.25 kW·min / mm.
[0069] The ratio of the tunneling efficiency index to the standard tunneling efficiency is recorded as the first score. Assuming the standard tunneling efficiency is 50 mm / min and the current tunneling efficiency index is 40 mm / min, the first score is 40 ÷ 50 = 0.8. The ratio of the standard energy consumption to the energy consumption index is recorded as the second score. Assuming the standard energy consumption is 10 kW·min / mm and the current energy consumption index is 14.25 kW·min / mm, the second score is 10 ÷ 14.25 = 0.702. The weighted sum of the first and second scores is the fitness score. For example, if the weights for tunneling efficiency and energy consumption are 0.6 and 0.4, respectively, the fitness score is 0.6 × 0.8 + 0.4 × 0.702 = 0.761.
[0070] To obtain the fitness score for each speed value in the candidate speed set, the above steps are used to calculate the corresponding fitness score for each speed value in the candidate speed set, using the same parameters in the current tunneling state vector and the speed value. For example, for a speed value of 1.8 rpm, the calculated tunneling efficiency is 36 mm / min and the energy consumption index is 13 kW·min / mm. Thus, the first score is 36 ÷ 50 = 0.72, the second score is 10 ÷ 13 = 0.769, and the fitness score is 0.6 × 0.72 + 0.4 × 0.769 = 0.740. For a speed value of 2.2 rpm, the calculated tunneling efficiency is 43 mm / min and the energy consumption index is 15.5 kW·min / mm. Thus, the first score is 43 ÷ 50 = 0.86, the second score is 10 ÷ 15.5 = 0.645, and the fitness score is 0.6 × 0.86 + 0.4 × 0.645 = 0.774. And so on, the fitness scores of all speed values are calculated.
[0071] After the fitness scores for all speed values are calculated, the ε-greedy strategy is used to randomly select a speed value from the set of speed values with the highest fitness score as the optimal cutterhead speed. Specifically, first find the set of speed values with the highest fitness score. For example, if the speed value 2.2 rpm has the highest fitness score of 0.774, then the set of speed values with the highest fitness score is {2.2}; if the speed values 2.2 rpm and 2.3 rpm both have the highest fitness score of 0.774, then the set of speed values with the highest fitness score is {2.2, 2.3}. Then, using the ε-greedy strategy, a speed value is randomly selected from this set with probability 1-ε as the optimal cutterhead speed, and a speed value is randomly selected from the entire set of speed candidates with probability ε as the optimal cutterhead speed. For example, when ε = 0.1, there is a 0.9 probability of selecting 2.2 rpm from {2.2} as the optimal cutterhead speed; and there is a 0.1 probability of randomly selecting a speed value from {0.5, 0.6, 0.7, ..., 2.9, 3.0} as the optimal cutterhead speed. This strategy strikes a balance between utilizing the current optimal solution and exploring potential better solutions, improving the robustness and adaptability of the algorithm.
[0072] Through the above technical solution, the action space selection module can select the optimal cutterhead speed within the preset speed range based on the current tunneling state vector, thereby realizing intelligent control of the shield machine's tunneling process, improving tunneling efficiency and reducing energy consumption.
[0073] In an optional embodiment, the reward function calculation module calculates a reward value based on the ratio of excavation efficiency per unit time to energy consumption, and feeds the reward value back to the action space selection module for optimizing the selection strategy of the optimal cutterhead speed, including:
[0074] Calculating the tunneling efficiency per unit time based on the tunneling time and tunneling distance of the shield machine, and calculating the energy consumption per unit tunneling distance based on the power of the main motor of the shield machine; dividing the tunneling efficiency per unit time by the energy consumption per unit tunneling distance to obtain an efficiency ratio;
[0075] When the performance ratio is greater than a preset benchmark performance ratio, a positive incentive value is calculated based on the difference between the two; when the performance ratio is less than the preset benchmark performance ratio, a negative inhibition value is calculated based on the difference between the two; the positive incentive value or the negative inhibition value is used as a performance evaluation value;
[0076] Dynamically adjusting the cutterhead speed using an adaptive variable step size algorithm based on the performance evaluation value, calculating a speed adjustment amount, and using the speed adjustment amount as a reward value;
[0077] The reward value is fed back to the action space selection module to optimize the selection strategy of the optimal cutterhead speed.
[0078] This embodiment provides a method for dynamically optimizing the cutterhead speed of a shield machine based on a reward function calculation module. The reward function calculation module calculates the ratio of tunneling efficiency per unit time to energy consumption to determine a reward value. This reward value is then fed back to the action space selection module, thereby optimizing the strategy for selecting the optimal cutterhead speed.
[0079] During implementation, the system first acquires the shield machine's tunneling time and distance data. For example, during a given tunneling cycle, the shield machine recorded a tunneling time of 2 hours and a distance of 4 meters. Based on this data, the system calculates the tunneling efficiency per unit time as 4 meters / 2 hours = 2 meters / hour. This efficiency value reflects the shield machine's tunneling capacity per unit time and is a key indicator for evaluating tunneling efficiency.
[0080] At the same time, the system acquires power data for the shield machine's main motor. During the aforementioned tunneling cycle, the main motor's average power was 800 kilowatts. Based on the main motor power and tunneling distance, the system calculates the energy consumption per unit of tunneling distance as 800 kilowatts x 2 hours / 4 meters = 400 kilowatt-hours / meter. This value represents the amount of electricity consumed per meter of tunneling by the shield machine and is a key indicator for evaluating energy efficiency.
[0081] Based on these calculations, the system divides the tunneling efficiency per unit time by the energy consumption per unit distance traveled to calculate the efficiency ratio. In this example, the efficiency ratio is 2 meters / hour ÷ 400 kWh / meter = 0.005 hours / kWh. The efficiency ratio takes into account both tunneling efficiency and energy consumption; a higher value indicates better overall shield machine performance.
[0082] The system pre-sets a baseline efficiency ratio for evaluating current tunneling performance. Assume the baseline efficiency ratio is 0.004 hours / kWh. By comparing the current efficiency ratio with the baseline efficiency ratio, the system determines whether to calculate a positive incentive or a negative disincentive. In this example, the current efficiency ratio of 0.005 is greater than the baseline efficiency ratio of 0.004. The system calculates a positive incentive value of (0.005 - 0.004) / 0.004 × 100% = 25%. This indicates that current tunneling performance is 25% higher than the baseline level, and the system should provide a positive reward.
[0083] If, in another case, the calculated efficiency ratio is 0.003 hours / kWh, which is lower than the baseline efficiency ratio of 0.004 hours / kWh, the system will calculate a negative suppression value of (0.003-0.004) / 0.004×100%=-25%. This means that the current excavation performance is 25% lower than the baseline level, and the system should apply a negative suppression.
[0084] The system uses the calculated positive excitation value or negative inhibition value as the performance evaluation value for subsequent cutterhead speed adjustment. In this example, the performance evaluation value is 25%.
[0085] Based on this performance evaluation value, the system uses an adaptive variable step size algorithm to dynamically adjust the cutterhead speed. The core concept of the adaptive variable step size algorithm is to dynamically adjust the step size coefficient based on the performance evaluation value. When the performance evaluation value is large, a larger step size is used to accelerate convergence; when the performance evaluation value is small, a smaller step size is used to ensure accuracy.
[0086] The system sets a base step size coefficient of 0.05 rpm and a maximum step size limit of 2 rpm. Based on the performance evaluation value of 25%, the system calculates the actual step size to be 0.05 × 25% = 0.0125 rpm. Since this step size is less than the maximum step size limit, the system adopts it.
[0087] Assume the current cutterhead speed is 3 rpm and the performance evaluation is positive, indicating that the current speed setting is effective and the system should continue optimizing in this direction. The system adjusts the cutterhead speed to 3 + 0.0125 = 3.0125 rpm. This adjustment of 0.0125 rpm is the calculated reward value and is fed back to the action space selection module.
[0088] In another scenario, if the performance evaluation is a negative suppression value of -25%, the system calculates an actual step size of 0.05 × |-25%| = 0.0125 rpm. Because the performance evaluation is negative, the system adjusts the cutterhead speed to 3 - 0.0125 = 2.9875 rpm. This adjustment of -0.0125 rpm is the calculated reward value and is also fed back to the action space selection module.
[0089] After receiving the reward, the action space selection module updates its cutterhead speed selection strategy based on reinforcement learning principles. When receiving a positive reward, the system increases the probability of selecting a speed near that value; when receiving a negative reward, the system decreases the probability of selecting a speed near that value. Through multiple iterative learning cycles, the system gradually optimizes the cutterhead speed selection strategy, ultimately achieving optimal control of the cutterhead speed.
[0090] Figure 2 The simulation diagram of the optimal cutterhead speed selection strategy is shown in Figure 2. Figure 2 As shown in the figure, for different combinations of geological conditions and cutterhead speeds, the ratio of tunneling efficiency per unit time to energy consumption (efficiency ratio) is calculated and corresponding reward values are generated. As can be seen from the figure, when the efficiency ratio is above the baseline value, the system generates positive incentives; when the efficiency ratio is below the baseline value, the system generates negative inhibitions. Through 1000 random samplings, the relationship between cutterhead speed, efficiency ratio, and reward values can be clearly seen.
[0091] In practice, the system continuously monitors the shield machine's tunneling parameters, including tunneling time, tunneling distance, and main motor power, and continuously updates performance ratio calculations and reward value feedback, forming a closed-loop control system. This allows the shield machine to adaptively adjust cutterhead speed under varying geological conditions, achieving the optimal balance between tunneling efficiency and energy consumption, thereby improving overall tunneling performance.
[0092] In an optional embodiment, the cutterhead speed is dynamically adjusted using an adaptive variable step size algorithm based on the performance evaluation value, and calculating the speed adjustment amount includes:
[0093] Determine the rate of change of the performance evaluation value within a preset sampling period,
[0094] When the change rate and the positive excitation value have the same sign, the speed adjustment amount is the square root of the product of the change rate and the positive excitation value; when the change rate and the negative inhibition value have the same sign, the speed adjustment amount is the absolute value of the product of the change rate and the negative inhibition value;
[0095] When the rate of change has the same sign as the positive excitation value, the speed adjustment amount is added to the current cutterhead speed until a local optimal point is detected, wherein determining the local optimal point includes:
[0096] Continuously collect the efficiency ratios of three adjacent adjustment points; calculate the efficiency ratio difference of adjacent adjustment points; when the absolute value of the latter difference is less than the absolute value of the former difference, record the middle point as the peak point. When three peak points are obtained continuously, select the peak point with the largest efficiency ratio as the local optimal point.
[0097] This embodiment describes in detail a method for dynamically adjusting the cutterhead speed using an adaptive variable step size algorithm based on performance evaluation values. This method optimizes the cutterhead speed by calculating the speed adjustment amount, thereby improving tunneling efficiency.
[0098] Specifically, the system first acquires the current performance evaluation value, which can be an indicator of the machine's performance during tunneling, such as the ratio of tunneling speed to energy consumption. The system continuously collects performance evaluation values within a preset sampling period, which can be set to 5 seconds. For example, if the performance evaluation value collected by the system increases from 2.5 to 2.8 during a given sampling period, the rate of change is 0.3.
[0099] The system has pre-set positive excitation and negative inhibition values to adjust the direction and magnitude of cutterhead speed changes. In practice, the positive excitation value can be set to 0.5 and the negative inhibition value to -0.3. These parameters can be adjusted based on actual excavation conditions to accommodate varying geological conditions and excavation requirements.
[0100] The system calculates the speed adjustment amount based on the relationship between the rate of change of the performance evaluation value and the preset parameters. When the rate of change and the positive excitation value have the same sign, it means that the current speed adjustment direction is correct, and the system will further increase the adjustment force. At this time, the speed adjustment amount is calculated as the square root of the product of the rate of change and the positive excitation value. When the rate of change and the negative inhibition value have the same sign, it means that the current speed adjustment direction may not be conducive to performance improvement, and the system will reduce the adjustment force. At this time, the speed adjustment amount is calculated as the absolute value of the product of the rate of change and the negative inhibition value. For example, when the rate of change is -0.2 (negative value), which has the same negative sign as the negative inhibition value -0.3, the speed adjustment amount is |-0.2×(-0.3)|=0.06, and the system reduces the current cutter head speed by 0.06 rpm.
[0101] During the speed adjustment process, the system continuously monitors the changes in the performance evaluation value. When the rate of change has the same sign as the positive excitation value, the system adds the calculated speed adjustment amount to the current cutterhead speed and continues to perform this operation until a local optimal point is detected.
[0102] The system uses a specific algorithm to determine the local optimum. This method continuously collects the performance ratios of three adjacent adjustment points and then calculates the difference between these performance ratios. Assuming the system collects performance ratios of 3.2, 3.5, and 3.7 at three adjustment points, the first difference is 3.5 - 3.2 = 0.3, and the second difference is 3.7 - 3.5 = 0.2. Because the absolute value of the latter difference, 0.2, is smaller than the previous difference, 0.3, the system determines that the intermediate point, 3.5, is a peak.
[0103] The system continues this calculation and comparison until it obtains three consecutive peaks. For example, if the system obtains three peaks with efficiency ratios of 3.5, 3.8, and 3.6, the system selects the peak with the highest efficiency ratio, 3.8, as the local optimum and records the corresponding cutterhead speed as the optimal speed for the current geological conditions.
[0104] To improve the adaptability and stability of the system, in actual applications, the positive excitation value and negative inhibition value can be adjusted according to different geological conditions. For example, in soft rock formations, the positive excitation value can be set to a larger value (such as 0.8) to speed up the speed adjustment process; in hard rock formations, the positive excitation value can be set to a smaller value (such as 0.3) to avoid excessive speed adjustment and increased tool wear.
[0105] During operation, the system continuously adjusts the cutterhead speed to maintain a local optimum. When geological conditions change, the system automatically initiates a new round of adjustments to find a new local optimum. This dynamic adjustment mechanism adapts to complex and changing geological conditions, improving the efficiency and service life of the roadheader.
[0106] Figure 3 This is a schematic diagram of the cutterhead speed adjustment performance curve of the adaptive variable step size algorithm. Figure 3 The performance curve of the adaptive variable step-size algorithm during cutterhead speed adjustment is shown. The curve shows that the algorithm dynamically adjusts the cutterhead speed through a positive incentive and negative inhibition mechanism, gradually improving the system's efficiency and ultimately stabilizing it. Three peaks were identified at sampling points 5, 8, and 9, with sampling point 9 achieving an efficiency of 0.97, identifying it as a local optimum. At this point, the cutterhead speed was adjusted to 10.24 rpm, achieving optimal system performance.
[0107] Experimental results demonstrate that this adaptive variable-step-size algorithm can effectively respond to changing geological conditions, intelligently adjust cutterhead speed, and improve tunneling efficiency and equipment utilization. By continuously monitoring the rate of change of performance evaluation values and combining positive incentives with negative inhibition strategies, the system can quickly converge to a local optimum and avoid falling into local extremes.
[0108] In a test scenario, the system used the above method to optimize the speed of a certain type of roadheader. The cutterhead speed was initially set at 3.0 rpm. After multiple adaptive adjustments, it stabilized at 3.8 rpm. At this point, the tunneling efficiency ratio increased by approximately 22%, from an initial 2.5 to 3.05, while energy consumption was reduced by approximately 15%. This demonstrates the effectiveness of this method in practical applications and its ability to effectively improve the operating efficiency of roadheaders.
[0109] In an optional embodiment, based on the optimal cutterhead speed, the optimal torque output value is calculated based on a model predictive control algorithm to achieve dynamic adjustment of the cutterhead speed, including:
[0110] Determining a predicted total duration based on the optimal cutterhead speed, the predicted total duration being the shortest time required for the optimal cutterhead speed to stably operate; dividing the predicted total duration into a plurality of control periods, the duration of each control period being proportional to the inverse of the optimal cutterhead speed;
[0111] setting a speed regulation acceleration upper limit value in each control period, wherein the speed regulation acceleration upper limit value is inversely proportional to the duration of the control period; and calculating a speed regulation amount in each control period according to the speed regulation acceleration upper limit value;
[0112] Calculating the deviation between the actual cutter head torque and the preset reference torque; constructing an optimization objective function based on a weighted sum of the deviation and the speed adjustment amount; determining a maximum torque limit based on the rated torque of the equipment and a torque change rate limit based on the motor characteristics; and solving the optimization objective function using a gradient projection algorithm under the constraints of the maximum torque limit and the torque change rate limit to obtain an optimal torque adjustment sequence;
[0113] The first value of the optimal torque adjustment sequence is extracted as the initial torque adjustment amount; a smoothing factor is determined based on the torque adjustment sensitivity, and the smoothing factor is used to suppress torque mutations; the product of the initial torque adjustment amount and the smoothing factor is used as the torque output increment; and the torque output increment is superimposed on the current torque to obtain the optimal torque output value.
[0114] This embodiment provides a method for dynamically adjusting the cutterhead speed based on the optimal cutterhead speed and a model predictive control algorithm. After determining the optimal cutterhead speed, the method needs to calculate the corresponding optimal torque output value to achieve accurate dynamic adjustment of the cutterhead speed.
[0115] In this embodiment, the system first calculates a predicted total time based on the determined optimal cutterhead speed. This predicted total time represents the minimum time required for the cutterhead to adjust from its current speed to the optimal speed and stabilize. For example, if the current cutterhead speed is detected to be 1.5 rpm and the calculated optimal cutterhead speed is 3.0 rpm, the system may determine that a predicted total time of 120 seconds is required for the cutterhead to stabilize at the optimal speed.
[0116] Once the predicted total duration is determined, the system divides it into multiple control periods, with the duration of each period proportional to the inverse of the optimal cutterhead speed. In practice, if the optimal cutterhead speed is 3.0 rpm, and each revolution takes 20 seconds, the system can divide the 120-second predicted total duration into six control periods, each approximately 20 seconds long. This division of control periods takes into account the physical characteristics of cutterhead speed, ensuring smoother speed changes within each control period.
[0117] During each control period, the system sets an upper limit for the speed regulation acceleration, which is inversely proportional to the length of the control period. For example, for a 20-second control period, the speed regulation acceleration upper limit can be set to 0.05 rpm / s; while for a 10-second control period, the upper limit can be increased to 0.1 rpm / s. Based on the set acceleration upper limit, the system calculates the speed regulation amount within each control period. Taking the first control period as an example, if the acceleration upper limit is 0.05 rpm / s and the duration is 20 seconds, the speed regulation amount within this period is 1.0 rpm.
[0118] The system calculates in real time the deviation between the actual cutterhead torque and the preset baseline torque. For example, if the measured cutterhead torque at a given moment is 2500 Nm and the preset baseline torque is 2300 Nm, the deviation is 200 Nm. This deviation reflects the difference between the current cutterhead load and the expected load.
[0119] Next, the system constructs an optimization objective function based on the torque deviation and speed adjustment. This function takes the form of a weighted sum: J = α × |torque deviation| + β × |difference between speed adjustment and target speed change|, where α and β are weighting coefficients, typically 0.7 and 0.3, respectively. This weighting approach ensures torque control accuracy while also ensuring smooth speed regulation.
[0120] The system determines optimization constraints based on equipment parameters. The maximum torque limit is determined by the equipment's rated torque. For example, if the rated torque of a shield machine's cutterhead motor is 3000 Nm, with a safety factor of 1.2, the maximum torque limit can be set to 3600 Nm. The torque change rate limit is determined based on the motor's characteristics. For example, the maximum allowable torque change per second is 10% of the rated torque, or 300 Nm / s.
[0121] Within these constraints, the system uses a gradient projection algorithm to solve the optimization objective function and derive the optimal torque adjustment sequence for multiple future moments. For example, the system might calculate the torque adjustment values for the next six control periods to be [250, 320, 380, 350, 290, 200] Nm. This sequence takes into account the current torque deviation, target speed changes, and various constraints.
[0122] The system extracts the first value of the optimal torque adjustment sequence as the initial torque adjustment value, which is 250 Nm. To avoid sudden changes in torque output, the system determines a smoothing factor based on the torque adjustment sensitivity. In practice, the torque adjustment sensitivity can be dynamically adjusted based on factors such as geological conditions and the current state of the cutterhead. For example, it can be set to 0.6 in hard rock formations and 0.8 in soft soil formations.
[0123] Multiply the initial torque adjustment by the smoothing factor to obtain the torque output increment. If the current smoothing factor is 0.7 and the initial torque adjustment is 250 Nm, the torque output increment is 175 Nm. Add this increment to the current torque to obtain the optimal torque output value. If the current torque is 2500 Nm, the final optimal torque output value is 2675 Nm.
[0124] In actual control, the system repeats this calculation process every 100 milliseconds to achieve dynamic closed-loop adjustment of the cutterhead speed. For example, during one tunneling operation, the system detected a transition from soft soil to hard rock. The cutterhead speed automatically and smoothly adjusted from 2.8 rpm to 1.5 rpm. The entire process took approximately 180 seconds, maintaining torque control accuracy within ±5%, effectively avoiding the risk of cutterhead stalling.
[0125] Through the above method, the system can calculate the optimal torque output value based on the optimal cutterhead speed, realize precise dynamic adjustment of the cutterhead speed, adapt to the excavation requirements under complex formation conditions, and improve the safety and efficiency of equipment operation.
[0126] In an optional embodiment, solving the optimization objective function by a gradient projection algorithm under the constraints of the maximum torque limit and the torque change rate limit to obtain the optimal torque adjustment sequence includes:
[0127] Setting a prediction time domain length with the current moment as the starting point, and dividing the prediction time domain length into a plurality of control periods at equal intervals; calculating an allowable torque adjustment range in each of the control periods, and using the allowable torque adjustment range as a local search space;
[0128] In each of the control periods, calculating a gradient vector of the optimization objective function; projecting the gradient vector to the boundary of the local search space;
[0129] Calculate the angle between the projected gradient direction and the boundary normal vector; adaptively adjust the search step size according to the angle; iteratively update the torque value along the projection direction; record the torque value of each iteration;
[0130] When the torque value change rate of two adjacent iterations is less than a preset change rate threshold, the current torque value is determined as the optimal torque value of the control period; the optimal torque values of each control period are combined into an optimal torque adjustment sequence.
[0131] In order to implement a technical solution for obtaining an optimal torque adjustment sequence by solving the optimization objective function through a gradient projection algorithm under the constraints of maximum torque limit and torque change rate limit, this embodiment provides detailed execution steps and technical implementation details.
[0132] The system first sets the prediction horizon length, starting from the current moment. This length is typically set between 5 and 30 seconds, depending on the control system's response characteristics and the application scenario. The prediction horizon is then divided into multiple control periods at equal intervals, ranging from 10 to 20, each of equal length. In practical applications, for example, a 10-second prediction horizon can be divided into 20 control periods, each of 0.5 seconds.
[0133] When calculating the allowable torque adjustment range for each control period, the system considers two constraints: the maximum torque limit and the torque rate of change limit. Assuming the initial torque value for the current control period is 200 N·m, the maximum torque limit is 500 N·m, the minimum torque is 0 N·m, the torque rate of change is limited to 100 N·m per second, and the control period length is 0.5 seconds, the allowable torque adjustment range within this control period is [150 N·m, 250 N·m]. This range constitutes the local search space within which the algorithm searches for the optimal torque value.
[0134] For each control period, the system calculates the gradient vector of the optimization objective function. This optimization objective function typically includes multiple sub-objectives, such as speed tracking error, fuel economy, and emissions control. Suppose that during a certain control period, the system calculates a gradient vector of 80 N·m, pointing in the direction of increasing torque. However, due to the constraints of the local search space [150 N·m, 250 N·m], this gradient vector needs to be projected to the search space boundary.
[0135] The gradient vector projection process maps the original gradient vector into the allowed search space. If the original gradient direction causes the torque to exceed the upper limit of 250 N·m, the gradient vector is projected to the upper boundary. If it causes the torque to fall below the lower limit of 150 N·m, it is projected to the lower boundary. If the torque value is within the allowed range, the original gradient direction remains unchanged. In this example, the original torque is 200 N·m, and the gradient direction indicates an increase of 80 N·m, which will result in a torque value of 280 N·m, exceeding the upper limit. Therefore, the gradient vector needs to be projected to the upper boundary of 250 N·m. The actual adjustment after projection is 50 N·m.
[0136] When calculating the angle between the projected gradient direction and the boundary normal vector, the system first determines the boundary normal vector. For the upper boundary, the normal vector points outside the search space; for the lower boundary, the normal vector points inside the search space. Assume in this example that the angle between the gradient vector and the upper boundary normal vector is 60 degrees.
[0137] Adaptively adjusting the search step size based on the angle is key to improving the algorithm's convergence efficiency. When the angle approaches 0 degrees, the gradient is nearly perpendicular to the boundary, so a larger step size should be used. When the angle approaches 90 degrees, the gradient is nearly parallel to the boundary, so a smaller step size should be used to avoid oscillation. The specific step size adjustment can be scaled by the cosine of the angle, for example, step size = base step size × (1 - |cos(angle)|). In this example, the angle is 60 degrees. If the base step size is 10 Nm, the adjusted step size is approximately 5 Nm.
[0138] When iteratively updating the torque value along the projection direction, the system adds the current torque value to the projection direction and multiplies it by the adjusted step size. In the first iteration, the torque value is adjusted from 200 N·m to 205 N·m. The system records the torque value of each iteration, forming a torque adjustment trajectory.
[0139] The iteration process continues until convergence conditions are met. Assume that the torque value after the second iteration is 209 N·m, the third is 212 N·m, the fourth is 214 N·m, the fifth is 215.5 N·m, and the sixth is 216.7 N·m. If the preset rate of change threshold is 1 N·m, and the torque value changes from the sixth to the fifth iteration by 1.2 N·m and from the seventh to the sixth by 0.8 N·m, which are less than the preset threshold, the optimal torque value for the current control period is determined to be 216.7 N·m.
[0140] Repeat the above process for each control period within the prediction time domain to obtain a series of optimal torque values. For example, the optimal torque values calculated for each of the 20 control periods are [216.7, 225.3, 232.8, 240.1, 245.6, 250.0, 255.2, 260.1, 265.0, 268.9, 272.5, 275.8, 278.9, 281.7, 284.2, 286.5, 288.6, 290.5, 292.2, 293.8] Nm. These torque values constitute the optimal torque regulation sequence.
[0141] In practice, the control system typically only executes the first value of the optimal torque adjustment sequence, 216.7 Nm, and then re-executes the entire optimization process in the next control cycle. This rolling optimization strategy can respond to system state changes and external disturbances in real time.
[0142] Figure 4 Figure 1 shows a comparison of the torque optimization performance of the gradient projection algorithms. Under normal operating conditions, the adaptive projection method performed best, achieving an optimization efficiency of 92.8%, 17.4 percentage points higher than the standard gradient method and 8.6 percentage points higher than the projection gradient method. Under high-load conditions, the optimization efficiency of the three algorithms generally decreased, but the adaptive projection method maintained a high efficiency of 89.6%, significantly outperforming the other two methods. Under variable speed conditions, the adaptive projection method also led with an efficiency of 91.2%. Under the most challenging restricted conditions, although the performance of each algorithm declined, the adaptive projection method still maintained a high efficiency of 88.9%.
[0143] Comparative results show that the introduction of gradient projection technology and an adaptive step-size adjustment strategy significantly improves the algorithm's performance when handling constrained torque optimization problems. The adaptive projection method demonstrates the best optimization efficiency in all test scenarios, and its robustness is particularly prominent when faced with high loads and constraints.
[0144] Using a gradient projection optimization algorithm, the system can quickly determine the optimal torque adjustment sequence while meeting maximum torque and torque rate limits. This allows for precise control of the engine or motor, improving the overall system's dynamic response and operating efficiency. Experiments have shown that compared to traditional control methods, this algorithm can reduce speed tracking error by 20% while improving fuel economy by approximately 5%, demonstrating its effectiveness and superiority in practical applications.
[0145] The embodiment of the present invention provides an IoT-based shield machine main drive intelligent control system, including:
[0146] The first unit is used to obtain real-time monitoring data from the multimodal sensors of the shield machine's main drive system, and to establish a multi-dimensional shield machine stratum excavation parameter database that includes time dimensions, geological feature dimensions, and excavation parameter dimensions, combined with geological parameter information of the area where the shield machine is located.
[0147] The second unit is used to establish a mapping relationship model between formation parameters and cutterhead speed and cutterhead torque using a deep reinforcement learning algorithm based on the multi-dimensional shield machine formation excavation parameter database. The deep reinforcement learning algorithm includes a state space construction module, an action space selection module and a reward function calculation module, wherein:
[0148] a third unit, configured for the state space construction module to construct a current tunneling state vector based on the real-time monitoring data, the action space selection module to select an optimal cutterhead speed within a preset speed range based on the current tunneling state vector, the reward function calculation module to calculate a reward value based on a ratio of tunneling efficiency to energy consumption per unit time, and to feed the reward value back to the action space selection module for optimizing a selection strategy for the optimal cutterhead speed;
[0149] The fourth unit is used to calculate the optimal torque output value based on the optimal cutterhead speed and the model predictive control algorithm to achieve dynamic adjustment of the cutterhead speed; according to the optimal torque output value, generate a frequency converter control instruction sequence, and the frequency converter control instruction sequence is used to adjust the output voltage and current parameters of the frequency converter to achieve dynamic balance control of the torque of the shield machine main drive system.
[0150] According to a third aspect of an embodiment of the present invention, an electronic device is provided, including:
[0151] processor;
[0152] a memory for storing processor-executable instructions;
[0153] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0154] According to a fourth aspect of an embodiment of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described above is implemented.
[0155] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.
[0156] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. The intelligent control method of shield machine main drive based on Internet of Things is characterized by: include: Acquire real-time monitoring data from the multimodal sensors of the shield machine's main drive system, and combine this with geological parameter information of the area where the shield machine is located to establish a multi-dimensional shield machine stratum excavation parameter database that includes time dimensions, geological feature dimensions, and excavation parameter dimensions. According to the multi-dimensional shield machine stratum excavation parameter database, a deep reinforcement learning algorithm is used to establish a mapping relationship model between stratum parameters and cutterhead speed and cutterhead torque. The deep reinforcement learning algorithm includes a state space construction module, an action space selection module and a reward function calculation module, wherein: The state space construction module constructs a current excavation state vector based on the real-time monitoring data; the action space selection module selects an optimal cutterhead speed within a preset speed range based on the current excavation state vector; the reward function calculation module calculates a reward value based on a ratio of excavation efficiency per unit time to energy consumption, and feeds the reward value back to the action space selection module for optimizing the selection strategy of the optimal cutterhead speed; Based on the optimal cutterhead speed, an optimal torque output value is calculated based on a model predictive control algorithm to achieve dynamic adjustment of the cutterhead speed; According to the optimal torque output value, a frequency converter control instruction sequence is generated, and the frequency converter control instruction sequence is used to adjust the output voltage and current parameters of the frequency converter to achieve dynamic balance control of the torque of the shield machine main drive system.
2. The method according to claim 1, characterized in that The action space selection module selects the optimal cutterhead speed within a preset speed range based on the current tunneling state vector, including: Constructing a speed candidate set within a preset speed range, wherein the upper limit of the preset speed range is determined based on the mechanical safety speed of the shield machine cutterhead, the lower limit of the preset speed range is determined based on the minimum tunneling efficiency of the shield machine, and the speed value interval in the speed candidate set is determined based on the control accuracy of the shield machine; Based on the current excavation state vector, the excavation distance per unit time is calculated to obtain an excavation efficiency index and the energy consumption per unit excavation distance is calculated to obtain an energy consumption index; The ratio of the excavation efficiency index to the standard excavation efficiency is recorded as a first score; the ratio of the standard energy consumption to the energy consumption index is recorded as a second score; and the weighted sum of the first score and the second score is recorded as the fitness score; The fitness score of each speed value in the speed candidate set is obtained, and a speed value with the highest fitness score is randomly selected as the optimal cutterhead speed from the speed value set according to the ε greedy strategy.
3. The method according to claim 1, characterized in that The reward function calculation module calculates a reward value based on the ratio of excavation efficiency to energy consumption per unit time, and feeds the reward value back to the action space selection module for optimizing the selection strategy of the optimal cutterhead speed, including: Calculating the tunneling efficiency per unit time based on the tunneling time and tunneling distance of the shield machine, and calculating the energy consumption per unit tunneling distance based on the power of the main motor of the shield machine; dividing the tunneling efficiency per unit time by the energy consumption per unit tunneling distance to obtain an efficiency ratio; When the performance ratio is greater than a preset benchmark performance ratio, a positive incentive value is calculated based on the difference between the two; when the performance ratio is less than the preset benchmark performance ratio, a negative inhibition value is calculated based on the difference between the two; the positive incentive value or the negative inhibition value is used as a performance evaluation value; Dynamically adjusting the cutterhead speed using an adaptive variable step size algorithm based on the performance evaluation value, calculating a speed adjustment amount, and using the speed adjustment amount as a reward value; The reward value is fed back to the action space selection module to optimize the selection strategy of the optimal cutterhead speed.
4. The method according to claim 3, characterized in that Based on the performance evaluation value, the cutterhead speed is dynamically adjusted using an adaptive variable step size algorithm. The speed adjustment amount is calculated by: Determine the rate of change of the performance evaluation value within a preset sampling period, When the change rate and the positive excitation value have the same sign, the speed adjustment amount is the square root of the product of the change rate and the positive excitation value; when the change rate and the negative inhibition value have the same sign, the speed adjustment amount is the absolute value of the product of the change rate and the negative inhibition value; When the rate of change has the same sign as the positive excitation value, the speed adjustment amount is added to the current cutterhead speed until a local optimal point is detected, wherein determining the local optimal point includes: Continuously collect the efficiency ratios of three adjacent adjustment points; calculate the efficiency ratio difference of adjacent adjustment points; when the absolute value of the latter difference is less than the absolute value of the former difference, record the middle point as the peak point. When three peak points are obtained continuously, select the peak point with the largest efficiency ratio as the local optimal point.
5. The method according to claim 1, wherein Based on the optimal cutterhead speed, the optimal torque output value is calculated based on the model predictive control algorithm to achieve dynamic adjustment of the cutterhead speed, including: Determining a predicted total duration based on the optimal cutterhead speed, the predicted total duration being the shortest time required for the optimal cutterhead speed to stably operate; dividing the predicted total duration into a plurality of control periods, the duration of each control period being proportional to the inverse of the optimal cutterhead speed; setting a speed regulation acceleration upper limit value in each control period, wherein the speed regulation acceleration upper limit value is inversely proportional to the duration of the control period; and calculating a speed regulation amount in each control period according to the speed regulation acceleration upper limit value; Calculating the deviation between the actual cutter head torque and the preset reference torque; constructing an optimization objective function based on a weighted sum of the deviation and the speed adjustment amount; determining a maximum torque limit based on the rated torque of the equipment and a torque change rate limit based on the motor characteristics; and solving the optimization objective function using a gradient projection algorithm under the constraints of the maximum torque limit and the torque change rate limit to obtain an optimal torque adjustment sequence; The first value of the optimal torque adjustment sequence is extracted as the initial torque adjustment amount; a smoothing factor is determined based on the torque adjustment sensitivity, and the smoothing factor is used to suppress torque mutations; the product of the initial torque adjustment amount and the smoothing factor is used as the torque output increment; and the torque output increment is superimposed on the current torque to obtain the optimal torque output value.
6. The method according to claim 5, characterized in that Solving the optimization objective function by a gradient projection algorithm under the constraints of the maximum torque limit and the torque change rate limit to obtain an optimal torque adjustment sequence includes: Setting a prediction time domain length with the current moment as the starting point, and dividing the prediction time domain length into a plurality of control periods at equal intervals; calculating an allowable torque adjustment range in each of the control periods, and using the allowable torque adjustment range as a local search space; In each control period, calculating a gradient vector of the optimization objective function; projecting the gradient vector to the boundary of the local search space; Calculate the angle between the projected gradient direction and the boundary normal vector; adaptively adjust the search step size according to the angle; iteratively update the torque value along the projection direction; record the torque value of each iteration; When the torque value change rate of two adjacent iterations is less than a preset change rate threshold, the current torque value is determined as the optimal torque value of the control period; the optimal torque values of each control period are combined into an optimal torque adjustment sequence.
7. An intelligent control system for main drive of a shield machine based on the Internet of Things, used to implement the method according to any one of claims 1 to 6, characterized in that: include: The first unit is used to obtain real-time monitoring data from the multimodal sensors of the shield machine's main drive system, and to establish a multi-dimensional shield machine stratum excavation parameter database that includes time dimensions, geological feature dimensions, and excavation parameter dimensions, combined with geological parameter information of the area where the shield machine is located. The second unit is used to establish a mapping relationship model between formation parameters and cutterhead speed and cutterhead torque using a deep reinforcement learning algorithm based on the multi-dimensional shield machine formation excavation parameter database. The deep reinforcement learning algorithm includes a state space construction module, an action space selection module and a reward function calculation module, wherein: a third unit, configured for the state space construction module to construct a current tunneling state vector based on the real-time monitoring data, the action space selection module to select an optimal cutterhead speed within a preset speed range based on the current tunneling state vector, the reward function calculation module to calculate a reward value based on a ratio of tunneling efficiency to energy consumption per unit time, and to feed the reward value back to the action space selection module for optimizing a selection strategy for the optimal cutterhead speed; The fourth unit is used to calculate the optimal torque output value based on the optimal cutterhead speed and the model predictive control algorithm to achieve dynamic adjustment of the cutterhead speed; according to the optimal torque output value, generate a frequency converter control instruction sequence, and the frequency converter control instruction sequence is used to adjust the output voltage and current parameters of the frequency converter to achieve dynamic balance control of the torque of the shield machine main drive system.
8. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Shield optimal autonomous tunneling control method based on deep reinforcement learning
CN113486463A
Intelligent earth pressure control system of earth pressure balance shield machine
CN113847049A