Intelligent scheduling method for tunnel boring machine collaborative manufacturing
By combining hybrid integer programming with deep reinforcement learning, an intelligent scheduling method is developed to address the limitations of manual scheduling and the problem of multi-source heterogeneous data fusion in tunnel boring machine (TBM) manufacturing. This method enables efficient, stable, and flexible intelligent scheduling in the TBM manufacturing process.
Patent Information
- Application Number
- CN202411727145.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-28
- Publication Date
- 2026-01-09
- Estimated Expiration
- 2044-11-28
AI Technical Summary
Traditional tunnel boring machine (TBM) manufacturing scheduling relies on manual experience, making it difficult to fully consider the constraints and optimization objectives in the manufacturing process. This leads to unreasonable planning, production interruptions, and increased costs. Furthermore, existing intelligent scheduling methods have failed to effectively address the unique characteristics of TBMs and the problem of fusion of multi-source heterogeneous data.
An intelligent scheduling method combining hybrid integer programming and deep reinforcement learning, along with robust optimization, is adopted to achieve real-time monitoring and dynamic optimization of the tunnel boring machine manufacturing process through data acquisition and preprocessing, intelligent scheduling modeling and solving, and monitoring and feedback optimization.
It improves the efficiency and stability of the tunnel boring machine manufacturing process, reduces costs, ensures product quality, enhances the ability to resist interference from uncertain factors, realizes semantic association and sharing of data, and improves the accuracy and flexibility of scheduling.
Smart Images

Figure CN119671136B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of tunnel boring machine manufacturing technology, in particular to an intelligent scheduling method for collaborative manufacturing of tunnel boring machines, aiming to solve the problem of multi-process and multi-resource collaborative scheduling in the manufacturing process of tunnel boring machines through innovative algorithms and data processing procedures, improve manufacturing efficiency, reduce costs and ensure product quality. BACKGROUND
[0002] With the large-scale development of tunnel engineering construction, the demand for tunnel boring machines is increasing, and its manufacturing process involves the processing, assembly of numerous parts and complex process flow. Traditional tunnel boring machine manufacturing scheduling mainly relies on manual experience and simple planning and scheduling tools. Production planning personnel manually develop manufacturing plans and scheduling schemes based on order demand, equipment capacity, personnel skills and other factors. This approach has many limitations: first, in the face of massive manufacturing data and complex process associations, manual consideration of various constraint conditions and optimization objectives is difficult, which can lead to unreasonable plans, such as coexistence of idle and overloaded equipment, and long production cycles; second, in the manufacturing process, once there are sudden situations such as equipment failure, raw material supply delay or order changes, manual scheduling is difficult to respond and adjust quickly, often causing production interruption or delay, increasing manufacturing costs and delivery risks; third, traditional scheduling methods lack real-time monitoring and data analysis capabilities for the manufacturing process, making it difficult to identify potential production bottlenecks and quality risks, and to continuously optimize the manufacturing process.
[0003] In recent years, intelligent scheduling technology has gradually been applied in manufacturing, but in the field of tunnel boring machine manufacturing, it still faces some challenges. Existing intelligent scheduling algorithms are mostly designed for common manufacturing problems and do not fully consider the particularity of tunnel boring machine manufacturing, such as the high precision requirements of large part processing, the coordination of multi-system assembly, and the impact of complex geological adaptability design on manufacturing processes. Moreover, existing scheduling systems have limitations in handling multi-source heterogeneous data (such as design drawings, process files, equipment operation data, quality detection data, etc.), making it difficult to achieve data-driven precise scheduling decisions. In addition, the robustness and adaptability of current intelligent scheduling methods in dealing with uncertainty need to be improved, and they cannot effectively ensure the stable operation of the tunnel boring machine manufacturing process in complex and variable environments. Therefore, there is an urgent need for an innovative intelligent scheduling system and method for collaborative manufacturing of tunnel boring machines to meet the needs of industry development. SUMMARY
[0004] The application provides a tunnel boring machine collaborative manufacturing intelligent scheduling method, comprising a data acquisition and preprocessing step, an intelligent scheduling modeling and solving step, a monitoring and feedback optimization step; the data acquisition and preprocessing step acquires multi-source heterogeneous data from each link of tunnel boring machine manufacturing and performs preprocessing; the intelligent scheduling modeling and solving step solves the static scheduling problem by using a mixed integer programming model and introduces a deep reinforcement learning algorithm for dynamic optimization, while a robust optimization method is used to enhance the robustness of the scheduling scheme; the monitoring and feedback optimization step performs real-time monitoring on the manufacturing process and optimizes the scheduling scheme according to the feedback information.
[0005] Preferably: the data acquisition and preprocessing step includes collecting data from design, process planning, production and material management links, preprocessing by using an ontology-based data fusion method, and processing data by using an outlier detection algorithm to ensure reliability.
[0006] Preferably: the intelligent scheduling modeling and solving step decomposes the manufacturing task into subtasks by using a mixed integer programming model, establishes an allocation matrix between the task and the resource, solves the initial scheduling scheme by using a multi-objective function, dynamically optimizes according to the real-time state information by using a deep reinforcement learning algorithm, and considers uncertain parameters by using a robust optimization method to enhance the stability of the scheduling scheme.
[0007] Preferably: the monitoring and feedback optimization step establishes a real-time monitoring system to display key information on a visual interface, collects feedback information including worker operation feedback, quality problem feedback and equipment maintenance records, and optimizes the scheduling scheme and updates the scheduling algorithm and model according to the feedback information.
[0008] Preferably: it further comprises training an echo state network, taking historical manufacturing data as input and outputting a prediction of future manufacturing state; then, in the ant colony optimization algorithm, the ants select the next task and resource for scheduling decision according to the pheromone concentration and the prediction result of the ESN; after each iteration, the pheromone concentration is updated according to the quality of the scheduling scheme, and the ESN is used to update the prediction of the new manufacturing state, guiding the ant colony to search for a better scheduling scheme.
[0009] Preferably: it further comprises encoding the equipment failure state as a binary vector and representing the production progress deviation as a percentage value, selecting actions by using a deep Q network, the actions including adjusting the process sequence, replacing the processing equipment, postponing or advancing the task start time; according to the execution result of the action, the environment gives the agent corresponding reward feedback, and the reward function comprehensively considers the change of the production cycle, the improvement of the equipment utilization rate and the reduction of the quality risk factors.
[0010] Preferably: use an intelligent scheduling system, including a data acquisition and integration module, an intelligent scheduling decision module, and a monitoring and feedback module; the data acquisition and integration module is used to collect multi-source heterogeneous data from each link of the tunnel boring machine manufacturing, and to integrate and preprocess; the intelligent scheduling decision module uses a combination of mixed integer programming and deep reinforcement learning algorithm for scheduling decision, and considers the uncertainty factors for robustness enhancement; the monitoring and feedback module monitors the manufacturing process in real time and collects feedback information to optimize the scheduling scheme.
[0011] Preferably: the manufacturing process is regarded as a dynamic system, and ESN is used to model it to learn the variation law of manufacturing tasks and resources in time series.
[0012] Preferably: the mixed integer programming model of the intelligent scheduling decision module takes minimizing production cycle, maximizing equipment utilization, and reducing production cost as the objective function, considers order delivery period, equipment capacity, personnel working hours, and material supply constraint conditions to solve the initial scheduling scheme, the deep reinforcement learning algorithm dynamically optimizes the initial scheme, and the robust optimization method enhances the anti-interference ability of the scheduling scheme to uncertain factors.
[0013] Preferably: the monitoring and feedback module displays production progress, equipment running state, and quality detection result key information in a visual interface, collects feedback information in the manufacturing process and feeds back to the intelligent scheduling decision module to optimize the scheduling algorithm and model.
[0014] Beneficial technical effects: Multi-source heterogeneous data is obtained from each link of the tunnel boring machine manufacturing, including design data, process data, production data, and material data, etc., which provides a rich data basis for intelligent scheduling. The ontology-based data fusion method is adopted to effectively solve the semantic and structural heterogeneity problems of multi-source heterogeneous data. By constructing the ontology model of the tunnel boring machine manufacturing field, the semantic association and sharing of different source data are realized, and the availability and understandability of the data are improved. The algorithm combining mixed integer programming and deep reinforcement learning is proposed for intelligent scheduling of tunnel boring machine collaborative manufacturing. The robust optimization method is introduced to enhance the robustness of the scheduling scheme, and the influence of uncertain factors such as raw material supply and equipment maintenance in the tunnel boring machine manufacturing process is considered. BRIEF DESCRIPTION OF DRAWINGS
[0015] Figure 1 Flow principle diagram. DETAILED DESCRIPTION
[0016] Example 1
[0017] Multi-source heterogeneous data is collected from various aspects of tunnel boring machine manufacturing. In the design stage, CAD design drawings, design change records, and product structure tree information are obtained, which are converted into processable data formats through specific data interfaces; in the process planning stage, CAPP process files are collected, including process flow, machining parameters, tooling fixture requirements, and other data; in the production and manufacturing stage, various sensors are installed, such as power sensors, vibration sensors, and temperature sensors for monitoring equipment operation status, high-precision measuring instruments for part machining size detection, and visual sensors for assembly process monitoring, to collect equipment operation data, part quality data, and assembly progress data; in the material management stage, barcodes, RFID, and other technologies are used to track raw materials, parts, storage location, and quantity information.
[0018] The collected multi-source heterogeneous data is integrated and preprocessed. An ontology-based data fusion method is used to construct a tunnel boring machine manufacturing domain ontology model, which maps data from different sources and formats to a unified ontology framework, achieving semantic association and sharing of data. For example, part names and numbers in design drawings are matched with process objects in process files and actual parts in production and manufacturing, solving the problem of data heterogeneity. At the same time, data is cleaned, denoised, and outlier processed to ensure data accuracy and integrity. For example, a statistical analysis-based outlier detection method is used to identify and eliminate abnormal fluctuations in equipment operation data, providing a reliable data foundation for subsequent analysis and decision-making.
[0019] Based on the integrated data, an innovative algorithm combining mixed integer programming and deep reinforcement learning is used for scheduling and decision-making. First, a mixed integer programming model is used to model the static scheduling problem of tunnel boring machine manufacturing, considering order delivery period, equipment capacity, personnel working hours, material supply, and other constraint conditions, with the objective function of minimizing production cycle, maximizing equipment utilization, and reducing production cost. The initial scheduling scheme is obtained by solving the model. For example, the assembly process and part machining process of the tunnel boring machine are regarded as tasks, and the machine tools and assembly stations are regarded as resources, and a model of the allocation relationship between tasks and resources is established, which is solved by an optimization algorithm.
[0020] Then, a deep reinforcement learning algorithm is introduced to dynamically optimize and adjust the initial scheduling scheme. An agent based on deep Q-network (DQN) is constructed, taking real-time state information in the manufacturing process (such as equipment failure information, order change information, production progress deviation, etc.) as input. Through interaction with the manufacturing environment, the agent learns the optimal scheduling strategy. For example, when a key processing device fails, the agent decides whether to adjust the process sequence, transfer production tasks to other devices, or wait for device repair, based on the current production task state, the idle status of other devices, and the material inventory status. During training, a priority experience replay-based technique is used to store and learn experience data with high value for improving the agent's strategy, improving learning efficiency and convergence speed.
[0021] In addition, considering the uncertainty factors in the tunnel boring machine manufacturing process, a robust optimization method is used to enhance the robustness of the scheduling scheme. Interval representation of uncertain parameters is introduced in the mixed integer programming model, such as fluctuation range of raw material supply time, uncertainty of equipment maintenance time, etc. By solving the robust optimization problem, a scheduling scheme that maintains good performance within a certain range of uncertainty is obtained, improving the anti-interference ability of the scheduling scheme to uncertain factors.
[0022] Real-time monitoring of the tunnel boring machine manufacturing process is performed, and key information such as production progress, equipment running status, and quality detection results is displayed through a visual interface. For example, the planned progress and actual progress of each process are displayed in the form of a Gantt chart, and the utilization rate and failure rate of equipment are displayed in the form of a dashboard, so that management personnel can intuitively understand the production situation.
[0023] Feedback information from the manufacturing process is collected, such as worker operation feedback, quality problem feedback, and equipment maintenance records, which is fed back to the intelligent scheduling decision module in a timely manner. For example, when workers find quality problems in a certain process, they report the information through the feedback system, and the intelligent scheduling decision module re-evaluates the scheduling scheme based on the feedback information, which may increase the quality inspection process or adjust the processing parameters of subsequent processes to ensure product quality. At the same time, based on monitoring and feedback information, the scheduling algorithm and model are continuously optimized and improved, continuously improving the performance and adaptability of the scheduling system.
[0024] II. Method flow
[0025] Multi-source heterogeneous data is collected from the design, process planning, production, and material management of tunnel boring machine manufacturing. In the design phase, geometric information and attribute information from CAD design drawings are extracted using a data conversion tool and integrated with design change records. In the process planning phase, CAPP process files are analyzed to extract data such as process steps, processing methods, and process parameters. In the production and manufacturing phase, real-time data collection is performed using a sensor network to collect equipment operation data and quality inspection data. For example, machine tool operation status data is collected every 1 second, and part processing precision data is collected immediately after each process step. In the material management phase, barcodes or RFID readers are used to scan material information at regular intervals to update material inventory and transfer data.
[0026] An ontology-based data fusion method is used to preprocess the collected data. An ontology model of tunnel boring machine manufacturing is constructed to define various entities (such as parts, equipment, and processes) and their attributes and relationships (such as assembly relationships and processing sequence relationships). Data from different sources is mapped and converted according to the ontology model to achieve semantic unification. For example, part entities in design drawings are associated with process entities in process files through a "processing object" relationship. In the data cleaning process, an outlier detection algorithm based on data distribution characteristics, such as the box plot method, is used to identify and process outliers in equipment operation data. Data that is more than 1.5 times the upper and lower quartiles of the box plot is considered an outlier and is corrected or replaced based on historical trends and adjacent data points to ensure data reliability.
[0027] A mixed integer programming model is used to model the static scheduling problem of tunnel boring machine manufacturing. The manufacturing tasks of the tunnel boring machine are divided into multiple sub-tasks, such as part processing tasks and assembly tasks, each with different processing times, resource requirements, and process constraints. Manufacturing resources such as machine tools, assembly workers, and tooling fixtures are classified and quantified, and an allocation matrix between tasks and resources is established. The objective functions include minimizing production cycle time, maximizing equipment utilization, and minimizing material inventory cost, while considering order delivery period constraints, equipment capacity constraints, and material supply constraints. A mixed integer programming model is constructed. For example, for the cutter processing task of a certain type of tunnel boring machine, it is necessary to determine which machine tools to use, the processing sequence, and the processing time arrangement, while meeting the requirements of machine tool processing precision, tool life limitations, and raw material supply time requirements.
[0028] An optimization algorithm is used to solve the mixed integer programming model to obtain an initial scheduling scheme. For example, a branch and bound algorithm is used to solve the model, and by continuously branching and pruning, the optimal task allocation scheme is searched. During the solving process, according to the size and complexity of the problem, the parameters of the algorithm are reasonably set, such as the branching strategy, the pruning condition, etc., to improve the solving efficiency.
[0029] A deep reinforcement learning algorithm is introduced to dynamically optimize the initial scheduling scheme. The real-time state information in the manufacturing process is digitized and characterized to construct the state space. For example, the device failure state is encoded as a binary vector, and the production progress deviation is represented as a percentage value. The agent selects actions according to the current state through a deep Q network, including adjusting the process sequence, replacing the processing equipment, delaying or advancing the task start time, etc. According to the execution result of the action, the environment gives the agent corresponding reward feedback, and the reward function considers factors such as changes in production cycle, improvements in equipment utilization, and reductions in quality risk. For example, if the agent's action shortens the production cycle and improves the equipment utilization, a higher reward is given; otherwise, if it causes quality problems or production delays, a penalty is given. Through a large number of training iterations, the agent learns the optimal scheduling strategy and realizes dynamic optimization of the manufacturing process.
[0030] Robust optimization methods are used to enhance the robustness of the scheduling scheme. In the mixed integer programming model, uncertain parameters such as raw material supply time and equipment maintenance time are represented as interval numbers. For example, the supply time of a certain raw material is expected to be between 3-5 days, represented as the interval [3,5]. By solving the robust optimization problem, a scheduling scheme is obtained that can still meet certain performance requirements within the fluctuation range of uncertain parameters. For example, even if the raw material supply time is delayed to 5 days, the scheduling scheme can still ensure the continuity of production and will not cause delays in critical processes.
[0031] A real-time monitoring system for the tunnel boring machine manufacturing process is established, and real-time data obtained through data acquisition and integration modules are displayed on the visual interface to show production progress, equipment running status, quality detection results, etc. For example, using big data visualization technology, a production progress curve is drawn to show the change of the completion rate of each process over time; the internal structure and running state of the equipment are displayed through a three-dimensional model, such as using different colors to represent the normal, failure, maintenance, etc. state of the equipment.
[0032] Feedback information from the manufacturing process is collected, including workers' operation experience feedback, quality problem feedback, equipment maintenance feedback, etc. For example, worker feedback terminals are set up on the assembly line, and workers can input problems found during assembly, such as loose parts, unreasonable assembly sequence, etc.; at the quality detection link, the detection results and quality problems found are fed back to the system in a timely manner.
[0033] Based on monitoring and feedback information, the intelligent scheduling decision module optimizes and adjusts the scheduling scheme. For example, if a piece of equipment is found to be frequently malfunctioning and the repair time is long, the scheduling decision module will reassess the task allocation on that equipment, transferring some tasks to other similar equipment and adjusting the planned time of subsequent processes. Simultaneously, this feedback information is used to update and improve the scheduling algorithm and model. For example, based on feedback on quality issues, the weight of quality control strategies in the scheduling model is adjusted to strengthen the monitoring and resource allocation of critical quality processes; based on equipment maintenance feedback, the arrangement of equipment maintenance plans in the scheduling scheme is optimized to improve equipment reliability and availability. Specifically:
[0034] (I) Task decomposition and resource quantification
[0035] 1. Task breakdown:
[0036] The manufacturing task of a tunnel boring machine (TBM) is decomposed into a series of sub-tasks, such as component processing tasks (e.g., cutterhead processing, shield processing, etc.) and assembly tasks (e.g., assembling the processed components into a complete machine). Let set (C = {C1, C2, ..., C...}) be defined. n Let} represent the set of all subtasks, where n is the total number of subtasks.
[0037] 2. Resource quantification:
[0038] Classify and quantify the various resources involved in the manufacturing process. For example, manufacturing resources include a set of machine tools M = {M1, M2, ..., M...} m} (where m is the number of machine tools), the set of assembly workers W = {W1, W2, ..., W...} w} (w is the number of assembly workers), the set of tooling fixtures F = {F1, F2, ..., F f (f represents the number of types of tooling fixtures, etc.)
[0039] 3. Objective function:
[0040] Minimize production cycle:
[0041] Let t i For the i-th subtask C i The processing time on the j-th machine tool M (if the subtask is not processed on this machine tool, then t) i =0),s i Let be the start time of the i-th subtask. The production cycle T can be expressed as the maximum completion time of all subtasks, i.e.:
[0042] T = max i∈C {s i +∑ j∈M t ij}
[0043] The objective is to minimize the production cycle, denoted as minT.
[0044] Maximize machine utilization:
[0045] For each machine M, let T be the total time available for processing, and let ∑ i∈C t ij be the total time actually spent on processing subtasks. Then the utilization U of machine M can be expressed as:
[0046]
[0047] We want to maximize the average utilization of all machines, and let the objective function be:
[0048]
[0049] The objective is to maximize Z1, i.e., maxz1.
[0050] Minimize material inventory cost:
[0051] Let I be the inventory level of the kth material, c be the unit inventory cost of the kth material, and K be the set of material types, K = {K1, K2, …, K k}. Then the material inventory cost C I can be expressed as:
[0052] C I = ∑ k∈K c k × I k
[0053] Material supply constraint:
[0054] Let r i be the unit time demand of the kth material by the ith subtask, and R be the supply rate of the kth material. Then we have: ∑ i∈C r ik t ij x ij ≤ R k
[0055] For all material types k ∈ K and all machines j ∈ M.
[0056] Human resource constraint:
[0057] For tasks that require manual operation, such as assembly tasks, let h iw be the number of man-hours required by the ith subtask for the wth assembly worker. If the total number of assembly workers is limited, let H be the total number of assembly workers, then we have:
[0058] ∑ i∈C h iw≤H
[0059] For all assembly workers w ∈ W.
[0060] (ii) Solution of mixed integer programming model
[0061] For the constructed mixed integer programming model, common solution methods include branch and bound algorithm, cutting plane algorithm, etc. Taking the branch and bound algorithm as an example:
[0062] 1. Initialization:
[0063] The original problem is taken as the root node, and the initial upper and lower bounds are set. The upper bound can quickly obtain a feasible solution through some heuristic algorithm, and the objective function value is calculated as the initial upper bound; the lower bound is usually set to negative infinity (because we want to minimize the objective function).
[0064] 2. Branching:
[0065] At the current node (starting from the root node), select an integer variable (such as decision variable; x i ), according to its possible values (0 or 1) to divide the current node into two sub-nodes, corresponding to different value cases.
[0066] 3. Bound:
[0067] For each newly generated sub-node, a lower bound is obtained by solving its corresponding linear programming relaxation problem (i.e. allowing integer variables to take real numbers). If this lower bound is greater than the current upper bound, the sub-node can be pruned (i.e. no longer continue to branch), because it cannot produce a better solution.
[0068] 4. Iteration:
[0069] Repeat the above branching and bounding operations, constantly update the upper and lower bounds, until the optimal solution is found or it is determined that there is no solution.
[0070] Through continuous branching and pruning, the optimal task allocation scheme is searched. During the solving process, according to the size and complexity of the problem, the parameters of the algorithm are set reasonably, such as branching strategy (selecting which integer variable to branch), pruning condition (such as comparing the lower bound with the upper bound to decide whether to prune), etc., to improve the solving efficiency.
[0071] After obtaining the initial scheduling scheme, a deep reinforcement learning algorithm is introduced to dynamically optimize it.
[0072] 1. Definition of state space:
[0073] Device state: Let F be the fault state of the jth machine tool, normal is 0, fault is 1. Encode the fault state of all machine tools into a binary vector F = [F1, F2,..., F m]。
[0074] Production progress status: Let P be the production progress, expressed as the percentage of completed subtasks over the total number of subtasks, i.e. where ncompleted is the number of completed subtasks. Let P target be the predetermined production progress target, then the production progress deviation can be expressed as d P = P - P target .
[0075] Material inventory status: Let I be the inventory level of the kth material, and let the inventory levels of all materials form a vector I = [I1, I2,..., In]. k ].
[0076] In summary, the state space S can be expressed as S = {F, d P , I}.
[0077] 2. Action space definition:
[0078] Adjusting process sequence: Let (ai be the action of adjusting the process sequence, for two subtasks (C i and C, the adjustment of the process sequence is achieved by swapping their execution order in the scheduling scheme. A action vector ai = [C i , C] can be defined.
[0079] Changing processing equipment: Let (a2 be the action of changing the processing equipment, for a subtask C i , it is transferred from the current processing equipment M to another equipment M. A action vector a2 = [C i , M, M] can be defined. i , its start time is delayed or advanced by At time units. A action vector a3 = [C i , At] can be defined.
[0080] The action space A can be expressed as A = ai, a2, a3.
[0081] 3. Reward function definition:
[0082] Production cycle reward: Let T rev be the production cycle before executing the action, and T new be the production cycle after executing the action. If T new < T prev , a reward is given for the shortening of the production cycle, and the reward value can be set as r T = a1(T prev - T new), where a1 is the reward coefficient for production cycle shortening. If T new ≥ T prev , a corresponding penalty is given, which can be set as r T = -a1(T new - T prev ).
[0083] Device utilization reward: Let U rev be the average device utilization before performing the action, and U new be the average device utilization after performing the action. If U new > U prev , a reward is given for the increase in device utilization, which can be set as r U = a2(U new - U prev ), where a2 is the reward coefficient for device utilization increase. If U new ≤ U prev , a corresponding penalty is given, which can be set as r U = -a2(U new - U prev ).
[0084] Quality risk reward: Let Q rev be the quality risk indicator before performing the action (which can be evaluated by the frequency and severity of quality issues), and Qnew be the quality risk indicator after performing the action. If Q new < Q prev , a reward is given for the reduction in quality risk, which can be set as r Q = a3(Q prev - Q new ), where a3 is the reward coefficient for quality risk reduction. If Q new ≥ Q prev , a corresponding penalty is given, which can be set as r Q = -a3(Q new - Q prev ).
[0085] In summary, the reward function can be represented as r = r T + r U + r Q .
[0086] Training process:
[0087] The agent selects an action a e A according to the current state S through a deep Q network (DQN). After performing the action, the environment gives the agent a corresponding reward feedback according to the reward function. Through a large number of training iterations, the agent learns the optimal scheduling strategy and realizes the dynamic optimization of the manufacturing process.
[0088] To enhance the robustness of the scheduling scheme, a robust optimization method is adopted.
[0089] 1. Uncertain parameter representation:
[0090] Uncertain parameters such as raw material supply time, equipment maintenance time, etc. are represented as interval numbers. For example, suppose the supply time of a certain raw material is expected to be between a-b, which is represented as interval [a, b]. Let the uncertain parameter vector be ξ = [ξ1, ξ2, …], where ξi is a different uncertain parameter.
[0091] 2. Robust optimization model construction:
[0092] Based on the mixed integer programming model, the influence of uncertain parameters is considered. For example, for the objective function of minimizing the production cycle, within the fluctuation range of uncertain parameters, the production cycle is required to meet certain robustness requirements. Let T(ξ) be the production cycle considering uncertain parameters, and the production cycle is required to be no more than a given robustness upper limit T R , i.e.:
[0093]
[0094] For other objective functions and constraints, corresponding adjustments and redefinitions need to be made according to the influence of uncertain parameters to ensure the robustness of the scheduling scheme in uncertain environment.
[0095] Inventive beneficial effects
[0096] 1. Innovative data collection and fusion method:
[0097] A comprehensive data collection system is constructed, which can obtain multi-source heterogeneous data from various aspects of tunnel boring machine manufacturing, including design data, process data, production data and material data, etc., providing rich data basis for intelligent scheduling. Compared with traditional data collection methods, the data collection range of the present invention is wider, the data types are more complete, and the actual situation of the manufacturing process can be more comprehensively reflected.
[0098] An ontology-based data fusion method is adopted, which effectively solves the problems of semantic heterogeneity and structural heterogeneity of multi-source heterogeneous data. By constructing the ontology model of tunnel boring machine manufacturing field, the semantic association and sharing of data from different sources are realized, and the usability and understandability of data are improved. This data fusion method can better mine the internal relationship between data, and provide more accurate and comprehensive information support for intelligent scheduling decision-making.
[0099] 2. Innovative intelligent scheduling algorithm based on mixed integer programming and deep reinforcement learning:
[0100] An algorithm combining mixed integer programming and deep reinforcement learning is proposed for intelligent scheduling of tunnel boring machine collaborative manufacturing. Mixed integer programming can accurately model and solve the static scheduling problem of the manufacturing process, considering various constraints and optimization objectives, to obtain an initial scheduling scheme. Deep reinforcement learning can optimize and adjust the initial scheme based on real-time dynamic information in the manufacturing process, effectively responding to uncertain factors and making dynamic decisions. This hybrid algorithm fully leverages the strengths of both methods, overcoming the limitations of traditional scheduling methods in handling static and dynamic problems, and improving the accuracy and flexibility of scheduling.
[0101] Robust optimization methods are introduced to enhance the robustness of the scheduling scheme, taking into account the impact of uncertain factors such as raw material supply and equipment maintenance in the tunnel boring machine manufacturing process. By solving the robust optimization problem, a scheduling scheme is obtained that maintains good performance within a certain range of uncertainty, effectively reducing the interference of uncertain factors on the manufacturing process and improving the stability and reliability of production.
[0102] 3. Real-time monitoring and feedback optimization mechanism innovation:
[0103] A complete monitoring and feedback system is established to monitor the tunnel boring machine manufacturing process in real time, visually displaying key information such as production progress, equipment operating status, and quality detection results. Through the visual interface, management personnel can promptly identify problems and abnormalities in the production process, providing timely information support for decision-making.
[0104] Feedback information from the manufacturing process, including workers' operating experience, quality problems, and equipment maintenance information, is fully utilized to optimize and adjust the scheduling scheme. This feedback optimization mechanism enables the scheduling system to continuously learn and improve, adapting to various changes and challenges in the manufacturing process, and improving the adaptability and intelligence of the scheduling system.
[0105] Example 2
[0106] Various types of sensors are deployed at each link in the manufacturing site, including device state sensors (monitoring operating parameters, fault information of processing equipment and assembly equipment), material sensors (inventory quantity and location information of raw materials and components), personnel positioning sensors (worker's working position and state), and environmental sensors (temperature, humidity, and other environmental parameters that affect manufacturing). At the same time, order information, production plans, process files, and other data are obtained from enterprise resource planning (ERP) systems, manufacturing execution systems (MES), and other information systems.
[0107] Data fusion techniques are used to integrate data from different sources, eliminating redundancy and inconsistencies. The collected data is pre-processed, including data cleaning (removing outliers and noise), data normalization (unifying the scale of different magnitudes), and data encoding (converting non-numerical data into a form suitable for computation), providing a high-quality data foundation for subsequent scheduling decisions.
[0108] Based on the design drawings and manufacturing process requirements of the tunnel boring machine, a manufacturing task model is established. The model includes manufacturing tasks, assembly tasks, and quality inspection tasks for each component, clearly defining the input and output, process parameters, quality standards, and time constraints of each task.
[0109] The overall manufacturing task is divided into multiple sub-tasks, and the sequence, parallel relationship, and resource dependency between sub-tasks are analyzed. A task relationship network is constructed to clearly represent the logical structure of the entire manufacturing process, providing constraint information at the task level for scheduling algorithms.
[0110] Various resources involved in the manufacturing process are classified and modeled, including production equipment (different types and specifications of machining tools, welding equipment, etc.), human resources (workers with different skill levels), material resources (raw materials, components, and supplier information), and auxiliary resources (such as tooling fixtures, transportation tools, etc.). The attributes of each resource are defined, such as the processing capacity of equipment, the skill range of personnel, the supply cycle of materials, etc.
[0111] Real-time tracking of resource state changes is performed through sensor data and information system feedback to update the availability, remaining workload, and location of resources in a timely manner. A resource state prediction model is established to predict the state changes of resources in the future based on historical data and current trends, providing forward-looking information for scheduling decisions.
[0112] Considering the manufacturing task and resource information, an intelligent scheduling model is constructed. The model aims to minimize the manufacturing cycle, maximize resource utilization, and ensure product quality, while considering complex constraints such as equipment maintenance plans, material supply limitations, and personnel working hours.
[0113] The following three algorithms are used to realize intelligent scheduling decision-making and generate detailed scheduling plans, including the start time, execution resources, completion time, and material distribution plan of each sub-task. The generated scheduling plan is converted into specific scheduling instructions, which are issued to the corresponding execution units through the interface with the manufacturing site's equipment control system, personnel management system, and material distribution system. This ensures that each link works in an orderly manner according to the scheduling plan. The execution of the manufacturing process is monitored in real-time, and by comparing the actual progress with the scheduling plan, any abnormalities such as production delays, equipment failures, or material shortages can be detected in a timely manner. When an abnormality occurs, the emergency handling mechanism is activated, and the dynamic scheduling algorithm is used to adjust the original scheduling plan, reallocate resources, and adjust the task sequence to ensure the continuity of the manufacturing process.
[0114] (Solution One) Scheduling Model Based on Super-Heuristic Evolutionary Fuzzy System: The super-heuristic evolutionary fuzzy system combines the advantages of super-heuristic algorithms and evolutionary fuzzy systems. Super-heuristic algorithms are used to select and combine low-level heuristic algorithms at a high level to adapt to different scheduling scenarios. Evolutionary fuzzy systems optimize fuzzy rules and membership functions through evolutionary algorithms, enabling them to handle the fuzziness and uncertainty in manufacturing scheduling, such as fuzzy evaluation of worker skills and fuzzy estimation of equipment failure risks.
[0115] First, a set of various low-level heuristic scheduling algorithms is constructed. The super-heuristic algorithm selects appropriate heuristic algorithms from the set for combination based on the characteristics of the current manufacturing tasks and resources. Meanwhile, evolutionary algorithms are used to perform evolutionary operations on the rule base and membership functions of the fuzzy system. During scheduling, the information of tasks and resources is input into the evolutionary fuzzy system, and scheduling decisions such as task allocation, resource selection, and time arrangement are obtained through fuzzy reasoning.
[0116] This model can adaptively select appropriate scheduling strategies in complex scheduling environments and effectively deal with various fuzzy and uncertain factors in tunnel boring machine manufacturing. Compared with traditional scheduling algorithms, it has stronger robustness and universality, and can improve the quality and efficiency of scheduling decisions.
[0117] (II) Solution two: dynamic scheduling model integrating echo state network and ant colony optimization Echo state network (ESN) is a special type of recurrent neural network that has the ability to model dynamic systems quickly. By treating the manufacturing process as a dynamic system, ESN is used to model it and learn the changing patterns of manufacturing tasks and resources over time. Ant colony optimization algorithm searches for the optimal solution by simulating the pheromone dissemination mechanism of ants during their search for food. In this model, ants search for scheduling solutions in the solution space composed of tasks and resources, and the update of pheromone is influenced by the output of ESN, so that the ant colony optimization algorithm can better adapt to the dynamic changes of the manufacturing process.
[0118] First, the echo state network is trained, using historical manufacturing data (including task progress, resource status, etc.) as input to predict future manufacturing states. Then, in the ant colony optimization algorithm, ants select the next task and resource for scheduling decisions based on pheromone concentration and ESN prediction results. After each iteration, the pheromone concentration is updated according to the quality of the scheduling solution, and the ESN is used to update the prediction of the new manufacturing state, guiding the ant colony to search for better scheduling solutions.
[0119] This integrated model can fully utilize the modeling capabilities of echo state networks for dynamic systems and the search capabilities of ant colony optimization algorithms to achieve dynamic and accurate scheduling of tunnel boring machine manufacturing processes. It can effectively respond to real-time changes in the manufacturing process, such as emergency order insertion and equipment failure, improving scheduling flexibility and response speed.
[0120] (III) Solution three: hybrid scheduling model based on quantum annealing and particle swarm optimization Quantum annealing algorithm uses quantum fluctuations to search for the optimal solution in the solution space, with the potential to quickly find the global optimal solution in complex energy landscapes. Particle swarm optimization algorithm simulates the foraging behavior of bird flocks, allowing particles to fly in the solution space and adjust their flight direction based on their own and group experience to find the optimal solution. By combining the two, quantum annealing is used to explore the high-quality solution region in the global range, and particle swarm optimization is used to perform fine search in the region found by quantum annealing, improving the accuracy of the solution.
[0121] In the quantum annealing part, the scheduling problem is mapped to the spin state space of quantum bits, and by setting appropriate Hamiltonians and annealing schedules, the system gradually converges to a low-energy state, i.e. a better scheduling solution, under the action of quantum fluctuations. Then, the solution obtained by quantum annealing is used as the initial population of the particle swarm optimization algorithm, and the velocity and position update rules of the particles are set to make them evolve towards better scheduling solutions in the solution space. Throughout the process, the quality of the solution is evaluated according to the scheduling objective function (such as manufacturing cycle, resource utilization, etc.), guiding the search direction of the algorithm.
[0122] The hybrid model combines the global search ability of quantum annealing algorithm and the local optimization ability of particle swarm optimization algorithm, and can more effectively search for the optimal scheduling scheme in the complex scheduling problem space. For the tunnel boring machine collaborative manufacturing scheduling problem with a large number of constraints and complex objectives, this hybrid algorithm can improve the quality of the scheduling scheme and reduce the calculation time.
[0123] Example 3
[0124] (I) Sensor data acquisition and transmission
[0125] According to the structure and operation characteristics of the shield machine, the installation positions of various sensors are determined. For example, displacement sensors are installed at key positions of the cutter head, thrust cylinder, shield body, etc. of the shield machine to accurately measure the displacement at different positions; angle sensors are installed near the rotating parts of the shield machine to monitor the rotation angle; pressure sensors are installed on the hydraulic system pipelines to obtain pressure data; speed sensors are installed at the traveling mechanism of the shield machine to monitor the traveling speed.
[0126] The parameters of each sensor are configured, including sampling frequency, measurement range, accuracy, etc. According to the speed of change of the shield machine operation data and the requirement for accuracy, the sampling frequency is set reasonably. For example, for data with relatively slow displacement and angle changes, the sampling frequency can be set to 1-10 times per second; while for data with relatively fast changes such as speed and pressure, the sampling frequency can be set to 10-100 times per second. At the same time, according to the range of actual operation parameters of the shield machine, the measurement range of each sensor is determined, and a sensor with appropriate accuracy is selected to ensure that the collected data can accurately reflect the operation state of the shield machine.
[0127] A transmission mode combining wired and wireless networks is established. Inside the shield machine, for sensors and local data acquisition nodes that are close to each other and require high data transmission, wired networks (such as Ethernet) are used for connection to ensure the stability and high speed of data transmission. For long-distance transmission between the shield machine and the central server, considering the complexity of the construction environment, wireless networks (such as 4G / 5G networks) are used for data transmission. The data transmission protocol is configured to ensure that the collected data can be accurately and completely transmitted to the central server. For example, TCP / IP protocol is used for network communication, the collected data is encapsulated at the data sending end, a packet header (containing source address, destination address, data length, etc.) and a check code (used to detect the integrity of data during transmission) are added, and the received data is unpacked and checked at the receiving end. If the check fails, the data is requested to be resent.
[0128] (II) Data cleaning and feature extraction
[0129] Invalid data removal: Identify invalid data by setting reasonable range of data and logical rules. For example, for displacement sensor data, if its value exceeds the actual possible displacement range of the shield machine (determined according to the design size of the shield machine and construction conditions), it is determined that the data is invalid and is removed. For pressure data, if its value is negative or exceeds the normal working pressure range of the hydraulic system, it is also considered as invalid data for processing.
[0130] Redundant data processing: Analyze the collected data sequence and find data segments that are almost constant or change very little at consecutive sampling points. These data may be redundant data due to the precision limit of the sensor or the relatively stable collection environment. For such redundant data, you can choose to retain some key data points (such as retaining one data point every certain time interval) or directly delete them according to specific circumstances to reduce the burden of data storage and processing.
[0131] Missing value processing: When detecting that there are missing values in the data, different processing methods are used according to the characteristics and missing conditions of the data. If the missing values are few and the data changes relatively smoothly, linear interpolation can be used to fill them, that is, the estimated value of the missing value is calculated according to the data points before and after the missing value through linear relationship. If the missing values are more and the data has certain periodicity or correlation, time series analysis-based methods (such as moving average method, seasonal decomposition method, etc.) can be used to fill them. If the missing value has little effect on subsequent analysis, you can also choose to delete the record corresponding to the missing value.
[0132] Feature extraction based on physical meaning: Directly extract physical features directly related to the running state of the shield machine, such as displacement, angle, speed, pressure, etc. mentioned in the previous text. These features can directly reflect the working state of the shield machine under different working conditions and are an important basis for subsequent model analysis.
[0133] Derivative feature extraction: In addition to basic physical features, some derivative features can be extracted by mathematical operation and combination of original data. For example, calculate the rate of change of displacement (i.e. the approximate value of speed), which is obtained by dividing the displacement difference between adjacent time points by the time interval. It can reflect the acceleration of the shield machine; calculate the ratio of pressure to flow (assuming that flow data can be obtained through other sensors or known conditions), which can reflect the working efficiency of the hydraulic system to some extent. These derivative features can provide more abundant information and help to better understand the running characteristics of the shield machine.
[0134] (Three) Model construction and training
[0135] 1. Deep learning model based on variational autoencoder (VAE):
[0136] Determine the input layer dimension of the VAE, which is consistent with the dimension of the feature vector after data cleaning and feature extraction. For example, if 10 features including displacement, angle, speed, pressure, etc. are extracted, the input layer node number is set to 10. Design the encoder and decoder structure of the VAE. The encoder usually consists of multiple hidden layers, each hidden layer uses different number of neurons (such as the first layer hidden layer is set to 128 neurons, the second layer hidden layer is set to 64 neurons, etc.), through the full connection layer and the activation function (such as ReLU activation function) to the input data is encoded step by step, the high-dimensional input data is compressed into low-dimensional latent variable representation (i.e. mean and variance vector). The decoder is the inverse process of the encoder, which restores the latent variable through multiple hidden layers to the same dimension output data as the input data.
[0137] The loss function of VAE consists of two parts: reconstruction loss and KL divergence loss. The reconstruction loss is used to measure the difference between the decoder output data and the original input data, and the mean square error (MSE) is usually used as the reconstruction loss function, that is, the average value of the square sum of the difference between the output data and the input data corresponding elements. The KL divergence loss is used to measure the difference between the latent variable distribution and the prior distribution (usually assumed to be a standard normal distribution), which is obtained by calculating the KL divergence between the mean and variance of the latent variable and the mean and variance of the standard normal distribution. The total loss function is the weighted sum of the reconstruction loss and the KL divergence loss, and by adjusting the weight parameter, the influence of the two in the training process can be balanced.
[0138] Use the labeled data (i.e. data labeled with shield machine posture changes) for training. Divide the data into training set and validation set according to a certain proportion (such as 70% for training, 30% for validation). In the training process, the feature vector in the training set is input into the VAE model, the gradient of the loss function with respect to the model parameters (such as the weights and biases of the neurons in each hidden layer) is calculated by the back propagation algorithm, and then the model parameters are updated according to the gradient descent algorithm (such as stochastic gradient descent algorithm and its variants, such as Adagrad, Adadelta, etc.), so that the loss function gradually decreases. After each training epoch (one complete traversal of the training data set), the loss function value and other evaluation indicators (such as reconstruction error, reasonableness of latent variable distribution, etc.) on the validation set are used to judge the training effect of the model, and the training parameters (such as learning rate, weight decay, etc.) are adjusted until the model reaches good performance on the validation set (such as the loss function value converges to a smaller value and does not decrease significantly).
[0139] Graph neural network model based on graph attention network (GAT):
[0140] According to the relationship between the component structure and the operation data of the shield machine, a graph structure is constructed. Each component (such as a cutter head, a pushing oil cylinder, a shield body, etc.) of the shield machine is regarded as a node of the graph, the connection relationship (such as a hydraulic transmission relationship, a mechanical transmission relationship, etc.) between the components and the correlation between the data (such as the relationship between the displacement data at different positions and the overall attitude) are regarded as edges of the graph. Each edge is assigned a corresponding weight, and the determination of the weight can be determined according to factors such as the importance between the components, the strength of the data correlation, etc. For example, a higher weight is assigned to the connection relationship between the components that have a greater influence on the attitude of the shield machine. The node feature vector of the graph is determined, and the dimension is consistent with the dimension of the feature vector after data cleaning and feature extraction, that is, the displacement, angle, speed, pressure, etc. extracted are taken as the node feature vector.
[0141] The core of the GAT is the attention mechanism, which calculates the attention degree of each node to other nodes. In the model architecture, multiple attention layers are set, and each attention layer contains multiple head attention mechanisms (such as 8 head attention mechanisms). The attention mechanism of each head calculates the attention weight between nodes, and performs weighted summation on the feature vector of the node to obtain a new feature vector. After processing by multiple attention layers, the information interaction and feature extraction ability between nodes are gradually enhanced.
[0142] After the last attention layer, a fully connected layer is set to map the processed node feature vector to the output layer, and the number of nodes of the output layer is determined according to the dimension of the predicted attitude change of the shield machine. For example, if the attitude change of the shield machine in the three-dimensional space is to be predicted, the number of nodes of the output layer can be set to 3. The cross-entropy loss function is used as the loss function of the GAT model. When predicting the attitude change of the shield machine, the prediction result is compared with the labeled actual attitude change, the cross-entropy between the prediction result and the actual result is calculated, the smaller the cross-entropy, the closer the prediction result to the actual result, and the better the performance of the model.
[0143] The labeled data is also divided into a training set and a validation set according to a certain proportion (such as 70% for training and 30% for verification).
[0144] During the training process, the constructed graph structure and the corresponding node feature vectors are input into the GAT model. The loss function is calculated by the backpropagation algorithm to obtain the gradient of the model parameters (such as the weights of the attention layer, the weights and biases of the fully connected layer, etc.). Then, the model parameters are updated according to the gradient descent algorithm, so that the loss function gradually decreases. After each training epoch, the training effect of the model is judged according to the loss function value and other evaluation indicators (such as prediction accuracy, recall rate, etc.) on the validation set. The training parameters (such as learning rate, number of attention layers, etc.) are adjusted until the model achieves good performance (such as the prediction accuracy reaches a certain threshold and does not decrease significantly) on the validation set.
[0145] 2. Integrated learning model based on mixed expert system (MoE):
[0146] Multiple expert networks of different types are constructed, each of which can be based on different machine learning algorithms or deep learning architectures. For example, an expert network based on decision trees, an expert network based on support vector machines, an expert network based on convolutional neural networks, etc. Each expert network learns and predicts different aspects or features of the tunneling machine operation data.
[0147] The input feature vector of each expert network is determined, which has the same dimension as the feature vector after data cleaning and feature extraction. Each expert network processes and learns the input feature vector differently according to its own characteristics and data processing method.
[0148] A gating network is constructed, which determines the participation degree of each expert network in the prediction process according to the features of the input data and the current operation situation. The gating network usually consists of multiple hidden layers, which process the input data through fully connected layers and activation functions (such as ReLU activation function) to output the weight coefficients of each expert network. These weight coefficients represent the importance of each expert network in the integrated prediction.
[0149] In prediction, first, the feature vector after data cleaning and feature extraction is input into the gating network, and the gating network calculates the weight coefficient of each expert network according to the input data. Then the same feature vector is input into each expert network respectively, and each expert network outputs its own prediction result according to its own learning and prediction ability. Finally, the prediction results of each expert network are weighted and summed according to the weight coefficients given by the gating network to obtain the final integrated prediction result. The mean square error (MSE) is used as the loss function of the MoE model. The final integrated prediction result is compared with the labeled actual result, and the mean square error between them is calculated. The smaller the MSE, the closer the integrated prediction result is to the actual result, and the better the performance of the model. The labeled data is divided into training set and validation set according to a certain proportion (such as 70% for training and 30% for validation). In the training process, the feature vectors in the training set are input into the gating network and each expert network in turn, and the gradient of the loss function with respect to the model parameters (such as the weights and biases of the hidden layers of the gating network, the weights and biases of the expert networks, etc.) is calculated by the back propagation algorithm. Then the model parameters are updated according to the gradient descent algorithm, so that the loss function gradually decreases. After each training epoch, the training effect of the model is judged according to the loss function value and other evaluation indicators (such as the accuracy and stability of integrated prediction) on the validation set, and the training parameters (such as learning rate, structure of gating network, etc.) are adjusted until the model achieves good performance (such as the accuracy of integrated prediction reaches a certain threshold and does not decrease significantly) on the validation set.
[0150] (iv) Posture control
[0151] During the operation of the shield machine, sensors continuously collect data and transmit them to the central server. At the central server end, the real-time received data is quickly processed, key features (such as displacement, angle, speed, pressure, etc.) are extracted, and compared with the pre-set normal operation range. For example, for displacement data, a reasonable displacement interval is set, if the real-time monitored displacement value exceeds this interval, it is determined that the posture of the shield machine may have changed, which needs to be further analyzed.
[0152] Using data visualization technology, the key operation data and state information of the shield machine are displayed in real time on the visualization interface of the human-computer interaction module, so that the on-site engineers can intuitively understand the real-time operation of the shield machine. For example, through the form of instrument panel, line chart, etc. to display the real-time change of displacement, angle, speed, pressure and other data, and through different colors or icon marks to mark the normal and abnormal state of the shield machine, etc.
[0153] The real-time monitoring extracts the key feature data into the already trained model (such as VAE, GAT, MoE-based model), and the model predicts the attitude change of the shield machine in the future (such as the next 10 minutes, 30 minutes, etc.) according to the input data and its learning ability. For example, predict the rotation angle change of the cutter head, the advance speed change of the advance cylinder, the displacement change of the shield body, etc., and the overall attitude change caused by these changes.
[0154] In order to improve the accuracy and reliability of the prediction, the prediction results of the model are fused. If multiple different models are used for prediction at the same time, weighted average, Bayesian fusion, etc. Method can be used to integrate the prediction results of each model to get a comprehensive prediction result. For example, suppose the prediction result based on the VAE model is A, the prediction result based on the GAT model is B, and the prediction result based on the MoE model is C. Using the weighted average method, suppose the weights are w1, w2, w3 (and w1+w2+w3=1), then the comprehensive prediction result D=w1A+w2B+w3C.
[0155] According to the results of attitude prediction and the pre-set optimization strategy, specific control measures are formulated to adjust the attitude of the shield machine. For example, if it is predicted that the rotation angle of the cutter head of the shield machine will exceed the normal range, which may cause uneven excavation, then the rotation speed of the cutter head drive motor is adjusted to make the rotation angle return to the normal range. If it is predicted that the advance speed of the advance cylinder is too fast, which may cause the attitude of the shield machine to be unstable, then the oil supply pressure of the advance cylinder is reduced to slow down the advance speed.
[0156] These control measures are executed through the control system of the shield machine (such as hydraulic control system, electrical control system, etc.) to realize real-time adjustment of the attitude of the shield machine. During the execution of the control measures, the running state of the shield machine is continuously monitored to observe whether the attitude is adjusted as expected. If the expected effect is not achieved, the prediction results and control strategies are re-evaluated for further adjustment.
[0157] A large amount of operation data is collected from multiple different shield machine construction projects, covering different geological conditions (such as soft soil stratum, hard rock stratum, composite stratum, etc.), different construction stages (such as starting, tunneling, receiving, etc.), and different equipment working conditions (such as different cutter head rotation speeds, advance speeds, etc.). The collected data can fully reflect the various situations of the shield machine in actual construction. Each set of data collected is divided according to the ratio of 80% for training and testing and 20% for final verification. Among them, the training set is used for the initial training of the model, the test set is used for the preliminary evaluation and adjustment of the model performance during the training process, and the verification set is used for the final comparison of the performance of different models.
[0158] Model comparison selection: Model one: traditional experience model: a model based on traditional engineering experience and simple mathematical formula. For example, by simple statistical analysis of historical construction data, a linear regression model is established to predict the attitude change of the shield machine, and a fixed control threshold and strategy are set according to experience for attitude control. This model does not involve complex deep learning or ensemble learning mechanism. Model two: single deep learning model (taking VAE as an example): only a deep learning model based on variational autoencoder (VAE) is used for attitude prediction and control. According to the VAE model construction and training method described above, operation is carried out, and no other model is fused or integrated. Model three: ensemble learning model (MoE+GAT+VAE fusion): an ensemble learning model based on hybrid expert system (MoE), combined with a graph neural network model based on graph attention network (GAT) and a deep learning model based on variational autoencoder (VAE) for fusion prediction. The prediction results of the three models are fused by weighted average method to obtain the final attitude prediction result, and the attitude control is carried out according to the result. The specific fusion weight can be adjusted and optimized according to the performance on the test set.
[0159] Evaluation metrics, accurately calculate the proportion of samples that correctly predict the shield machine pose changes in the total number of samples. For each sample, compare the model's predicted pose changes (including displacement, angle, velocity, etc.) of the shield machine at a specific future time (such as the next 10 minutes) with the actual pose changes that occur. If the predicted value is within a pre-set error range (such as displacement error not exceeding ±5 cm, angle error not exceeding ±2 degrees, velocity error not exceeding ±0.1 m / s, etc.) of the actual value, it is considered correct prediction. By comparing all validation set samples in this way, the prediction accuracy index is obtained, and the higher the index, the stronger the model's prediction ability for shield machine pose changes. For each sample, calculate the squared difference between the model's predicted shield machine pose change value (such as predicted displacement value, angle value, velocity value, etc.) and the actual value. Then take the average of the squared differences of all validation set samples to get the mean squared error (MSE). The smaller the MSE value, the higher the prediction accuracy of the model, that is, the closer the predicted value is to the actual value. Pose stability: Observe the shield machine's pose fluctuations during the entire construction process after controlling according to the model's prediction results. Quantify the stability of the pose by calculating the standard deviation of the pose parameters (such as displacement, angle, velocity, etc.). The smaller the standard deviation, the more stable the shield machine's pose and the better the control effect. For example, within a continuous construction time, calculate the standard deviation of the displacement data, and if the value is less than a certain threshold (such as 5 cm), it indicates that the displacement pose is relatively stable. Construction quality indicators: Evaluate the construction quality based on the actual results after construction. For example, for tunnel excavation accuracy, the deviation between the actual tunnel profile and the designed profile can be measured to evaluate the accuracy, and the smaller the deviation, the higher the excavation accuracy; for tunnel flatness, a laser flatness instrument can be used to measure the flatness of the tunnel inner surface, and the smaller the flatness value, the better the construction quality. Combine these construction quality indicators to give a quantitative construction quality score, with a full score of 10 points, and deduct points according to the actual deviation. Safety risk assessment: Monitor whether there are any potential safety issues during construction, such as tunnel collapse, abnormal wear of shield machine components, groundwater leakage, etc. If there are no safety problems, the safety risk assessment score is 10 points; if there are minor safety hazards (such as local small-scale groundwater leakage), the score is 6 points; if there are serious safety problems (such as tunnel collapse risk or critical component severe wear), the score is 2 points. Combine the pose stability, construction quality indicators, and safety risk assessment scores to get the comprehensive score of the control effect evaluation, and the higher the score, the better the control effect. Record the computing resources consumed by each model during training and prediction, including CPU usage, memory usage, and the time required for training or prediction. By comparing the computing resource consumption of different models under the same data volume and hardware environment, evaluate the running efficiency of the models.Lower computational resource consumption means that the model can run on more common hardware devices, or can complete training and prediction tasks faster on the same hardware, with better practicality.
[0160] The test results show that the prediction accuracy of Model One (traditional empirical model) is about 55% on average. This is because the traditional empirical model is based on simple linear relationships and fixed thresholds, making it difficult to accurately capture the complex nonlinear relationships in the shield machine operation data and the changing patterns under different working conditions. The prediction accuracy of Model Two (single deep learning model - VAE) is about 70% on average. Although the VAE model can learn some nonlinear features in the data, the limitations of a single model lie in its inability to fully consider various factors involved in the shield machine posture changes, such as the mutual influence between different components and the comprehensive effects of complex factors such as geological conditions. The prediction accuracy of Model Three (ensemble learning model - MoE+GAT+VAE fusion) is about 85% on average. By integrating different types of models and performing fusion prediction, the advantages of each model can be fully utilized, such as the expert system of MoE which can effectively process different aspects of data, GAT which can better capture the relationships in graph structure data, and VAE which can learn the underlying features of the data, thus more accurately predicting the posture changes of the shield machine. The mean square error (MSE) of Model One is about 0.25 on average. Due to its relatively low prediction accuracy, the deviation from the actual value is large, resulting in a high MSE value. The MSE of Model Two is about 0.15 on average. Although it has decreased compared to the traditional empirical model, there is still some prediction error, indicating that there is room for improvement in the single deep learning model in accurately fitting the data. The MSE of Model Three is about 0.08 on average. By integrating the advantages of multiple models, the predicted value is closer to the actual value, effectively reducing the mean square error, and demonstrating the advantages of ensemble learning models in improving prediction accuracy. The shield machine posture stability under Model One control is poor, with standard deviations of displacement, angle, and velocity of about 8 cm, 3 degrees, and 0.2 m / s on average. This is because the traditional empirical model is difficult to dynamically adjust the control strategy based on real-time data, resulting in large posture fluctuations. The shield machine posture stability under Model Two control has improved, with standard deviations of posture parameters of about 5 cm, 2 degrees, and 0.15 m / s on average. Although the VAE model can provide some prediction information for control, the limitations of a single model leave room for improvement in control effectiveness. The shield machine posture stability under Model Three control performs well, with standard deviations of posture parameters of about 3 cm, 1 degree, and 0.1 m / s on average. The ensemble learning model can develop more effective control strategies by more accurately predicting and comprehensively considering various factors, thus making the shield machine posture more stable. Construction quality indicators: The construction quality indicator score corresponding to Model One is about 6 on average. Due to its inaccurate control of the shield machine posture, there is a certain deviation in excavation accuracy and tunnel flatness, affecting the construction quality. The construction quality indicator score corresponding to Model Two is about 7 on average. The application of the VAE model has improved the construction quality to some extent, but it has not yet reached the best effect.The average score of the construction quality indicators corresponding to Model Three is about 9. The ensemble learning model effectively improves the excavation accuracy and tunnel flatness and other construction quality indicators by better controlling the shield machine posture, resulting in a significant improvement in construction quality. The safety risk assessment score of Model One is about 6 on average. The limitations of the traditional experience model may lead to safety hazards in some complex working conditions, such as improper posture control that may cause tunnel collapse risks or abnormal wear of components. The safety risk assessment score of Model Two is about 7 on average. Although the VAE model helps to reduce safety risks, a single model still has limitations in dealing with complex situations. The safety risk assessment score of Model Three is about 9 on average. The ensemble learning model can minimize safety risks and ensure construction safety by accurately predicting and effectively controlling the shield machine posture. The comprehensive control effect evaluation score of Model One is about 6 on average. Considering the posture stability, construction quality indicators, and safety risk assessment scores, the traditional experience model performs poorly in overall control effect. The comprehensive control effect evaluation score of Model Two is about 7 on average. Although the single deep learning model has improved in some aspects, the overall control effect still needs to be improved. The comprehensive control effect evaluation score of Model Three is about 9 on average. The ensemble learning model performs well in prediction accuracy and control effect by integrating the advantages of multiple models, effectively controlling the shield machine posture, improving construction quality, and ensuring construction safety. Model One consumes relatively less computing resources during training and prediction, with an average CPU usage of about 20% and an average memory usage of about 2GB, and the training or prediction time is relatively short, averaging about 10 minutes. This is because the traditional experience model has lower computational complexity and does not require extensive data learning and complex calculations. Model Two (VAE) has an average CPU usage of about 60% and an average memory usage of about 8GB during training, and the training time is relatively long, averaging about 60 minutes. Due to the need for extensive data learning and complex neural network operations, deep learning models consume more computing resources. Model Three (MoE+GAT+VAE fusion) has an average CPU usage of about 80% and an average memory usage of about 12GB during training, and the training time is the longest, averaging about 90 minutes. The ensemble learning model involves the fusion and training of multiple models, resulting in higher computational complexity and more computing resource consumption. However, despite the higher resource consumption, the significant advantages in prediction accuracy, control effect evaluation, and other aspects make it worthwhile, as it can provide more accurate posture prediction and more effective control for shield machine construction, thereby improving construction quality and safety.
[0161] Through the above comparative synergistic test, it can be clearly seen that the scheme of integrating learning model based on mixed expert system (MoE) and combining graph neural network model based on graph attention network (GAT) and deep learning model based on variational autoencoder (VAE) for fusion prediction has obvious advantages in shield machine posture prediction, control effect, construction quality and safety guarantee and the like, although the calculation resource consumption is relatively high, but the comprehensive benefit is significantly better than the traditional experience model and the single deep learning model.
[0162] The above has made a detailed description of the present application in combination with the embodiments, but those skilled in the art can understand that various specific parameters in the above embodiments can be changed to form multiple specific embodiments without departing from the purpose of the present application, which are within the common variation range of the present application, and will not be described one by one in detail.
Claims
1. A tunnel boring machine collaborative manufacturing intelligent scheduling method, characterized in that, The method comprises a data acquisition and preprocessing step, an intelligent scheduling modeling and solving step, a monitoring and feedback optimization step; the data acquisition and preprocessing step acquires multi-source heterogeneous data from each link of the tunnel boring machine manufacturing and performs preprocessing; the intelligent scheduling modeling and solving step solves a static scheduling problem by using a mixed integer programming model and introduces a deep reinforcement learning algorithm for dynamic optimization, and simultaneously uses a robust optimization method to enhance the robustness of the scheduling scheme; the monitoring and feedback optimization step performs real-time monitoring on the manufacturing process and optimizes the scheduling scheme according to feedback information; further comprising training an echo state network, taking historical manufacturing data as input, and outputting a prediction of the future manufacturing state; Then, in the ant colony optimization algorithm, the ants select the next task and resource for scheduling decision according to the pheromone concentration and the prediction result of the ESN. After each iteration, the pheromone concentration is updated according to the quality of the scheduling scheme, and the ESN is used to update the prediction of the new manufacturing state, guiding the ant colony to search for a better scheduling scheme; the intelligent scheduling modeling and solving step decomposes the manufacturing task into subtasks by using a mixed integer programming model, establishes an allocation matrix between the tasks and resources, solves an initial scheduling scheme by using a multi-objective function, the deep reinforcement learning algorithm performs dynamic optimization according to real-time state information, and the robust optimization method considers uncertain parameters to enhance the stability of the scheduling scheme.
2. The method of claim 1, wherein, The data acquisition and preprocessing step comprises acquiring data from the design, process planning, production manufacturing and material management links, preprocessing the data by using an ontology-based data fusion method, and processing the data by using an outlier detection algorithm to ensure reliability.
3. The method of claim 1, wherein, The monitoring and feedback optimization step establishes a real-time monitoring system to display key information on a visual interface, collects feedback information including worker operation feedback, quality problem feedback and equipment maintenance records, and optimizes the scheduling scheme and updates the scheduling algorithm and model according to the feedback information.
4. The method of claim 1, wherein, Further comprising encoding the equipment failure state into a binary vector and representing the production progress deviation as a percentage value, selecting actions by using a deep Q network, the actions including adjusting the process sequence, replacing the processing equipment, delaying or advancing the task start time; according to the execution result of the action, the environment gives the agent corresponding reward feedback, and the reward function comprehensively considers the change of the production cycle, the improvement of the equipment utilization rate and the reduction of the quality risk factors.
5. An intelligent dispatch system for implementing the method of any one of claims 1-4, characterized by The method comprises a data acquisition and integration module, an intelligent scheduling decision module and a monitoring and feedback module; the data acquisition and integration module is used for acquiring multi-source heterogeneous data from each link of the tunnel boring machine manufacturing, and performing integration and preprocessing; the intelligent scheduling decision module uses an algorithm combining mixed integer programming and deep reinforcement learning to make scheduling decisions, and considers uncertain factors for robustness enhancement; the monitoring and feedback module performs real-time monitoring on the manufacturing process and collects feedback information to optimize the scheduling scheme.
6. The intelligent dispatch system of claim 5, wherein, The manufacturing process is regarded as a dynamic system, and the ESN is used to model the manufacturing process and learn the variation law of the manufacturing task and resource in the time sequence.
7. The intelligent dispatch system of claim 5, wherein, The mixed integer programming model of the intelligent scheduling decision module takes minimizing production cycle, maximizing equipment utilization, and reducing production cost as the objective function, considers order delivery period, equipment capacity, personnel working hours, and material supply constraint conditions to solve the initial scheduling scheme, the deep reinforcement learning algorithm dynamically optimizes the initial scheme, and the robust optimization method enhances the anti-interference ability of the scheduling scheme to uncertain factors.
8. The intelligent dispatch system of claim 5, wherein, The monitoring and feedback module displays production progress, equipment running state, and quality detection result key information in a visual interface, collects feedback information in the manufacturing process, and feeds back to the intelligent scheduling decision module to optimize the scheduling algorithm and model.
Citation Information
Patent Citations
Dynamic collaborative scheduling method for smart factory
CN116880396A
Intelligent factory decision support system and method thereof
CN118982437A
Industrial production intelligent scheduling system
CN119005609A