A production line fault prediction method and system based on large language model complex signal reasoning
Patent Information
- Application Number
- CN202611029723.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-10
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2046-07-10
AI Technical Summary
[0005]有鉴于此,本发明提供了一种基于大语言模型复杂信号推理的生产线故障预测方法及系统,有助于解决现有技术中针对生产线监控预测模型过度依赖结果评估、缺乏严密推理步骤验证的问题
本发明方法评估科学性强,通过严密的MCTS框架和DAPO优化算法,成功解决了评估模型缺乏推导约束的缺陷;本发明方法在训练数据合成方面,深入利用了节点扩展与反向传播的数学关系,计算准确的节点期望收益,提供了更高质量的监督信号;本发明方法克服了基于结果的评估指标常常高估模型能力的缺陷,实现了步骤级的深度验证,并利用算法对模型性能进行精确的归因分配,对大语言模型的发展及优化具有重要的指导价值。
Smart Images

Figure CN122529111B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the intersection of industrial control and artificial intelligence, and more specifically to a method and system for predicting production line faults based on complex signal reasoning using a large language model. Background Technology
[0002] With the increasing popularity of Large Language Models (LLMs) in the field of intelligent manufacturing, they are gradually being widely applied to the operation monitoring and predictive maintenance of production line equipment. In these industrial scenarios oriented towards monitoring and prediction, solving complex problems usually requires a close combination of physical intuition and rigorous mathematical derivation, forming a long chain of reasoning.
[0003] However, existing evaluation methods for such large-scale industrial forecasting models have significant drawbacks: they rely excessively on indicators based on the final prediction results. A correct final prediction does not necessarily indicate an effective reasoning and diagnostic process, and incorrect predictions often contain a large amount of correct intermediate analysis work. This makes it difficult for traditional evaluation methods to distinguish between "real fault diagnosis reasoning" and "correct alarms obtained due to coincidence or incorrect causes," hindering the provision of reliable capability assessments and system-level diagnostic analysis in actual production line supervision. Furthermore, existing process reward models (PRM) face two major challenges in evaluating actual production line forecasting systems: first, the diagnostic trajectories generated by different large models vary greatly in format and granularity, making unified step-level verification difficult; second, the annotation cost of complex signal processing step-level supervision data is extremely high.
[0004] Therefore, how to provide an automated process-level assessment and prediction method that can fit the actual situation of industrial production supervision, statistical algorithms and formulas, has high reliability and can provide fine-grained diagnostic analysis is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0005] In view of this, the present invention provides a method and system for predicting production line faults based on complex signal reasoning of large language models, which helps to solve the problems of over-reliance on result evaluation and lack of rigorous reasoning verification in existing production line monitoring and prediction models.
[0006] To achieve the above objectives, the present invention adopts the following technical solution: This invention discloses a production line fault prediction method based on complex signal reasoning using a large language model, comprising: S1. Collect raw data and problems of production line equipment, preprocess them, and build a benchmark test set; S2. Based on the benchmark test set, use the Monte Carlo tree search algorithm to generate step-level inference trajectory sample data to obtain the training dataset; S3. Perform multi-view data filtering on the sample data of the training dataset; S4. Use the training samples filtered by S3 to train the reward model during the process. S5. Use the large language model to predict production line failures on the samples of the benchmark test set, obtain the original inference prediction trajectory generated by the large language model, and convert it into a standardized sequence of steps. S6. Input the standardized step sequence into the process reward model trained in S4 to obtain the correctness label and evaluation reason for each step; S7. Based on the verification results of S6, the comprehensive reasoning ability of the large language model in production line supervision and prediction is quantitatively evaluated. S8. Based on the process completion rate and the first fatal error attribution result of the quantitative evaluation output in S7, the large language model is optimized in a targeted manner, and the optimized large language model is deployed in the edge computing device of the production line to perform real-time online inference on the original sensor complex signals of the production line equipment, so as to realize the production line fault prediction.
[0007] Furthermore, the preprocessing in S1 includes: cleaning the original data and questions, removing irrelevant data; then performing standardization and multimodal alignment; and finally performing domain expert verification to obtain the inference diagnosis step data.
[0008] Furthermore, the specific steps of S2 are as follows: S2.1: Starting from the root node, select child nodes using the UCT rules. The selected formula is: ; in, This represents the best child node selected using the UCT rule. Represents a node v The set of child nodes This indicates the number of times a node has been visited. Represents child nodes u The current expected accuracy estimate, An exploration coefficient is used to control the intensity of exploration. S2.2: Based on selected child nodes The next step, expanding to K candidates, is formulated as follows: ; in, Indicates the first k The next step for each candidate Representational strategy large language model, Indicates input m Multimodal raw data or multimodal context conditions, This represents the complex signal problem proposed for production line failure prediction. This represents the distance from the root node to the selected child node. The historical reasoning trajectory; each new step creates a child node. u k The corresponding part of the reasoning trajectory is obtained through a concatenation operation: ; S2.3: For each extended node... u Supplement M A complete trajectory, the formula is: ; in, Indicates the first The complete reasoning trajectory needs to be supplemented. This indicates the node to be simulated for expansion. u Part of the reasoning trajectory; Extract the final answer and the benchmark answer Compare and compute nodes u The expected accuracy is calculated and converted into a binary hard label to be assigned to the node, using the following formula: ; ; in, M Represents the total number of simulated trajectories or Rollout Number of samplings; This indicates an indicator function that takes the value 1 when the condition inside the parentheses is true, and 0 otherwise. This represents the function for checking and verifying the consistency of the answer. Indicates assignment to a node u Binary hard tags; S2.4: For each ancestor node of the extended node, update its visit count and recalculate the expected accuracy using the following formula: ; in, Represents ancestor nodes v Number of visits, This indicates the subset of terminal nodes whose final answer has been verified as correct.
[0009] Furthermore, the Monte Carlo tree search algorithm incorporates textbook-guided solutions and self-reflection algorithms. The textbook's answer guidance specifically refers to providing expert-verified solutions. The reasoning is divided into blocks with a unified objective, as shown in the formula: The expert trunk branch in the search tree is initialized based on the reasoning block with a unified goal, and Monte Carlo tree search expansion is performed on the expert trunk branch. The self-reflection algorithm specifically involves: before extending to the next step in S2.2, first analyzing the historical partial trajectory... Generate reflection content The formula is: Then, the reflection content will be... Introduced as additional context, new candidate steps are generated, as shown in the formula: .
[0010] Furthermore, the training of the process reward model in S4 includes supervised fine-tuning and decoupled pruning and dynamic sampling strategy optimization stages; The overall reward function in the decoupled pruning and dynamic sampling strategy optimization phase The formula is: ; in, Indicates an accuracy-based bonus. Indicates a formatted reward. and These represent the corresponding coefficients; The objective function for maximizing the decoupled pruning and dynamic sampling strategy optimization phase is expressed as: ; in, Let represent the objective function for optimizing the decoupling pruning and dynamic sampling strategies. Represents the mathematical expectation. The sample set representing the prefix inference trajectory. This represents a dataset of step-level trajectory samples generated by the Monte Carlo tree search algorithm. This represents the set of candidate responses within the same sampled group. This represents the total number of candidate responses sampled within each group. j Indicates the index of candidate responses within the group. This represents the old policy model before the update. Indicates the first j The total number of intermediate steps contained in each candidate response. Indicates the current policy model at the th... j The first trajectory t The step-level reward value output by each step. This represents a clipping function that restricts input values to a specified range. This indicates the lower bound of the clipping hyperparameter boundary. This indicates the upper bound of the clipping hyperparameter boundary. This represents the group advantage calculated based on the normalized mean and standard deviation of the reward. ; in, Indicates the first j The total reward score corresponding to each candidate reasoning trajectory. This represents the average reward score of all candidate inference trajectories in this group. This represents the standard deviation of the reward scores for all candidate inference trajectories in this group.
[0011] Furthermore, in S5, the original inference prediction trajectory is converted into a standardized sequence of steps using a step granularity alignment module; the step granularity alignment module uses a general large language model to perform non-evaluative parsing of the original response, and inserts clear step boundary markers into the original response text according to the principles of single purpose, logical coherence and clear transition.
[0012] Furthermore, the quantitative evaluation in S7 includes: based on the step verification results of S6, calculating and outputting the process completion rate and the attribution of the first fatal error; the process completion rate is the proportion of correct steps before the occurrence of the first fatal error; the attribution of the first fatal error classifies the main failure sources into one of four dimensions: problem understanding, knowledge application, symbolic derivation, and numerical calculation.
[0013] This invention also discloses a production line fault prediction system based on complex signal reasoning using a large language model, comprising: Data acquisition and processing module: Collects raw data and problems from production line equipment, performs preprocessing, and constructs a benchmark test set; Inference trajectory module: Based on the benchmark test set, step-level inference trajectory sample data is generated using the Monte Carlo tree search algorithm to obtain the training dataset; Sample filtering module: performs multi-view data filtering on the sample data of the training dataset; Process reward model module: Trains the process reward model using filtered training samples; Fault Prediction and Inference Sequence Module: Utilizes a large language model to predict production line faults on samples in the benchmark test set, obtains the original inference prediction trajectory generated by the large language model, and converts it into a standardized sequence of steps; Verification module: Input the standardized step sequence into the trained process reward model to obtain the correctness label and evaluation reason for each step; Evaluation module: Based on the step-by-step verification results, quantitatively evaluate the comprehensive reasoning ability of the large language model in production line supervision and prediction; Fault prediction feedback module: It is used to optimize and adjust the large language model based on the process completion rate and the first fatal error attribution result output by the evaluation module, and use the optimized large language model to perform online inference on the real-time sensor signals of the target production line, and output the production line fault prediction result.
[0014] As can be seen from the above technical solution, compared with the prior art, the present invention provides a production line fault prediction method and system based on complex signal reasoning of large language models, which has the following beneficial effects: The method of this invention has strong scientific evaluation. Through the rigorous MCTS framework and DAPO optimization algorithm, it successfully solves the defect of lack of derivation constraints in the evaluation model. In terms of training data synthesis, the method of this invention makes in-depth use of the mathematical relationship between node expansion and backpropagation to calculate accurate expected node returns and provide higher quality supervision signals. The method of this invention overcomes the defect that result-based evaluation metrics often overestimate model capabilities, realizes step-level deep validation, and uses algorithms to accurately attribute model performance, which has important guiding value for the development and optimization of large language models. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0016] Figure 1 This is a schematic diagram of the overall process of the present invention.
[0017] Figure 2 This is a schematic diagram of the hierarchical taxonomy and evaluation criteria provided by the present invention.
[0018] Figure 3 This is a schematic diagram illustrating the operating principle of the data synthesis and filtering module algorithm provided by the present invention. Detailed Implementation
[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] This invention discloses a production line fault prediction method based on complex signal reasoning using a large language model, such as... Figure 1As shown, it includes: S1. Collect raw data and problems of production line equipment, preprocess them, and build a complex signal benchmark test set for production line monitoring and fault prediction. S2. Based on the benchmark test set, the Monte Carlo tree search algorithm is used to generate step-level inference trajectory sample data to obtain the training dataset. S3. Perform multi-view data filtering on the sample data of the training dataset to remove noisy data with low consistency and retain high-quality training samples. S4. Using the training samples filtered by S3, train a process reward model for the production line prediction domain. S5. Use the large language model to predict production line failures on the benchmark test set samples, obtain the original inference prediction trajectory generated by the large language model for production line problems, and convert it into a standardized step sequence. S6. Input the standardized step sequence into the process reward model trained in S4, and verify each intermediate diagnostic reasoning step of the large language model layer by layer to obtain the correctness label and evaluation reason for each step. S7. Based on the verification results of S6, the comprehensive reasoning ability of the large language model in production line supervision and prediction is quantitatively evaluated. S8. Based on the process completion rate and the attribution result of the first fatal error in the quantitative evaluation output of S7, the large language model is optimized in a targeted manner: when the first fatal error is attributed to the problem understanding or knowledge application dimension, the strategy of the large language model is adjusted by introducing relevant textbook domain knowledge or multimodal examples in the context of the prompt words; when the first fatal error is attributed to the symbolic derivation or numerical calculation dimension, the large language model is configured with external computing plugins or symbolic solvers for tool call optimization; subsequently, the optimized large language model is deployed in the edge computing device of the production line to perform real-time online inference on the complex signals of the original sensors of the production line equipment, so as to achieve highly reliable production line fault prediction.
[0021] In a specific embodiment, such as Figure 2 As shown, the preprocessing in S1 includes: cleaning the original data and questions, removing irrelevant data; then standardizing and multimodal alignment; and finally, domain expert verification to obtain the inference and diagnostic step data.
[0022] Specifically, the construction of the benchmark set in S1 includes three stages: S1.1 collecting and cleaning the original materials, and masking irrelevant metadata; S1.2 using a visual language model to transcribe the original documents, mathematical formulas and geometric figures into structured LaTeX markup, and referencing high-resolution graphics through contextual anchors; S1.3 being manually reviewed by domain experts to correct transcription errors and decompose unstructured answers into steps, serving as the gold standard basis for process evaluation.
[0023] In a specific embodiment, such as Figure 3 As shown, the specific steps of S2 are as follows: S2.1 Node Selection: Starting from the root node, select child nodes using UCT rules. The selected formula is: ; in, This represents the best child node selected using the UCT rule. Represents a node v The set of child nodes This indicates the number of times a node has been visited. Represents child nodes u The current expected accuracy estimate, An exploration coefficient is used to control the intensity of exploration. S2.2 Node Expansion and Self-Reflection: Based on Selected Child Nodes The next step, expanding to K candidates, is formulated as follows: ; in, Indicates the first k The next step for each candidate Representational strategy large language model, Indicates input m Multimodal raw data or multimodal context conditions, This represents the complex signal problem proposed for production line failure prediction. This represents the distance from the root node to the selected child node. The historical reasoning trajectory; each new step creates a child node. uk The corresponding part of the reasoning trajectory is obtained through a concatenation operation: ; S2.3 Simulation (Rollout): This involves rotating each extended node... u Supplement M A complete trajectory, the formula is: ; in, Indicates the first The complete reasoning trajectory needs to be supplemented. This indicates the node to be simulated for expansion. u Part of the reasoning trajectory; Extract the final answer and the benchmark answer Compare and compute nodes u The expected accuracy is calculated and converted into a binary hard label to be assigned to the node, using the following formula: ; ; in, M Represents the total number of simulated trajectories or Rollout Number of samplings; This indicates an indicator function that takes the value 1 when the condition inside the parentheses is true, and 0 otherwise. This represents the function for checking and verifying the consistency of the answer. Indicates assignment to a node u Binary hard tags; S2.4 Feedback Propagation: For each ancestor node of the extended node, update its visit count and recalculate the expected accuracy, using the following formula: ; in, Represents ancestor nodes v Number of visits, This indicates the subset of terminal nodes whose final answer has been verified as correct.
[0024] In one specific embodiment, the Monte Carlo tree search algorithm incorporates textbook-guided solutions and self-reflection algorithms; The textbook's answer guidance specifically includes: providing expert-verified solutions. The reasoning is divided into blocks with a unified objective, as shown in the formula: The expert trunk branch in the search tree is initialized based on the reasoning block with a unified goal, and Monte Carlo tree search expansion is performed on the expert trunk branch. The self-reflection algorithm works as follows: before extending to the next step in S2.2, it first considers the historical partial trajectory. Generate reflection content The formula is: Then, the reflection content will be... Introduced as additional context, new candidate steps are generated, as shown in the formula: .
[0025] In a specific embodiment, step-level multi-perspective majority voting filtering is performed on the synthetic data in S3, validating it from five complementary perspectives: formula and arithmetic validity, theoretical and hypothetical basis validity, theoretical and hypothetical advanced validity, global coherence validity, and intra-step objective validity. A majority voting validation integrator composed of open-source and closed-source multimodal models is used; a sample is discarded only if multiple perspectives simultaneously disagree with the automatically generated hard labels.
[0026] In a specific embodiment, the training of the process reward model in S4 includes supervised fine-tuning (SFT) and decoupled pruning and dynamic sampling strategy optimization (DAPO) stages; wherein, supervised fine-tuning refers to: using the high-quality training samples retained after filtering the multi-view data to perform supervised conditional probability alignment training on the initial language model. Specifically, given multimodal input data... Complex signal problems related to production line faults (Q) and historical prefix inference trajectories preceding the current step. Given the conditional context, the verification reasons for the natural language thought chain, whether verified by experts or synthesized with high quality, will be used. Corresponding step-level binary correctness labels As the ground truth, the difference between the model output and the ground truth is minimized through the maximum likelihood estimation (MLE) loss function, so that the trained process reward model can jointly output a logically consistent natural language verification reason given any prefix trajectory. and accurate prediction of binary correctness labels The ability to generate a certain number of times is characterized by the conditional probability generation formula: This supervised fine-tuning phase provides an initial strategy model with baseline alignment and rationale generation capabilities for the subsequent Decoupled Pruning and Dynamic Sampling Strategy Optimization (DAPO) phase.
[0027] The goal of the process reward model is to address a given multimodal problem. and prefix trajectory In this case, generate the reason for the thought chain. And predict step-level correctness labels The formula is: ; The overall reward function in the decoupling pruning and dynamic sampling strategy optimization phase The formula is: ; in, Indicates an accuracy-based bonus. Indicates a formatted reward. and These represent the corresponding coefficients; The objective function for maximizing the decoupled pruning and dynamic sampling strategy optimization phase is expressed as: ; in, Let represent the objective function for optimizing the decoupling pruning and dynamic sampling strategies. Represents the mathematical expectation. The sample set representing the prefix inference trajectory. This represents a dataset of step-level trajectory samples generated by the Monte Carlo tree search algorithm. This represents the set of candidate responses within the same sampled group. This represents the total number of candidate responses sampled within each group. j Indicates the index of candidate responses within the group. This represents the old policy model before the update. Indicates the first j The total number of intermediate steps contained in each candidate response. Indicates the current policy model at the th... j The first trajectory t The step-level reward value output by each step. This represents a clipping function that restricts input values to a specified range. This indicates the lower bound of the clipping hyperparameter boundary. This indicates the upper bound of the clipping hyperparameter boundary. This represents the group advantage calculated based on the normalized mean and standard deviation of the reward. ; in, Indicates the first j The total reward score corresponding to each candidate reasoning trajectory. This represents the average reward score of all candidate inference trajectories in this group. This represents the standard deviation of the reward scores for all candidate inference trajectories in this group.
[0028] In a specific embodiment, in S5, the Step Granularity Alignment (SGA) module is used to convert the original inference prediction trajectory into a standardized sequence of steps. The Step Granularity Alignment (SGA) module uses a general large language model to perform non-evaluative parsing of the original response. Without changing the original inference content, based on the three principles of single purpose, logical coherence and clear transition, a clear step boundary marker "stepn:" is inserted into the original response text.
[0029] In one specific embodiment, the process reward model in S6 receives the multimodal problem and the prefix trajectory. The natural language verification reason xi and the binary correctness label are jointly output. The formula is: .
[0030] In a specific embodiment, the quantitative evaluation in S7 includes: based on the step verification results of S6, calculating and outputting the process completion rate (PCR) and the first fatal error attribution (FFEA); the process completion rate is the proportion of correct steps before the first fatal error occurs, used to quantify the effective reasoning depth that the model can advance before the error occurs; the first fatal error attribution classifies the main failure sources into one of four dimensions: problem understanding, knowledge application, symbolic derivation, and numerical computation.
[0031] Specifically, the rubric-based attribution scoring mechanism in S7 continuously scores the reasoning trajectory across four dimensions: problem comprehension, knowledge application, symbolic derivation, and numerical computation. The evaluator simultaneously calculates Process Completion Rate (PCR) and First Fatal Error Attribution (FFEA). Process Completion Rate is defined as the proportion of correct steps taken before the first fatal error occurs; the First Fatal Error Attribution categorizes the main root causes of reasoning failure into one of the four capability dimensions mentioned above.
[0032] In one specific embodiment, a benchmark test set was constructed and a large amount of trajectory data was synthesized using the algorithm. Ablation experiments and evaluations of multimodal large models were conducted on various key technical modules. The data synthesis configuration and output quality are shown in Table 1, demonstrating that the complete improved algorithm integrating Textbook Guidance (TG) and Self-Reflection (SR) can significantly improve the effective output rate of correct trajectories.
[0033] Table 1. Impact of Improved MCTS Components on Data Synthesis Quality: Ablation Experiment Data Synthesis Configuration
[0034] Regarding the impact of the SGA module, as shown in Table 2, comparing the results of direct verification of long text with verification after SGA alignment, SGA completely eliminates the evaluation bias caused by inconsistent formatting.
[0035] Table 2 Step-Granular Alignment (SGA) Ablation Experiment
[0036] For the Process Reward Model (PRM) that has been trained, this embodiment performs a meta-evaluation on the step correctness classification task. As shown in Table 3, the method of joint training with SFT and DAPO achieves extremely high expert alignment accuracy.
[0037] Table 3 Meta-evaluation of the process reward model in step-level correctness classification
[0038] Finally, the examples verified the alignment between the automated evaluation metrics and human domain expert scores, as shown in Table 4. The multi-dimensional rubric scores proposed in this invention show a high degree of consistency with the results of blind human evaluation, far exceeding the accuracy of OA based solely on the final results.
[0039] Table 4. Alignment between Automatic Evaluation Indicators and Human Expert Scores (Automatic Evaluation)
[0040] This invention also discloses a production line fault prediction system based on complex signal reasoning using a large language model, comprising: Data acquisition and processing module: Collects raw data and problems from production line equipment, performs preprocessing, and constructs a benchmark test set; Inference Trajectory Module: Based on the benchmark test set, step-level inference trajectory sample data is generated using the Monte Carlo tree search algorithm to obtain the training dataset; Sample filtering module: performs multi-view data filtering on the sample data of the training dataset; Process reward model module: Trains the process reward model using filtered training samples; Fault Prediction and Inference Sequence Module: Utilizes a large language model to predict production line faults on samples in the benchmark test set, obtains the original inference prediction trajectory generated by the large language model, and converts it into a standardized sequence of steps; Validation module: Input the standardized sequence of steps into the trained process reward model to obtain the correctness label and evaluation reason for each step; Evaluation module: Based on the step-by-step verification results, quantitatively evaluate the comprehensive reasoning ability of the large language model in production line supervision and prediction; Fault prediction feedback module: It is used to optimize and adjust the large language model based on the process completion rate and the first fatal error attribution result output by the evaluation module, and use the optimized large language model to perform online inference on the real-time sensor signals of the target production line, and output the production line fault prediction result.
[0041] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.
[0042] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A production line fault prediction method based on complex signal reasoning using a large language model, characterized in that, include: S1. Collect raw data and problems of production line equipment, preprocess them, and build a benchmark test set; S2. Based on the benchmark test set, use the Monte Carlo tree search algorithm to generate step-level inference trajectory sample data to obtain the training dataset; S3. Perform multi-view data filtering on the sample data of the training dataset; S4. Use the training samples filtered by S3 to train the reward model during the process. S5. Use the large language model to predict production line failures on the samples of the benchmark test set, obtain the original inference prediction trajectory generated by the large language model, and convert it into a standardized sequence of steps. S6. Input the standardized step sequence into the process reward model trained in S4 to obtain the correctness label and evaluation reason for each step; S7. Based on the verification results of S6, the comprehensive reasoning ability of the large language model in production line supervision and prediction is quantitatively evaluated. S8. Based on the process completion rate and the first fatal error attribution result of the quantitative evaluation output in S7, the large language model is optimized in a targeted manner, and the optimized large language model is deployed in the edge computing device of the production line to perform real-time online inference on the original sensor complex signals of the production line equipment to realize production line fault prediction. The specific steps of S2 are as follows: S2.1: Starting from the root node, select child nodes using the UCT rules. The selected formula is: ; in, This represents the best child node selected using the UCT rule. Represents a node v The set of child nodes This indicates the number of times a node has been visited. Represents child nodes u The current expected accuracy estimate, An exploration coefficient is used to control the intensity of exploration. S2.2: Based on selected child nodes The next step, expanding to K candidates, is formulated as follows: ; in, Indicates the first k The next step for each candidate Representational strategy large language model, Indicates input m Multimodal raw data or multimodal context conditions, This represents the complex signal problem proposed for production line failure prediction. This represents the distance from the root node to the selected child node. The historical reasoning trajectory; each new step creates a child node. u k The corresponding part of the reasoning trajectory is obtained through a concatenation operation: ; S2.3: For each extended node... u To supplement with M complete trajectories, the formula is: ; in, Indicates the first m The complete reasoning trajectory needs to be supplemented. This indicates the node to be simulated for expansion. u Part of the reasoning trajectory; Extract the final answer and the benchmark answer Compare and compute nodes u The expected accuracy is calculated and converted into a binary hard label to be assigned to the node, using the following formula: ; ; in, This indicates the total number of simulated trajectories or the number of Rollout samples; This indicates an indicator function that takes the value 1 when the condition inside the parentheses is true, and 0 otherwise. This represents the function for checking and verifying the consistency of the answer. Indicates assignment to a node u Binary hard tags; S2.4: For each ancestor node of the extended node, update its visit count and recalculate the expected accuracy using the following formula: ; in, Represents ancestor nodes v Number of visits, This indicates the subset of terminal nodes whose final answer was verified as correct; The training of the process reward model in S4 includes supervised fine-tuning and decoupled pruning and dynamic sampling strategy optimization stages; The overall reward function in the decoupled pruning and dynamic sampling strategy optimization phase The formula is: ; in, Indicates an accuracy-based bonus. Indicates a formatted reward. and These represent the corresponding coefficients; The objective function for maximizing the decoupled pruning and dynamic sampling strategy optimization phase is expressed as: ; in, Let represent the objective function for optimizing the decoupling pruning and dynamic sampling strategies. Represents the mathematical expectation. The sample set representing the prefix inference trajectory. This represents a dataset of step-level trajectory samples generated by the Monte Carlo tree search algorithm. This represents the set of candidate responses within the same sampled group. This represents the total number of candidate responses sampled within each group. Indicates the index of candidate responses within the group. This represents the old policy model before the update. Indicates the first j The total number of intermediate steps contained in each candidate response. Indicates the current policy model at the th... j The first trajectory t The step-level reward value output by each step. This represents a clipping function that restricts input values to a specified range. This indicates the lower bound of the clipping hyperparameter boundary. This indicates the upper bound of the clipping hyperparameter boundary. This represents the group advantage calculated based on the normalized mean and standard deviation of the reward: ; in, Indicates the first j The total reward score corresponding to each candidate reasoning trajectory. This represents the average reward score of all candidate inference trajectories in this group. This represents the standard deviation of the reward scores for all candidate inference trajectories in this group.
2. The production line fault prediction method based on complex signal reasoning using a large language model according to claim 1, characterized in that, The preprocessing in S1 includes: cleaning the original data and questions, removing irrelevant data; then standardizing and multimodal alignment; and finally, domain expert verification to obtain the inference and diagnostic step data.
3. The production line fault prediction method based on complex signal reasoning using a large language model according to claim 2, characterized in that, The Monte Carlo tree search algorithm incorporates textbook-guided solutions and self-reflection algorithms. The textbook's answer guidance specifically refers to providing expert-verified solutions. The reasoning is divided into blocks with a unified objective, as shown in the formula: The expert trunk branch in the search tree is initialized based on the reasoning block with a unified goal, and Monte Carlo tree search expansion is performed on the expert trunk branch. The self-reflection algorithm specifically involves: before extending to the next step in S2.2, first analyzing the historical partial trajectory... Generate reflection content The formula is: Then, the reflection content will be... Introduced as additional context, new candidate steps are generated, as shown in the formula: .
4. The production line fault prediction method based on complex signal reasoning using a large language model according to claim 3, characterized in that, In step S5, the original inference prediction trajectory is converted into a standardized sequence of steps using the step granularity alignment module. The step granularity alignment module uses a general large language model to perform non-evaluative parsing of the original response and inserts clear step boundary markers into the original response text based on the principles of single purpose, logical coherence, and clear transition.
5. The production line fault prediction method based on complex signal reasoning using a large language model according to claim 4, characterized in that, The quantitative evaluation in S7 includes: based on the step verification results of S6, calculating and outputting the process completion rate and the attribution of the first fatal error; the process completion rate is the proportion of correct steps before the occurrence of the first fatal error; the attribution of the first fatal error classifies the main failure sources into one of four dimensions: problem understanding, knowledge application, symbolic derivation, and numerical calculation.
6. A production line fault prediction system based on complex signal reasoning using a large language model, employing the production line fault prediction method based on complex signal reasoning using a large language model as described in any one of claims 1 to 5, characterized in that, include: Data acquisition and processing module: Collects raw data and problems from production line equipment, performs preprocessing, and constructs a benchmark test set; Inference trajectory module: Based on the benchmark test set, step-level inference trajectory sample data is generated using the Monte Carlo tree search algorithm to obtain the training dataset; Sample filtering module: performs multi-view data filtering on the sample data of the training dataset; Process reward model module: Trains the process reward model using filtered training samples; Fault Prediction and Inference Sequence Module: Utilizes a large language model to predict production line faults on samples in the benchmark test set, obtains the original inference prediction trajectory generated by the large language model, and converts it into a standardized sequence of steps; Verification module: Input the standardized step sequence into the trained process reward model to obtain the correctness label and evaluation reason for each step; Evaluation module: Based on the step-by-step verification results, quantitatively evaluate the comprehensive reasoning ability of the large language model in production line supervision and prediction; Fault prediction feedback module: It is used to optimize and adjust the large language model based on the process completion rate and the first fatal error attribution result output by the evaluation module, and use the optimized large language model to infer the real-time sensor signals of the target production line, and output the online production line fault prediction results and fault analysis report.
Citation Information
Patent Citations
Method for improving reasoning ability of large language model based on Monte Carlo tree search
CN119831051A
Measuring instrument quality intelligent analysis and data processing method based on AI large model
CN121145117A