Reliability analysis method and system for production manufacturing system, and electronic device

CN122779639APending Publication Date: 2026-09-18SHANGHAI JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611224153.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-12
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

部分方案能够根据历史故障数据或报警信息识别单一设备的异常状态,但通常难以刻画多工序、多设备之间的失效传播关系;另一类方案通过贝叶斯网络、故障树或专家规则进行可靠性建模,能够实现一定程度的风险推断,但往往停留在概率计算或风险评估层面,缺少与实时工况识别、知识库匹配和运维建议生成的联动

Benefits of technology

[0008]In this invention, firstly, a working condition identification agent identifies the current operating status of the production line, equipment participation, task cycle time, and process flow to determine whether the production line is currently in a state of full-process processing, single processing, assembly flow, or mixed operation. Secondly, a failure probability analysis module, combined with probabilistic reasoning methods such as two-layer Bayesian networks, models, injects, and updates the real-time failure probabilities of each subsystem, failure mode, and top event, forming a failure rate distribution under different working conditions. Then, a risk identification agent identifies high-risk nodes, potential failure paths, and reliability degradation trends based on the failure rate distribution, risk thresholds, and system correlations. Finally, an operation and maintenance decision-making agent calls upon a knowledge base, historical fault case library, and operation and maintenance rule base to generate targeted inspection, maintenance, adjustment, or early warning suggestions, thus forming an intelligent reliability analysis closed loop of "working condition identification—probability analysis—risk identification—path tracing—operation and maintenance suggestions."

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122779639A_ABST
    Figure CN122779639A_ABST
Patent Text Reader

Abstract

This invention provides a reliability analysis method, system, and electronic device for manufacturing systems, relating to the field of manufacturing analysis technology. The method includes: determining the current failure rate distribution information of the manufacturing system based on a failure event point table within a target time interval; parsing user-input task requirements to obtain user analysis tasks, including analysis objects, analysis objectives, and output requirements; identifying current high-risk nodes within the manufacturing system based on the user analysis tasks and real-time failure rate distribution information, where a high-risk node is any one of the following: failure modes within a subsystem and the top event of the subsystem; performing path tracing within the manufacturing system based on the high-risk nodes to obtain at least one failure propagation path, and generating a risk identification result, including the current high-risk node and the failure propagation path.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of manufacturing analysis technology, specifically to a reliability analysis method and system for manufacturing systems, and electronic equipment. Background Technology

[0002] With the development of intelligent manufacturing and mixed-line production, production lines are increasingly characterized by the collaborative operation of multiple devices, processes, and systems. In actual production, subsystems such as AGV (Automated Guided Vehicle) systems, CNC (Computer Numerical Control) systems, assembly units, gantry robots, loading and unloading channels, and MES (Manufacturing Execution System) are tightly coupled through process flow, logistics transfer, task scheduling, and data interaction. Any malfunction in any local device or functional module can be transmitted along the process chain and affect the subsequent production cycle, leading to task interruption, decreased efficiency, quality fluctuations, or even complete line shutdown.

[0003] In the field of existing production line reliability analysis, related research and practice mainly focus on single-equipment fault diagnosis, statistical failure rate analysis, failure mode and effects analysis (FMEA) / fault tree analysis (FTA), and rule-based equipment maintenance methods. Some solutions can identify the abnormal state of a single device based on historical fault data or alarm information, but they are usually difficult to characterize the failure propagation relationship between multiple processes and multiple devices. Another type of solution uses Bayesian networks, fault trees, or expert rules for reliability modeling, which can achieve a certain degree of risk inference, but often remains at the level of probability calculation or risk assessment, lacking linkage with real-time operating condition identification, knowledge base matching, and maintenance suggestion generation. Summary of the Invention

[0004] The purpose of this invention is to provide a reliability analysis method, system, and electronic equipment for manufacturing systems, which can identify high-risk nodes and failure propagation paths under multi-process and multi-equipment coupling based on real-time failure rate distribution, thereby improving the risk assessment capability at the production line level.

[0005] To achieve the above objectives, the present invention provides a reliability analysis method for a manufacturing system, comprising: Based on the failure event point table of the production and manufacturing system within the target time interval, the current failure rate distribution information of the production and manufacturing system is determined. The current failure rate distribution information includes: the real-time failure rate distribution of each subsystem within the production and manufacturing system and the real-time failure rate of the production and manufacturing system. The task requirement information input by the user is parsed to obtain the user analysis task, which includes: analysis object, analysis objective and output requirements; Based on the user analysis task and the real-time failure rate distribution information, the current high-risk nodes in the production and manufacturing system are determined. The high-risk nodes are any one of the following: failure modes and top events of the subsystem. Based on the high-risk nodes, path tracing is performed within the production system to obtain at least one failure propagation path, and a risk identification result is generated. The risk identification result includes: the current high-risk node and the failure propagation path. This invention also provides a reliability analysis system. This includes: a demand understanding agent, a risk identification agent, and a smart operation and maintenance agent; The demand understanding agent is used to: determine the current failure rate distribution information of the production and manufacturing system based on the failure event point table of the production and manufacturing system within the target time interval. The current failure rate distribution information includes: the real-time failure rate distribution of each subsystem within the production and manufacturing system and the real-time failure rate of the production and manufacturing system. The agent also parses the task demand information input by the user to obtain the user analysis task, which includes: the analysis object, the analysis target, and the output requirements. The risk identification agent is used to: determine the current high-risk node in the production and manufacturing system based on the user analysis task and the real-time failure rate distribution information, wherein the high-risk node is any one of the following: failure modes and top events in the subsystem; perform path tracing in the production and manufacturing system based on the high-risk node to obtain at least one failure propagation path, and generate a risk identification result, wherein the risk identification result includes: the current high-risk node and the failure propagation path; The intelligent operation and maintenance agent is used to retrieve data based on the user analysis task and the risk identification results, and generate an operation and maintenance solution for the production and manufacturing system.

[0006] The present invention also provides an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform a reliability analysis method for a manufacturing system as described above.

[0007] The present invention also provides a computer-readable storage medium, which is a non-volatile or non-transient storage medium, on which a computer program is stored, and the computer program is executed by a processor to perform the reliability analysis method of the manufacturing system described above.

[0008] In this invention, firstly, a working condition identification agent identifies the current operating status of the production line, equipment participation, task cycle time, and process flow to determine whether the production line is currently in a state of full-process processing, single processing, assembly flow, or mixed operation. Secondly, a failure probability analysis module, combined with probabilistic reasoning methods such as two-layer Bayesian networks, models, injects, and updates the real-time failure probabilities of each subsystem, failure mode, and top event, forming a failure rate distribution under different working conditions. Then, a risk identification agent identifies high-risk nodes, potential failure paths, and reliability degradation trends based on the failure rate distribution, risk thresholds, and system correlations. Finally, an operation and maintenance decision-making agent calls upon a knowledge base, historical fault case library, and operation and maintenance rule base to generate targeted inspection, maintenance, adjustment, or early warning suggestions, thus forming an intelligent reliability analysis closed loop of "working condition identification—probability analysis—risk identification—path tracing—operation and maintenance suggestions."

[0009] Compared to existing technologies, the advantages are as follows: First, it can expand the reliability analysis of production lines from single-point fault judgment to system-level risk analysis under multi-process and multi-equipment coupling conditions, improving the ability to identify cross-system failure propagation; second, it can use real-time failure probability and failure rate distribution to dynamically assess the reliability status of production lines, improving risk prediction and early warning capabilities; third, it can modularize the division of labor for operating condition identification, risk judgment, path tracing, and maintenance suggestion generation through a multi-agent collaborative mechanism, improving the automation and interpretability of the analysis process; fourth, it can combine knowledge bases and historical cases to transform failure probability analysis results into actionable maintenance and support suggestions, reducing reliance on human experience judgment and improving the efficiency of production line maintenance decisions.

[0010] This invention has the following significant features: First, it introduces a multi-agent collaborative mechanism, breaking down the production line reliability analysis task into sub-tasks such as operating condition identification, failure probability acquisition, risk identification, propagation path analysis, and maintenance suggestion generation, which are completed collaboratively by different agents according to their respective responsibilities. Second, it combines the probabilistic reasoning method of two-layer Bayesian networks to model the failure modes within each subsystem and the propagation relationships between systems, enabling the system to obtain real-time failure rate distributions under different operating conditions, providing a quantitative basis for risk identification and reliability assessment. Third, through a knowledge base, historical fault case database, and maintenance rule base invocation mechanism, the high-risk nodes, potential propagation paths, and reliability decline trends identified by the agents are transformed into targeted maintenance and assurance suggestions, thereby improving the automation and interpretability of production line risk analysis, prediction and early warning, and maintenance decision-making.

[0011] In one embodiment, the method further includes: performing a retrieval based on the user analysis task and the risk identification results to generate an operation and maintenance solution for the production and manufacturing system.

[0012] In one embodiment, based on the user analysis task and the real-time failure rate distribution information, determining the current high-risk nodes within the production and manufacturing system includes: Based on the current production line status of the manufacturing system, determine the target subsystems to participate in the operation; Based on the real-time failure rate distribution of each subsystem within the generative manufacturing system, N candidate high-risk nodes are identified, where N > 1. Based on the subsystem to which each candidate high-risk node belongs, the target subsystem corresponding to the current production line status, and the association strength between each candidate high-risk node and its neighboring nodes, the current high-risk node is selected from the N candidate high-risk nodes.

[0013] In one embodiment, path tracing is performed within the production system based on the high-risk node to obtain at least one failure propagation path, including: For each of the aforementioned high-risk nodes: Obtain the specified subsystem to which the high-risk node belongs; Based on the causal dependencies within the specified subsystem contained in the two-layer Bayesian network of the production and manufacturing system, the contribution of the root cause node and / or failure mode node within the specified subsystem to the high-risk node is obtained. Based on the propagation relationship between the top events of the subsystems contained in the two-layer Bayesian network, the downstream subsystems of the specified subsystem affected by the high-risk node and the production line-level consequences are determined. Based on the contribution of the root cause node and / or failure mode node in the specified subsystem to the high-risk node, the downstream subsystems affected by the high-risk node, and the production line-level consequences, the failure propagation path corresponding to the high-risk node is determined.

[0014] In one embodiment, based on the user analysis task and the risk identification results, a retrieval is performed to generate an operation and maintenance solution for the production and manufacturing system, including: Using the high-risk factors in the risk identification results as the main search objects, extract the search-related information corresponding to each main search object, and construct a search question for preset knowledge data based on the main search objects and the search-related information. Based on the retrieval questions corresponding to each of the high-risk factors and the user analysis task, the target knowledge content is obtained by retrieval within the preset knowledge data; Based on the target knowledge content corresponding to each of the high-risk factors and the real-time failure rate of each of the high-risk factors, tiered treatment suggestions are generated for each of the high-risk factors.

[0015] In one embodiment, the subsystems of the manufacturing system include: an automated guided vehicle system, a loading and unloading system, a computer numerical control machine tool system, a gantry robot system, an assembly system, and a manufacturing execution system; Based on the current production line status of the manufacturing system, identify the target subsystems to participate in the operation, including: If the current production line status of the manufacturing system is full-process processing, the target subsystems involved in the operation are identified as the automated guided vehicle system, loading and unloading system, computer numerical control machine tool system, gantry robot system, assembly system, and manufacturing execution system. If the current production line status of the manufacturing system is a single processing state, the target subsystems involved in the operation are identified as the gantry robot system, the loading and unloading system, and the computer numerical control machine tool system. If the current production line status of the manufacturing system is assembly flow, the target subsystems involved in the operation are identified as the automated guided vehicle system, the assembly system, and the manufacturing execution system.

[0016] In one embodiment, based on the failure rate distribution of each target subsystem, N candidate high-risk nodes are determined, including: For each node within each target subsystem, if the failure rate of the node meets a preset condition, then the node is determined as a candidate high-risk node; the preset condition is any one of the following: the failure rate of the node is greater than the corresponding failure rate threshold, or the failure rate of the node shows an abnormal trend in multiple consecutive time windows within the target time interval.

[0017] In one embodiment, after determining the failure propagation path corresponding to each of the high-risk nodes, the method further includes: If the high-risk node corresponds to multiple failure propagation paths, then: Based on the failure rate, conditional propagation probability, contribution of each node in each failure propagation path, and the participation of the subsystem to which the high-risk node belongs in the current production line state, the path risk weight of each failure propagation path is calculated. Based on the path risk weight of each failure propagation path, the failure propagation path corresponding to the high-risk node is selected from the multiple failure propagation paths. Attached Figure Description

[0018] Figure 1 This is a schematic diagram of the failure propagation structure of the AGV system within the production and manufacturing system in the first embodiment of the present invention; Figure 2 This is a schematic diagram of a two-layer Bayesian network of the manufacturing system in the first embodiment of the present invention; Figure 3This is a flowchart illustrating the reliability analysis method for a manufacturing system in the first embodiment of the present invention. Figure 4 This is a schematic diagram of the failure probability data of the AGV system within the production and manufacturing system in the first embodiment of the present invention; Figure 5 This is a schematic diagram of the subsystems participating in the operation of the production and manufacturing system under different production line states in the first embodiment of the present invention; Figure 6 This is a schematic diagram of the failure propagation path obtained based on high-risk node tracing in the first embodiment of the present invention; Figure 7 This is a flowchart illustrating the reliability analysis process in the first embodiment of the present invention; Figure 8 This is a schematic diagram of the real-time failure rate distribution of task interruption TI (failure mode) within the lower-level Bayesian network CPT of the AGV system in the first embodiment of the present invention. Figure 9 This is a schematic diagram of the user-inputted task requirements and corresponding response content in the first embodiment of the present invention. Specific Implementation

[0019] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings to provide a clearer understanding of the purpose, features, and advantages of the present invention. It should be understood that the embodiments shown in the drawings are not intended to limit the scope of the present invention, but are merely illustrative of the essential spirit of the technical solution of the present invention.

[0020] In the following description, certain specific details are set forth for the purpose of illustrating various disclosed embodiments in order to provide a thorough understanding of the various disclosed embodiments. However, those skilled in the art will recognize that embodiments may be practiced without one or more of these specific details. In other instances, well-known apparatuses, structures, and techniques associated with this application may not have been shown or described in detail to avoid unnecessarily obscuring the description of the embodiments.

[0021] Unless the context requires otherwise, throughout the specification and claims, the word “comprising” and its variations, such as “including” and “having”, shall be understood to have an open, inclusive meaning, that is, to be interpreted as “including, but not limited to”.

[0022] Throughout this specification, references to "an embodiment" or "an embodiment" indicate that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. Therefore, the appearance of "in an embodiment" or "an embodiment" in various places throughout the specification does not necessarily refer to the same embodiment. Furthermore, a particular feature, structure, or characteristic may be combined in any manner in one or more embodiments.

[0023] The singular forms “a” and “the” used in this specification and the appended claims include plural references unless otherwise expressly stated herein. It should be noted that the term “or” is generally used to include the meaning of “or / and” unless otherwise expressly stated herein.

[0024] In the following description, in order to clearly demonstrate the structure and working method of the present invention, a number of directional terms will be used. However, terms such as "front", "back", "left", "right", "outside", "inside", "outward", "inward", "up", and "down" should be understood as convenient terms and not as limiting terms.

[0025] Existing methods for reliability analysis of production lines still have several technical shortcomings in practical applications, such as: On the one hand, traditional equipment fault diagnosis and reliability analysis methods often focus on single devices or local systems, relying mainly on alarm records, manual inspections, and historical failure statistics for judgment. Since manufacturing lines typically contain multiple subsystems, and these subsystems are coupled through process flows, logistics, task scheduling, and data interaction, local anomalies in a single device can propagate along the process chain and gradually accumulate, ultimately leading to task interruptions, efficiency reductions, or even line-wide shutdowns. If only static analysis is performed on single-point faults, it is difficult to promptly identify cross-system failure propagation paths and production line-level reliability degradation trends.

[0026] On the other hand, while existing reliability modeling methods based on rules, fault trees, or Bayesian networks can infer some failure relationships, they typically focus more on calculating failure probabilities or judging risk levels. They lack comprehensive utilization of real-time operating conditions, failure rate distributions, historical cases, and operational knowledge, and struggle to further transform risk identification results into actionable operational support recommendations. Therefore, existing methods tend to remain at the level of "detecting anomalies" or "calculating probabilities," failing to form a complete intelligent analysis loop from real-time probability acquisition, risk identification, path tracing to maintenance decisions.

[0027] To address the aforementioned shortcomings, this invention proposes the following reliability analysis method for manufacturing systems.

[0028] The first embodiment of this invention relates to a reliability analysis method for a manufacturing system, applied to electronic devices, such as servers or common computer devices, such as desktop computers and laptops. The manufacturing system includes subsystems such as automated guided vehicle (AGV) systems, loading and unloading systems, computer numerical control (CNC) machine tool systems, gantry robot systems, assembly systems, and manufacturing execution systems.

[0029] When conducting reliability analysis, corresponding failure propagation structures are constructed in advance around typical subsystems within the production and manufacturing system, including: Automated Guided Vehicle (AGV) system, Computer Numerical Control (CNC) machine tool system, loading and unloading system, gantry robot system, assembly system, and Manufacturing Execution System (MES). For each subsystem, based on historical failure point tables, alarm records, rectification records, and process flow information, the failure nodes and top events (system failures) within that subsystem can be determined. Failure nodes include: failure modes and root causes within each failure mode. For example, taking an AGV system as an example, its top event is AGV system failure. Failure modes causing AGV system failure include: Task Interruption (TI) and Inability to Start UTS. Its root causes include Mechanical Failure (MF), Scheduling Communication Failure (SCF), Perception and Positioning Failure (PLF), and Docking Failure (DCF). Each failure mode is caused by its subordinate root causes, ultimately reflecting as AGV system failure. Specifically, the root causes under Task Interruption (TI) include Mechanical Failure (MF), Scheduling Communication Failure (SCF), Perception and Positioning Failure (PLF), and Docking Failure (DCF), while the root causes under Inability to Start UTS include Mechanical Failure (MF) and Scheduling Communication Failure (SCF). Figure 1 As shown; similarly, the failure node relationships among CNC systems, assembly systems, material handling systems, gantry robot systems, and manufacturing execution systems can be established in the same way, as illustrated in [reference]. Figure 2 .

[0030] Based on the above process, the failure propagation structure between subsystems of the manufacturing system and the relationship between failure nodes within subsystems have been constructed. Probabilistic modeling can then be performed; in this embodiment, a two-layer Bayesian network can be used as the basis for failure rate propagation and inference. Figure 2 As shown, the constructed two-layer Bayesian network of the manufacturing system includes a lower-layer Bayesian network and a top-layer Bayesian network, specifically: The lower-level Bayesian network is used to describe the relationship between failure nodes within a single subsystem, namely the probabilistic dependency between failure root causes, failure modes, and the top event of the subsystem. Specifically, it infers the failure mode based on the lower-level failure root causes, and then infers the failure probability of the top event of the subsystem based on the failure mode.

[0031] Top-level Bayesian networks are used to describe the propagation relationship between top events of different subsystems. They can analyze the impact of the failure of a certain subsystem on the reliability state of other subsystems and the entire system, and follow the correlation relationship of subsystems in actual production.

[0032] By using the two-layer Bayesian network structure of the production system, it is possible to simultaneously obtain the failure rate distribution within each subsystem and the failure propagation results at the production line level. This can not only help identify high-risk systems, but also further analyze the failure propagation path and potential impact.

[0033] like Figure 3 The diagram shown is a detailed flowchart of the reliability analysis method for the production and manufacturing system in this embodiment.

[0034] Step 101: Based on the failure event point table of the production and manufacturing system within the target time interval, determine the current failure rate distribution information of the production and manufacturing system. The current failure rate distribution information includes: the real-time failure rate distribution of each subsystem within the production and manufacturing system and the real-time failure rate of the production and manufacturing system.

[0035] Specifically, when using a two-layer Bayesian network in a production system, a failure rate is first injected into the two-layer Bayesian network to enable it to have initial inference capabilities. Specifically, a failure event point table within the target time interval is first obtained. This failure event point table includes information such as: failure occurrence time, subsystem to which it belongs, failure phenomenon, failure mode, corrective measures, downtime impact, and processing results.

[0036] First, the failure events are statistically analyzed and screened based on the failure event point table. Anomalies (failure events) that are highly sporadic, have incomplete records, or have weak engineering significance are removed, while frequent, high-impact failure events or failure events that are clearly related to production conditions are retained.

[0037] Subsequently, by combining the average annual working hours within the target time interval, the frequency of failure events (after screening), the frequency of failure modes, historical rectification records, and expert experience, the prior failure probability of each failure node (root cause node and failure mode) is calculated.

[0038] For example, if the target time interval is 3 years, calculate the total working time of the production and manufacturing system within 3 years and divide it by 3 to get the average annual working time. Then, based on the average annual working time and the length of a single observation time window, count the average number of observations per year. For example, if the average annual working time is 4000 hours and the length of a single observation time window is 4 hours, then the average number of observations per year is 1000. The frequency of failure events is the number of times various failure events occur within a year. Each failure event corresponds to a root cause, thus the number of occurrences of each root cause in each subsystem (annual average) can be calculated. If there are no historical rectification records or expert experience, the occurrence frequency of each type of root cause in each subsystem can be divided by the annual average number of observations to obtain the prior failure probability of each root cause. If there are historical rectification records and / or expert experience, the rectification weight corresponding to the historical rectification records and / or the expert weight corresponding to the expert experience can be obtained. For example, if there is a rectification record, a rectification weight less than 1 is selected; if there is no rectification for a long time, a rectification weight greater than 1 is selected; if there is no rectification for a short time, a rectification weight of 1 is selected. The expert weight is a preset fixed value. For example, a greater than 1 expert weight can be set for common root causes, and a less than 1 expert weight can be set for root causes with fast rectification cycles. The quotient obtained by dividing the occurrence frequency of each type of root cause in each subsystem by the annual average number of observations is then multiplied by the corresponding rectification weight and / or expert weight to obtain the prior failure probability of each root cause. Based on the above process, the prior failure probability of each type of failure root cause under each subsystem can be obtained; For each failure mode of each subsystem, a conditional probability table is obtained based on the occurrence frequency of the failure mode and the occurrence frequency of each type of failure root cause under it. This conditional probability table indicates the combination of all failure root causes under the failure mode and the conditional probability of the corresponding failure mode.

[0039] For each failure mode under each subsystem, the prior failure probability of each failure mode can be calculated based on the prior failure probabilities of various failure root causes under that failure mode and the conditional probability table of that failure mode.

[0040] Furthermore, for each subsystem, a conditional probability table can be obtained based on the occurrence frequency of the failure modes under that subsystem. This conditional probability table indicates the combination of all failure modes under the subsystem and the corresponding conditional probability of the subsystem.

[0041] If there is a lack of sufficient statistical samples for conditional probability relationships, failure modes and conditional probability tables for subsystems can be supplemented by typical distributions, expert rules, or historical cases of similar equipment.

[0042] For example, such as Figure 4As shown, it displays the failure probability data of the AGV system, including: a priori failure probability variation graphs for various failure root causes, conditional probability tables for each failure mode (Task Interruption TI and Unable to Start UTS), and a conditional probability table for the AGV system. "Normal" indicates that no failure has occurred.

[0043] Subsequently, based on the failure frequency of each subsystem within the target time interval, the real-time failure probability of each subsystem and the overall failure probability of the manufacturing system are calculated; furthermore, the conditional probability table of the manufacturing system can be obtained.

[0044] Based on the above, the failure probability data of all subsystems calculated above, along with the real-time failure probability data of each subsystem and the failure probability data of the manufacturing system, can be injected into a two-layer Bayesian network as raw data.

[0045] In some embodiments, for a manufacturing system, since different subsystems participate in the operation of mixed-line production at different times within the target time interval, that is, the production line status is different at different times, it is necessary to determine the target subsystems participating in reliability analysis based on the different production line statuses, thus obtaining a set of subsystems formed by the target subsystems corresponding to the production line status; for example, please refer to Figure 5 The details are as follows: If the current production line status of the manufacturing system is full-process processing, the target subsystems involved in the operation are identified as the automated guided vehicle system, loading and unloading system, computer numerical control machine tool system, gantry robot system, assembly system, and manufacturing execution system. If the current production line status of the manufacturing system is a single processing state, the target subsystems involved in the operation are identified as the gantry robot system, the loading and unloading system, and the computer numerical control machine tool system. If the current production line status of the manufacturing system is assembly flow, the target subsystems involved in the operation are identified as the automated guided vehicle system, the assembly system, and the manufacturing execution system.

[0046] For example, suppose the manufacturing system includes Each subsystem, and has Production line state, N>1, K>1; at any time The production line status is represented as For the first Production line status Determine the set of target subsystems corresponding to the production line status, and define... The first subsystem Each subsystem is in production line status The subsystem state participation quantity is ,exist At that time, the first Each subsystem belongs to the production line status. The corresponding set of target subsystems; in At that time, the first This subsystem is not part of the production line status. The corresponding set of target subsystems.

[0047] Among them, the subsystem state participation quantity is used to indicate whether the corresponding subsystem participates in the reliability analysis under the current production line state.

[0048] Furthermore, the runtime of each production line within the target time interval can be statistically analyzed. Let the first... The runtime of the production line in this state is The total working time of the production system within the target time interval is Then the first The weights for the occurrence of different production line states are as follows:

[0049] When the states of each production line are mutually exclusive and cover the entire working time of the manufacturing system, the sum of the weights of each production line state is 1. These production line state weights are used to calculate the contribution of each production line state to the overall failure risk of the manufacturing system within the target time interval, and are not used to replace the subsystem state participation.

[0050] For the Each subsystem, its real-time failure rate at any time t. Converted to a uniform observation time window Real-time failure probability within:

[0051] in, Indicates from t to ( A continuously changing time variable.

[0052] When the observation time window When the real-time failure probability within the time frame is approximately constant, the real-time failure probability is:

[0053] In the case where the failures of each subsystem are independent of each other, and the failure of any participating target subsystem will lead to the failure of the corresponding production line state, the first... Production line status The overall failure probability of the manufacturing system at any time t for:

[0054] No. Production line status The weighted contribution of any time t to the overall failure risk of the target time interval is:

[0055] The weighted overall failure probability of the manufacturing system at any time t within the target time interval, considering the frequency of occurrence of different production line states, is:

[0056] For the Production line status The next to participate in the operation For each target subsystem, the posterior failure probability of each failure root cause node, failure mode node, and subsystem top event node is calculated using the corresponding lower-level Bayesian network.

[0057] For any node (For the root cause node, failure mode node, or subsystem top event node), in the... Production line status State-related failure probability at any next time t for:

[0058] in, Represents a node The set of parent nodes, This represents a state combination of the parent node set. This represents the real-time operational data or observational evidence obtained at any given time t. (The rest of the text appears to be a fragment and requires further context for accurate translation.) The state-related failure probabilities of each node within a target subsystem determine the state of that target subsystem on the production line. The real-time failure probability distribution is shown below.

[0059] The failure rate injection described above can provide real-time failure probabilities for each subsystem, failure mode, and subsystem top event.

[0060] Step 102: Parse the task requirement information input by the user to obtain the user analysis task, which includes: analysis object, analysis target and output requirements.

[0061] Specifically, task requirement information is information input by the user through natural language or structured commands. This information clarifies the analytical objectives the user wants the system to achieve, such as identifying high-risk aspects of the current production line, analyzing the causes of anomalies in a subsystem, assessing potential downtime risks, or generating corresponding maintenance and support recommendations. Examples include: "Are there any high-risk aspects in the current production line?", "What are the main reasons for the increased AGV failure rate?", "Is there a risk of downtime under the current operating conditions?", and "What maintenance measures should be taken to address CNC task scheduling anomalies?"

[0062] The intelligent agent can receive task requirement information input by the user through the requirement understanding agent, parse the task objectives, analysis objects, time range, working condition type and expected output format, and combine it with the real-time failure rate distribution input during the failure rate injection process to form the analysis tasks that can be executed by the subsequent risk identification agent and operation and maintenance decision agent.

[0063] Based on steps 101 and 102 above, the following two types of content are achieved: First, the failure rate distribution generated by the two-layer Bayesian network, including the failure probability of each subsystem, the probability of major failure modes, and the overall failure probability of the production line, and able to identify specific operating conditions (corresponding to the production line status); Second, the user analysis task obtained by the demand understanding agent, including the analysis object (which may include: production system, subsystem, failure mode, root cause of failure, etc.), analysis objective, and output requirements. These results together serve as input to the multi-agent reliability analysis system, providing a foundation for subsequent high-risk identification, failure path tracing, data analysis, and the generation of operation and maintenance support recommendations.

[0064] Step 103: Based on the user analysis task and the real-time failure rate distribution information, determine the current high-risk nodes in the production and manufacturing system. The high-risk nodes are any one of the following: failure modes and top events in the subsystem.

[0065] Step 104: Based on the high-risk node, perform path tracing within the production and manufacturing system to obtain at least one failure propagation path and generate a risk identification result, which includes: the current high-risk node and the failure propagation path.

[0066] Steps 103 and 104 are executed by a risk identification intelligent agent. After completing the failure rate injection and user requirement understanding, it further identifies and judges the production line operation risks of the manufacturing system based on each subsystem, failure mode, and real-time failure rate distribution. The specific process includes: Step S11: Based on the current production line status of the manufacturing system, determine the target subsystems involved in the operation. The current production line status indicates the current working condition, i.e., the scope of risk analysis is determined based on the current working condition. Specifically: Since the subsystems involved in the operation differ under different production line statuses, the objects to be analyzed can be selected based on the working condition participation relationships of the type of working condition to be analyzed. For example, if the current production line status to be analyzed is a full-process processing state, the target subsystems involved in the operation include the automated guided vehicle system, loading and unloading system, CNC machine tool system, gantry robot system, assembly system, and manufacturing execution system; if the current production line status to be analyzed is a single processing state, the target subsystems involved in the operation include the gantry robot system, loading and unloading system, and CNC machine tool system; if the current production line status to be analyzed is an assembly flow state, the target subsystems involved in the operation include the automated guided vehicle system, assembly system, and manufacturing execution system.

[0067] The risk analysis scope (target subsystem) determined based on the current operating conditions can avoid interfering with the judgment of irrelevant subsystems, making the subsequent risk analysis results more consistent with the current production line status.

[0068] Step S12: Based on the real-time failure rate distribution of each subsystem within the manufacturing system, determine N candidate high-risk nodes, where N > 1. Specifically: for each node within each target subsystem, if the failure rate of the node meets a preset condition, then the node is determined as a candidate high-risk node. The preset condition is any one of the following: the failure rate of the node is greater than the corresponding failure rate threshold; or the failure rate of the node shows an abnormal trend within multiple consecutive time windows in the target time interval. When determining high-risk nodes, the failure rate and failure rate trend can be judged only for the failure mode nodes and the top event of the target subsystem.

[0069] The nodes of the subsystems mentioned here refer to the failure modes and top events of the subsystems. Dynamic failure rate thresholds can be set for the failure modes and top events of each subsystem. This enables the identification of high-risk nodes based on adaptive thresholds and failure rate change trends. For each node of a subsystem, the failure rate threshold can be dynamically determined by combining its historical failure rate distribution, current production line status, node importance, and node change trends within a continuous time window.

[0070] Therefore, for each node within the target subsystem, if the real-time failure rate of the node is greater than the corresponding failure rate threshold, or if the failure rate of the node gradually increases, suddenly increases, or fluctuates abnormally (representing an abnormal trend) over multiple consecutive time windows, the node is identified as a candidate high-risk node. Here, the time window is a preset duration within the target time interval, which can be divided into multiple time windows. When the failure rate of a node continuously increases over multiple consecutive time windows, the increase in failure rate between adjacent time windows exceeds the set change threshold, or abnormal fluctuations occur such as a sudden increase followed by a sudden decrease, or a sudden decrease followed by a sudden increase, it is determined that its trend is abnormal, and the node is also identified as a candidate high-risk node.

[0071] Furthermore, risk assessment intervals for different risk levels can be formed based on the historical average failure rate, failure rate fluctuation range, preset quantile threshold, or safety boundary value set by experts for each node. Thus, the risk level of each node can be determined based on the risk assessment interval in which its real-time failure rate falls.

[0072] Based on the above, candidate high-risk nodes within each target subsystem can be identified; Step S13: Based on the subsystem to which each candidate high-risk node belongs, the target subsystem corresponding to the current production line state, and the association strength between each candidate high-risk node and its neighboring nodes, the current high-risk node is selected from the N candidate high-risk nodes. Specifically: For all candidate high-risk nodes, the importance of each candidate high-risk node can be ranked based on its level (i.e., whether it is a top event of the subsystem or a failure mode), the state participation of the current production line state, and the association strength between the candidate high-risk node and its upstream and downstream nodes (neighboring nodes). For example, candidate high-risk nodes that are top events of the subsystem have higher priority, while candidate high-risk nodes that are failure modes have lower priority. For multiple candidate high-risk nodes that are also top events of the subsystem, and multiple candidate high-risk nodes that are also failure modes, the priority is determined according to the working condition participation of each candidate high-risk node. The quantities are sorted from largest to smallest. If multiple candidate high-risk nodes have the same priority and operating condition participation, they are further sorted from largest to smallest according to the correlation strength between these candidate high-risk nodes and their upstream and downstream nodes. The operating condition participation is the percentage of the candidate high-risk node's runtime in the total runtime of the current production line. The correlation strength between candidate high-risk nodes and their upstream and downstream nodes is determined as follows: for example, for a candidate high-risk node that is the top event of a subsystem, the more other target subsystems it is adjacent to, the stronger the correlation strength is determined; for a candidate high-risk node that is a failure mode, the higher its contribution to the top event of its target subsystem, the stronger its correlation strength is determined.

[0073] Step S14: Sort by importance from high to low, select the top M candidate high-risk nodes from all candidate high-risk nodes as the final high-risk nodes, 1≤M≤N.

[0074] Step S15: Based on the two-layer Bayesian network of each high-risk node in the production and manufacturing system, perform reverse tracing and forward analysis to obtain the failure propagation path of each high-risk node; specifically: For each of the aforementioned high-risk nodes: Obtain the specified subsystem to which the high-risk node belongs; Based on the causal dependencies within the specified subsystem contained in the two-layer Bayesian network of the production and manufacturing system, the contribution of the root cause node and / or failure mode node within the specified subsystem to the high-risk node is obtained. Based on the propagation relationship between the top events of the subsystems contained in the two-layer Bayesian network, the downstream subsystems of the specified subsystem affected by the high-risk node and the production line-level consequences are determined. Based on the contribution of the root cause node and / or failure mode node in the specified subsystem to the high-risk node, the downstream subsystems affected by the high-risk node, and the production line-level consequences, the failure propagation path corresponding to the high-risk node is determined.

[0075] In other words, the risk identification agent performs path tracing based on the two-layer Bayesian network of the production system, including forward tracing and reverse tracing. If a high-risk node is a high-risk top event or an abnormal failure mode, then the agent starts from the high-risk node and traces backward along the lower-layer Bayesian network of the subsystem to which the high-risk node belongs, in order to achieve subsystem tracing and determine the contribution of the root cause node and failure mode node within the subsystem to which the high-risk node belongs to the failure of the subsystem. Then, forward analysis is performed along the top-level propagation relationship between the process flow and the current production line status to determine the downstream target subsystems that the target subsystem corresponding to the high-risk node may affect, as well as the production line-level consequences. The production line-level consequences indicate the contribution of the target subsystem corresponding to the high-risk node and the downstream target subsystems that may be affected to the production line failure in the current production line status.

[0076] For example, please refer to Figure 6 It shows the failure propagation path obtained by tracing the source of a high-risk node, including the failure probability P of each node. Furthermore, it also shows the risk level grade of the node, which can be one of three risk levels: 1, 2, and 3.

[0077] From the above, we can obtain the nodes within the subsystem associated with each high-risk node in its respective target subsystem, as well as the downstream target subsystem associated with it in the current production line state. For each high-risk node, by connecting the high-risk node, its associated nodes within the subsystem, and the downstream target subsystem, we can obtain one or more corresponding failure propagation paths. Furthermore, if the high-risk node corresponds to multiple failure propagation paths, then: Based on the failure rate, conditional propagation probability, contribution of each node in each failure propagation path, and the participation of the subsystem to which the high-risk node belongs in the current production line state, the path risk weight of each failure propagation path is calculated.

[0078] Specifically, let the high-risk node correspond to the first The failure propagation path is , including 1 node Indicates the first The first failure propagation path The node will be at time t. The real-time failure rate of each node within the failure propagation path is converted into a unified observation time window. Failure probability within and the Failure propagation path number Conditional propagation probability between adjacent nodes within a failure propagation path and the contribution of each node .

[0079] For the For each failure propagation path, first calculate the weighted average of the contribution of each node's failure probability in the path. Then multiply this average by the product of the conditional propagation probabilities between adjacent nodes on the path and the path condition participation of the path under the current production line state to obtain the failure propagation path at time t. Unnormalized path risk weights of failure propagation paths :

[0080] in, Indicates the first The failure propagation path is in the first Production line status The path condition participation amount under the current production line status. The path condition participation amount can be determined based on the condition participation amount of each subsystem involved in the failure propagation path under the current production line status, for example, taking the minimum value among the condition participation amounts of each involved subsystem.

[0081] Furthermore, the unnormalized path risk weights of each failure propagation path are normalized:

[0082] in, This indicates the number of candidate failure propagation paths corresponding to high-risk nodes. Indicates the normalized i-th The path risk weights of the failure propagation paths are determined. One or more candidate failure propagation paths are selected from R candidate failure propagation paths in descending order of their normalized path risk weights, and these selected paths are designated as the primary failure propagation paths corresponding to the high-risk nodes.

[0083] Based on the path risk weight of each failure propagation path, the failure propagation path corresponding to the high-risk node is selected from the multiple failure propagation paths. For example, one or more failure propagation paths are selected as the final failure propagation path according to the path risk weight from largest to smallest, which is the main failure path.

[0084] This forms an explainable risk chain, i.e., the failure propagation path, from high-risk nodes to potential root causes, and then to the affected objects and production line consequences.

[0085] Step S16 outputs the identified high-risk nodes and corresponding failure propagation paths as risk identification results. This also includes data analysis and interpretation of the risk identification results. The data analysis includes the current high-risk subsystems, main failure factors, failure rate trends, failure propagation paths, impact scope, and risk level. This result serves as the basis for subsequent knowledge base matching and maintenance support recommendations.

[0086] Based on the above process, real-time failure rate data can be transformed into understandable risk assessment results, which can not only identify high-risk systems, but also explain the source of risk and identify the root cause of failure.

[0087] Step 105: Based on the user analysis task and the risk identification results, a search is performed to generate an operation and maintenance solution for the production and manufacturing system.

[0088] This step is executed by the intelligent operation and maintenance agent, which receives the risk identification results output by the risk identification agent, including: the current high-risk subsystem, main failure factors, failure rate trend, failure propagation path, scope of impact, and risk level. It then combines these results with the production line knowledge base, historical event case library, operation and maintenance rule base, and factor graph contained in the preset knowledge data to perform retrieval enhancement generation (RAG) and generate an operation and maintenance handling plan for the production and manufacturing system. Specifically, it uses the high-risk factors in the risk identification results as the main retrieval objects, extracts the retrieval-related information corresponding to each main retrieval object, and constructs a retrieval question based on the main retrieval objects and the retrieval-related information, targeting the preset knowledge data. Based on the retrieval questions corresponding to each of the high-risk factors and the user analysis task, the target knowledge content is obtained by retrieval within the preset knowledge data; Based on the target knowledge content corresponding to each of the high-risk factors and the real-time failure rate of each of the high-risk factors, tiered treatment suggestions are generated for each of the high-risk factors.

[0089] More details: The intelligent operation and maintenance agent determines the search targets based on the risk identification results. For subsystem top events or failure modes that have been identified as high-risk, it uses these high-risk nodes as the main search objects, extracting the corresponding equipment names, failure types, abnormal manifestations, upstream and downstream related nodes, and possible consequences to construct search-related information. It then uses the high-risk nodes and search-related information to construct a search question oriented towards the knowledge base. For example, when the system identifies a high probability of AGV sensing and positioning failure, which may further lead to docking failure and task interruption, the intelligent operation and maintenance agent will perform a knowledge search around keywords such as "AGV sensing and positioning anomaly," "docking failure," "task interruption," "positioning sensor check," and "path calibration," looking up corresponding handling methods, past maintenance cases, and standard operating procedures from the knowledge base.

[0090] The intelligent operation and maintenance agent utilizes RAG to enhance its knowledge retrieval capabilities. The production line knowledge base can store equipment manuals, historical fault cases, maintenance records, alarm handling rules, process flow documents, spare parts information, operation and maintenance manuals, and expert experience. Based on current high-risk nodes and user analysis tasks, the intelligent operation and maintenance agent retrieves relevant content from the production line knowledge base and generates structured operation and maintenance suggestions by combining it with a large model. To avoid the problem of ordinary text retrieval relying solely on keyword matching, a graph-based relationship based on failure factors can be constructed, linking "root cause—failure mode—subsystem—operating condition participation—impact consequences—remedial measures." This factor graph can assist in determining failure paths in the aforementioned risk tracing and can also be used to expand the search scope in RAG retrieval.

[0091] Then, the intelligent operation and maintenance agent generates tiered handling suggestions based on the risk status of different high-risk nodes. For high-risk nodes with already high failure rates that may directly affect current production tasks, immediate handling suggestions are output, including shutdown inspection, manual review, temporary bypass, reduced load operation, task rescheduling, inspection of key components or replacement of spare parts, etc., and the priority targets and expected impacts are clearly defined. For subsystems that have not yet reached a serious risk level but are affected by upstream failures, preventative suggestions are output, including increasing monitoring frequency, checking the status of upstream and downstream interfaces, reviewing communication links, confirming process cycle time, reserving maintenance windows or conducting partial joint debugging tests to prevent the risk from continuing to spread. For nodes whose current failure rate has not increased significantly but are in critical process links, preventative maintenance suggestions are output, including regular inspections, sensor calibration, software log checks, lubrication maintenance, spare parts status confirmation and threshold tracking, to ensure that they do not become new risk amplification points in subsequent production processes.

[0092] Finally, the operation and maintenance recommendations, cited knowledge sources, failure propagation paths, and handling results are recorded to form a traceable operation and maintenance closed loop. On the one hand, users can perform inspection, maintenance, or adjustment operations based on the output recommendations; on the other hand, the actual handling results can be fed back to the knowledge base and case library for rapid matching and optimization of recommendations for similar risks in the future.

[0093] Therefore, the intelligent operation and maintenance agent can transform the "failure rate distribution and risk identification results" into an "explainable, traceable, and executable" operation and maintenance support solution, enabling the system to have complete intelligent operation and maintenance capabilities from risk discovery to problem location, from path tracing to maintenance decision-making.

[0094] Please refer to Figure 7 The execution flow of the above process can be summarized as follows: First, the user inputs their reliability analysis requirements and receives the real-time failure rate distribution. The requirement understanding agent parses the user's requirements, extracts the analysis object, analysis target, operating condition range, and expected output content, and generates a structured task JSON.

[0095] Secondly, the risk identification agent reads the task JSON and failure rate distribution JSON, and based on the current operating conditions, node thresholds, failure rate trends and failure propagation networks, identifies high-risk subsystems, main failure modes, potential root causes and propagation paths, and outputs the risk identification JSON.

[0096] Then, the intelligent operation and maintenance agent reads the risk identification results, calls the RAG knowledge base, historical case library, operation and maintenance rule library and factor graph, retrieves relevant processing methods and similar cases, and generates JSON operation and maintenance suggestions containing processing priorities, inspection items, preventive measures and subsequent observation indicators.

[0097] Finally, the system integrates the above results into a user-oriented reliability analysis report, outputting the current risk status, problem location, potential impact, and operation and maintenance support recommendations.

[0098] This embodiment uses structured JSON format as the input and output interface between intelligent agents, and standardizes and encapsulates user analysis tasks, failure rate distribution, risk identification results and operation and maintenance suggestions, so that each intelligent agent can receive tasks, output results and make subsequent calls according to unified fields; so that the information transmission between different intelligent agents has uniformity, parsability and traceability.

[0099] Through the aforementioned JSON structure, different intelligent agents can complete task transfer and result retrieval in a unified format, avoiding information omissions or inconsistencies in fields that occur in natural language interactions. Simultaneously, the system can retain each round of input, analysis results, knowledge base references, and operational suggestions, forming a traceable reliability analysis record. Therefore, this invention can link failure rate injection, risk identification, path tracing, and intelligent operational suggestions into a complete closed loop, realizing an intelligent reliability analysis method and system for production lines.

[0100] For example, using a one-year failure event log of an automotive parts manufacturing line as the basic data, the log includes information such as failure occurrence time, subsystem, failure phenomenon, failure mode, corrective measures, downtime impact, and handling results. The system first cleans and categorizes this log, filtering out events with high frequency or incomplete records. Then, the system calculates the frequency of different failure modes, combines this with average annual working hours and operational participation relationships, calculates the base failure rate for each subsystem, failure mode, and top event, and injects this into the failure propagation network to form a failure rate distribution that can be invoked by the intelligent agent, such as... Figure 8 The figure shows the real-time failure rate distribution of task interruption TI (failure mode) within the lower-level Bayesian network CPT of the AGV system. TI_t=Failure indicates that a task interruption failure has occurred, and TI_t=Normal indicates that a task interruption failure has not occurred. Various failure root causes are also shown, with Normal indicating no failure and Failure indicating a failure.

[0101] After the failure rate injection is completed, the user can submit specific analysis requirements. The multi-agent system can then complete the analysis task. The system outputs a reliability analysis report in a unified result format, including current operating conditions, major high-risk subsystems, key failure modes, possible propagation paths, potential impact consequences, handling priorities, and operational support recommendations.

[0102] For example, please refer to Figure 9After the failure rate injection is completed, the user inputs the task requirement information as "Please identify the current production line working status, analyze whether there is a potential downtime risk based on the real-time failure rate distribution, and provide operation and maintenance support suggestions."; and the response content given by the multi-agent system is shown.

[0103] This example demonstrates that this invention does not merely statistically analyze historical failure events, nor does it simply output failure probabilities. Instead, it integrates the failure rate baseline formed by a one-year failure point table, real-time operating conditions, user task requirements, multi-agent risk identification, and knowledge base-enhanced maintenance recommendations to form a complete reliability analysis closed loop from "data input—failure rate calculation—risk identification—path tracing—maintenance decision-making." Specifically, it can identify high-risk subsystems and major failure modes under different operating conditions, pinpoint key issues that may lead to task interruptions, abnormal cycle times, or reduced capacity, and generate corresponding inspection, maintenance, and early warning recommendations based on the knowledge base and historical cases.

[0104] This embodiment first establishes a failure rate distribution based on a failure event point table, real-time operational data, and operational condition participation relationships. Second, it determines the current analysis object, operational condition range, and output target according to the user's analysis requirements. Third, the risk identification agent judges high-risk subsystems, major failure modes, and propagation paths, and identifies existing risks, related risks, and potential risks. Finally, the intelligent operation and maintenance agent generates tiered handling suggestions by combining a knowledge base and past cases, and uses structured JSON to realize task transfer, result invocation, and process traceability between agents.

[0105] Through the above mechanism, a shift from "single-point fault diagnosis" to "multi-agent collaborative reliability analysis and intelligent operation and maintenance" has been achieved. Compared with existing technologies, its advantages are mainly reflected in three aspects: First, it can identify high-risk nodes and propagation paths under multi-process and multi-equipment coupling based on real-time failure rate distribution, improving the production line-level risk assessment capability; Second, it can combine RAG knowledge base and factor graph retrieval processing methods, historical cases and operation and maintenance rules to improve the pertinence and interpretability of operation and maintenance suggestions; Third, it can form a closed-loop process from data input, risk identification to operation and maintenance decision-making through a multi-agent serial mechanism, reducing reliance on human experience and improving the automation level, response efficiency and engineering application value of production line reliability analysis.

[0106] The second embodiment of the present invention relates to a reliability analysis system, which is a multi-agent system, comprising: a demand understanding agent, a risk identification agent, and a smart operation and maintenance agent; The demand understanding agent is used to: determine the current failure rate distribution information of the production and manufacturing system based on the failure event point table of the production and manufacturing system within the target time interval. The current failure rate distribution information includes: the real-time failure rate distribution of each subsystem within the production and manufacturing system and the real-time failure rate of the production and manufacturing system. The agent also parses the task demand information input by the user to obtain the user analysis task, which includes: the analysis object, the analysis target, and the output requirements. The risk identification agent is used to: determine the current high-risk node in the production and manufacturing system based on the user analysis task and the real-time failure rate distribution information, wherein the high-risk node is any one of the following: failure modes and top events in the subsystem; perform path tracing in the production and manufacturing system based on the high-risk node to obtain at least one failure propagation path, and generate a risk identification result, wherein the risk identification result includes: the current high-risk node and the failure propagation path; The intelligent operation and maintenance agent is used to retrieve data based on the user analysis task and the risk identification results, and generate an operation and maintenance solution for the production and manufacturing system.

[0107] In this embodiment, a multi-agent collaborative architecture is used to process production line reliability analysis tasks in stages. The requirement understanding agent clarifies the analysis object, current operating condition, and output target based on the user-input analysis requirements and failure rate distribution. The risk identification agent, combining failure rate thresholds, trends, and failure propagation networks, identifies high-risk subsystems, major failure modes, and potential failure paths. The intelligent operations and maintenance agent further utilizes the RAG knowledge base, historical case library, operations and maintenance rule base, and factor graph to retrieve similar problems and solutions, generating hierarchical operations and maintenance support recommendations for already high-risk nodes, affected nodes, and potentially risky nodes. This collaborative approach achieves a closed-loop transformation from "failure rate calculation" to "risk identification" and then to "operations and maintenance decision-making."

[0108] Based on failure rate injection and user-inputted task requirement understanding, the system transforms real-time failure rate distribution and user needs into structured analysis tasks. It employs methods for high-risk node screening, threshold judgment, trend analysis, and failure path tracing based on risk identification agents. A smart operation and maintenance suggestion generation method based on the RAG knowledge base and factor graph is used to retrieve processing methods, historical cases, and operation and maintenance rules based on risk identification results. A multi-agent concatenation mechanism based on JSON structured input and output is used to achieve information transmission, result retrieval, and process traceability between different agents. This system forms an integrated architecture encompassing failure rate injection, task understanding, risk identification, path tracing, knowledge retrieval, and operation and maintenance suggestion output. Functionally, it integrates failure probability, production line operating conditions, knowledge base resources, and agent decision-making capabilities, providing an implementable, scalable, and traceable technical solution for production line reliability analysis and smart operation and maintenance.

[0109] Since the first embodiment corresponds to this embodiment, this embodiment can be implemented in conjunction with the first embodiment. The relevant technical details mentioned in the first embodiment remain valid in this embodiment, and the technical effects achievable in the first embodiment can also be achieved in this embodiment. To reduce repetition, they will not be repeated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied to the first embodiment.

[0110] The third embodiment of the present invention relates to an electronic device, such as a server, or a common computer device, such as a desktop computer or laptop. The reliability analysis system of the second embodiment can be deployed in this electronic device to implement the reliability analysis method of the manufacturing system in the first embodiment.

[0111] The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the reliability analysis method of the manufacturing system in the first embodiment.

[0112] It should be noted that the memory may include random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage device. Similarly, the processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0113] Since the first embodiment corresponds to this embodiment, this embodiment can be implemented in conjunction with the first embodiment. The relevant technical details mentioned in the first embodiment remain valid in this embodiment, and the technical effects achievable in the first embodiment can also be achieved in this embodiment. To reduce repetition, they will not be repeated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied to the first embodiment.

[0114] The fourth embodiment of the present invention relates to a computer-readable storage medium, which is a non-volatile storage medium or a non-transient storage medium, and stores a computer program thereon. The computer program is executed by a processor to perform the reliability analysis method of the manufacturing system in the first embodiment.

[0115] The preferred embodiments of the present invention have been described in detail above, but it should be understood that, if necessary, aspects of the embodiments can be modified to utilize aspects, features, and concepts from various patents, applications, and publications to provide other embodiments.

[0116] In light of the detailed description above, these and other changes can be made to the embodiments. Generally, the terminology used in the claims should not be considered limited to the specific embodiments disclosed in the specification and claims, but should be understood to include all possible embodiments together with the full scope of equivalents enjoyed by these claims.

Claims

1. A reliability analysis method for a manufacturing system, characterized in that, include: Based on the failure event point table of the production and manufacturing system within the target time interval, the current failure rate distribution information of the production and manufacturing system is determined. The current failure rate distribution information includes: the real-time failure rate distribution of each subsystem within the production and manufacturing system and the real-time failure rate of the production and manufacturing system. The task requirement information input by the user is parsed to obtain the user analysis task, which includes: analysis object, analysis objective and output requirements; Based on the user analysis task and the real-time failure rate distribution information, the current high-risk nodes in the production and manufacturing system are determined. The high-risk nodes are any one of the following: failure modes and top events of the subsystem. Based on the high-risk node, path tracing is performed within the production and manufacturing system to obtain at least one failure propagation path, and a risk identification result is generated. The risk identification result includes: the current high-risk node and the failure propagation path.

2. The reliability analysis method for a manufacturing system according to claim 1, characterized in that, The method further includes: performing a retrieval based on the user analysis task and the risk identification results to generate an operation and maintenance solution for the production and manufacturing system.

3. The reliability analysis method for a manufacturing system according to claim 1, characterized in that, Based on the user analysis task and the real-time failure rate distribution information, the current high-risk nodes within the production and manufacturing system are identified, including: Based on the current production line status of the manufacturing system, determine the target subsystems to participate in the operation; Based on the real-time failure rate distribution of each subsystem within the generative manufacturing system, N candidate high-risk nodes are identified, where N > 1. Based on the subsystem to which each candidate high-risk node belongs, the target subsystem corresponding to the current production line status, and the association strength between each candidate high-risk node and its neighboring nodes, the current high-risk node is selected from the N candidate high-risk nodes.

4. The reliability analysis method for a manufacturing system according to claim 1, characterized in that, Based on the high-risk nodes, path tracing is performed within the production system to obtain at least one failure propagation path, including: For each of the aforementioned high-risk nodes: Obtain the specified subsystem to which the high-risk node belongs; Based on the causal dependencies within the specified subsystem contained in the two-layer Bayesian network of the production and manufacturing system, the contribution of the root cause node and / or failure mode node within the specified subsystem to the high-risk node is obtained. Based on the propagation relationship between the top events of the subsystems contained in the two-layer Bayesian network, the downstream subsystems of the specified subsystem affected by the high-risk node and the production line-level consequences are determined. Based on the contribution of the root cause node and / or failure mode node in the specified subsystem to the high-risk node, the downstream subsystems affected by the high-risk node, and the production line-level consequences, the failure propagation path corresponding to the high-risk node is determined.

5. The reliability analysis method for a manufacturing system according to claim 2, characterized in that, Based on the user analysis task and the risk identification results, a retrieval is performed to generate an operation and maintenance solution for the production and manufacturing system, including: Using the high-risk factors in the risk identification results as the main search objects, extract the search-related information corresponding to each main search object, and construct a search question for preset knowledge data based on the main search objects and the search-related information. Based on the retrieval questions corresponding to each of the high-risk factors and the user analysis task, the target knowledge content is obtained by retrieval within the preset knowledge data; Based on the target knowledge content corresponding to each of the high-risk factors and the real-time failure rate of each of the high-risk factors, tiered treatment suggestions are generated for each of the high-risk factors.

6. The reliability analysis method for a manufacturing system according to claim 3, characterized in that, The subsystems of the manufacturing system include: automated guided vehicle system, loading and unloading system, computer numerical control machine tool system, gantry robot system, assembly system, and manufacturing execution system; Based on the current production line status of the manufacturing system, identify the target subsystems to participate in the operation, including: If the current production line status of the manufacturing system is full-process processing, the target subsystems involved in the operation are identified as the automated guided vehicle system, loading and unloading system, computer numerical control machine tool system, gantry robot system, assembly system, and manufacturing execution system. If the current production line status of the manufacturing system is a single processing state, the target subsystems involved in the operation are identified as the gantry robot system, the loading and unloading system, and the computer numerical control machine tool system. If the current production line status of the manufacturing system is assembly flow, the target subsystems involved in the operation are identified as the automated guided vehicle system, the assembly system, and the manufacturing execution system.

7. The reliability analysis method for a manufacturing system according to claim 3, characterized in that, Based on the failure rate distribution of each target subsystem, N candidate high-risk nodes are identified, including: For each node within each target subsystem, if the failure rate of the node meets a preset condition, then the node is determined as a candidate high-risk node; the preset condition is any one of the following: the failure rate of the node is greater than the corresponding failure rate threshold, or the failure rate of the node shows an abnormal trend in multiple consecutive time windows within the target time interval.

8. The reliability analysis method for a manufacturing system according to claim 4, characterized in that, After determining the failure propagation path corresponding to each of the high-risk nodes, the method further includes: If the high-risk node corresponds to multiple failure propagation paths, then: Based on the failure rate, conditional propagation probability, contribution of each node in each failure propagation path, and the participation of the subsystem to which the high-risk node belongs in the current production line state, the path risk weight of each failure propagation path is calculated. Based on the path risk weight of each failure propagation path, the failure propagation path corresponding to the high-risk node is selected from the multiple failure propagation paths.

9. A reliability analysis system, characterized in that, include: Demand understanding intelligent agent, risk identification intelligent agent, intelligent operation and maintenance intelligent agent; The demand understanding agent is used to: determine the current failure rate distribution information of the production and manufacturing system based on the failure event point table of the production and manufacturing system within the target time interval. The current failure rate distribution information includes: the real-time failure rate distribution of each subsystem within the production and manufacturing system and the real-time failure rate of the production and manufacturing system. The agent also parses the task demand information input by the user to obtain the user analysis task, which includes: the analysis object, the analysis target, and the output requirements. The risk identification agent is used to: determine the current high-risk node in the production and manufacturing system based on the user analysis task and the real-time failure rate distribution information, wherein the high-risk node is any one of the following: failure modes and top events in the subsystem; perform path tracing in the production and manufacturing system based on the high-risk node to obtain at least one failure propagation path, and generate a risk identification result, wherein the risk identification result includes: the current high-risk node and the failure propagation path; The intelligent operation and maintenance agent is used to retrieve data based on the user analysis task and the risk identification results, and generate operation and maintenance solutions for the production and manufacturing system.

10. An electronic device, characterized in that, include: At least one processor; And a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the reliability analysis method for the manufacturing system as described in any one of claims 1 to 8.

11. A computer-readable storage medium, said computer-readable storage medium being a non-volatile storage medium or a non-transient storage medium, having stored thereon a computer program, characterized in that, The computer program is executed by the processor to perform the reliability analysis method for the manufacturing system as described in any one of claims 1 to 8.