Failure transfer method, equipment, medium and product of equipment automation manufacturing system

By performing change event detection and heartbeat packet detection on the equipment automation manufacturing system, and combining the situation and quantity of the downward airport to formulate a failover strategy, the problem of single failover strategy in the existing technology is solved, and finer granularity control and higher adaptability and robustness are achieved.

CN119415325BActive Publication Date: 2025-05-13上海朋熙半导体股份有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510018421.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-13
Estimated Expiration
2045-01-07

AI Technical Summary

Technical Problem

The failover strategy in the prior art is too single to cope with the diverse and complex failure modes in the automated manufacturing systems of equipment in the semiconductor industry, resulting in the possibility of losing key events during the downtime and affecting business processes.

Method used

By detecting the change event of the child nodes under the current server persistent node in the equipment automation manufacturing system, determine whether the automation manufacturing program is down; detecting the survival status of the automation manufacturing equipment and daemons based on the heartbeat package to determine whether it is down; formulating different failover strategies based on the airport situation and the number of downtime will be achieved to achieve automatic transfer and restore to the pre-downtime status.

Benefits of technology

It realizes finer granular control, improves the adaptability and robustness of the equipment automation manufacturing system, ensures that it can be automatically transferred and restored to the previous state in the event of downtime, and reduces the impact of business processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119415325B_ABST
    Figure CN119415325B_ABST
Patent Text Reader

Abstract

The embodiments of the present application relate to the field of computer technology, and disclose a method, device, medium, and product for the failover of an equipment automated manufacturing system. The method includes detecting change events of subnodes under the persistent node of the current server in the equipment automated manufacturing system to determine whether the automated manufacturing program is down; detecting the survival status of the automated manufacturing equipment and the daemon program based on the heartbeat packet to determine whether the automated manufacturing equipment or the daemon program is down; formulating different failover strategies based on the downtime scenario and the number of downtimes to automatically transfer and restore the equipment automated manufacturing system to the state before the downtime. The solution of the present invention solves the problem of downtime of a single automated manufacturing device, multiple automated manufacturing devices, and servers in the equipment automated manufacturing system of the semiconductor industry by customizing failover strategies that meet different downtime scenarios, achieving more fine-grained control, and improving the adaptability and robustness of the equipment automated manufacturing system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a fault transfer method, device, medium and product of an equipment automation manufacturing system. Background Art

[0002] Failover, or Fail Over, is a common fault-tolerance mechanism that aims to automatically switch to a backup system or component to continue providing services when a primary system or component fails. Failover can be implemented in a variety of ways and can be customized based on different technical backgrounds and application scenarios.

[0003] However, the inventors found that there are at least the following technical problems in the related art:

[0004] Different business scenarios have different requirements for failover, and the current failover strategy may be too simple to meet the needs of various business scenarios. With the development of technology, new hardware and software architectures may require new failover mechanisms to adapt to their characteristics, and the failure modes that the system may encounter are becoming more and more diverse and complex; for example, in real-time data output tasks, such as equipment downtime during automated manufacturing operations, key events may be lost during failover, affecting business processes. Summary of the invention

[0005] One purpose of the present application is to provide a system that at least solves the problem of downtime of a single automated manufacturing device, multiple automated manufacturing devices, and a server in an automated manufacturing system for equipment in the semiconductor industry, achieves finer-grained control, and improves the adaptability and robustness of the automated manufacturing system for equipment.

[0006] To achieve the above objectives, some embodiments of the present application provide the following aspects:

[0007] In a first aspect, some embodiments of the present application further provide a method for failure transfer of an equipment automation manufacturing system, characterized in that the method comprises:

[0008] Perform change event detection on the sub-nodes under the current server persistent node in the equipment automated manufacturing system to determine whether the automated manufacturing program is down;

[0009] Detect the survival status of the automated manufacturing equipment and the daemon program based on the heartbeat packet to determine whether the automated manufacturing equipment or the daemon program is down;

[0010] Different failover strategies are formulated according to the downtime scenarios and the number of downtimes to perform automatic transfer and restore the equipment automated manufacturing system to the state before the downtime; the downtime scenarios include at least one of the automated manufacturing program downtime, the automated manufacturing equipment downtime scenario and the daemon downtime scenario.

[0011] In a second aspect, some embodiments of the present application further provide an electronic device, comprising: one or more processors; and a memory storing computer program instructions, wherein the computer program instructions, when executed, cause the processor to perform the steps of the method described above.

[0012] In a third aspect, some embodiments of the present application further provide a computer-readable medium having computer program instructions stored thereon, wherein the computer program instructions can be executed by a processor to implement the method as described above.

[0013] In a fourth aspect, some embodiments of the present application further provide a computer program product, comprising a computer program / instruction, which implements the steps of the method described above when executed by a processor.

[0014] Compared with the related art, the solution provided in the embodiment of the present application determines whether the automated manufacturing program is down by detecting the change event of the subnode under the current server persistent node in the equipment automated manufacturing system; detects the survival status of the automated manufacturing equipment and the daemon according to the heartbeat packet to determine whether the automated manufacturing equipment or the daemon is down; formulates different failover strategies according to the downtime scenario and the number of downtimes to automatically transfer and restore the equipment automated manufacturing system to the state before the downtime. The technical solution of the embodiment of the present invention solves the problem of downtime of a single automated manufacturing device, multiple automated manufacturing devices and servers in the equipment automated manufacturing system of the semiconductor industry by customizing the failover strategy that meets different downtime scenarios, realizes more fine-grained control, and improves the adaptability and robustness of the equipment automated manufacturing system. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] One or more embodiments are exemplarily described by pictures in the corresponding drawings, and these exemplified descriptions do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings represent similar elements, and unless otherwise stated, the figures in the drawings do not constitute proportional limitations.

[0016] Figure 1 An exemplary flow chart of a method for failure transfer of an equipment automation manufacturing system provided according to some embodiments of the present application;

[0017] Figure 2 A system architecture diagram of a fault transfer method for an equipment automation manufacturing system provided according to some embodiments of the present application;

[0018] Figure 3 An exemplary flow chart of a method for determining whether an automated manufacturing program is down according to some embodiments of the present application;

[0019] Figure 4An exemplary flow chart of a method for determining whether an automated manufacturing device or a daemon is down according to some embodiments of the present application;

[0020] Figure 5 A schematic diagram of a life cycle architecture of an equipment automation manufacturing system provided according to some embodiments of the present application;

[0021] Figure 6 A schematic diagram of an exemplary structure of a failover strategy provided according to some embodiments of the present application;

[0022] Figure 7 This is a schematic diagram of an exemplary structure of an electronic device provided according to some embodiments of the present application. DETAILED DESCRIPTION

[0023] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0024] First embodiment

[0025] The first embodiment of the present application relates to a fault transfer method for an equipment automation manufacturing system. Figure 1 As shown, the method may include the following steps:

[0026] S110 , performing change event detection on the sub-nodes under the current server persistent node in the equipment automated manufacturing system to determine whether the automated manufacturing program is down.

[0027] Among them, Equipment Automation Programming (EAP) can refer to a system solution that serves as an integrated middleware between the operation layer and the control layer on an automated production line, converting production instructions into equipment control information through protocols, automatically collecting data, automatically loading and unloading machines, and equipment management warnings, and can also achieve information interaction with upper and lower layer systems. EAP is a name for equipment automation management systems in the electronics and semiconductor industries. Due to the increasing cost and complexity of the manufacturing process, equipment automation technology has become a necessary part of semiconductor manufacturing.

[0028] A persistent node can refer to an important data storage method. Once a persistent node is created, it will always exist on the server unless there is an explicit deletion operation. Persistent nodes are suitable for long-term storage of data, including but not limited to configuration information and naming services.

[0029] The child nodes of a persistent node include temporary nodes and temporary sequential nodes. A temporary node is a type of node in the registration center, and the life cycle of a temporary node is bound to the client session. In other words, if the client session that created the temporary node fails, for example, the client crashes or closes the connection with the registration center, then the temporary node will be automatically cleared without the need to actively delete it. In an embodiment of the present invention, when the automated manufacturing program crashes, the temporary node will be automatically deleted; in an embodiment of the present invention, by detecting change events on the child nodes under the persistent node of the current server in the equipment automated manufacturing system, it is determined whether the automated manufacturing program is down.

[0030] S120: Detect the survival status of the automated manufacturing equipment and the daemon program based on the heartbeat packet to determine whether the automated manufacturing equipment or the daemon program is down.

[0031] Among them, the heartbeat packet can refer to a mechanism used to maintain the connection status in network communication, which detects the status of the other party by sending simple command words at regular intervals. The heartbeat packet is a custom command word sent regularly between the client and the server. These command words contain the status information of the sender. If no response is received from the other party within a certain period of time, it can be judged that the other party may be offline or there is a problem with the network connection.

[0032] In the embodiment of the present invention, the heartbeat packets regularly sent by the automated manufacturing equipment and the daemon are detected to determine the survival status of the automated manufacturing equipment and the daemon, thereby determining whether the automated manufacturing equipment and the daemon are down.

[0033] S130. Different failover strategies are formulated according to the downtime scenarios and the number of downtimes to automatically transfer and restore the equipment automation manufacturing system to the state before the downtime.

[0034] The downtime scenario includes at least one of an automated manufacturing program downtime, an automated manufacturing equipment downtime scenario, and a daemon program downtime scenario.

[0035] The equipment automated manufacturing system includes several servers, and one server includes several automated manufacturing program downtime, automated manufacturing equipment downtime scenarios, and daemon downtime. The downtime number may refer to the number of automated manufacturing program downtime, the number of automated manufacturing equipment downtime, and the number of daemon downtime in the current server. When the number of downtime is different, different failover strategies may be selected.

[0036] Failover is a common fault-tolerance mechanism that aims to automatically switch to a backup system or component to continue providing services when a primary system or component fails. For example, in a distributed file system, failover is implemented by using two naming nodes. The two namespaces include a naming node in progress and a backup naming node. When a naming node fails, the backup naming node can take over immediately to ensure high availability of the system. Flink is a popular stream processing framework. Its failover mechanism allows automatic recovery when a task fails. Flink's failover strategy includes regional failover, which divides tasks into multiple regions based on topological connections. When any task fails within a region, all tasks in the region are redeployed. In addition, Flink also involves heartbeat message synchronization and work status synchronization between working nodes and control nodes to handle node failures and task failures.

[0037] In an optional solution of the embodiment of the present invention, all automated manufacturing equipment in the equipment automated manufacturing system are equipped with a watchdog, which can wake up any automated manufacturing equipment when it fails.

[0038] The failover strategy includes, but is not limited to, transferring the faulty program or device to a backup server or other main server, restarting the server, and restarting the automated manufacturing program. In the embodiment of the present invention, different failover strategies are formulated according to the downtime scenario and the number of downtimes to automatically transfer and restore the equipment automated manufacturing system to the state before the downtime. Figure 2 When the daemon 1 of host 1 goes down, by formulating a corresponding failover strategy, the daemon 2 in host 2 is used to execute the operations in the daemon 1 of host 1, so that the equipment automation manufacturing system is restored to the state before the downtime.

[0039] It is not difficult to find that compared with the related art, the solution provided in the embodiment of the present application determines whether the automated manufacturing program is down by detecting the change event of the subnode under the current server persistent node in the equipment automated manufacturing system; detects the survival status of the automated manufacturing equipment and the daemon according to the heartbeat packet to determine whether the automated manufacturing equipment or the daemon is down; formulates different failover strategies according to the downtime scenario and the number of downtimes to automatically transfer and restore the equipment automated manufacturing system to the state before the downtime. The technical solution of the embodiment of the present invention solves the problem of downtime of a single automated manufacturing device, multiple automated manufacturing devices and servers in the equipment automated manufacturing system of the semiconductor industry by customizing the failover strategy that meets different downtime scenarios, realizes more fine-grained control, and improves the adaptability and robustness of the equipment automated manufacturing system.

[0040] Second embodiment

[0041] The second embodiment of the present application relates to a method for fault transfer of an equipment automated manufacturing system. The second implementation is an improvement on the first embodiment. The specific improvement is: a detailed description of determining whether the automated manufacturing program is down is provided, such as Figure 3 As shown, the following steps may be included:

[0042] S310: Build a registration center cluster, and connect the automated manufacturing program and the daemon program to the registration center cluster.

[0043] Among them, through the construction of the registration center cluster, even if some registration center nodes fail, the overall service is still available, which can ensure the high stability of the equipment automation manufacturing system. Each node in the registration center cluster can share the request pressure and improve the overall processing capacity and response speed of the equipment automation manufacturing system.

[0044] S320: Determine whether a persistent node identified by the current server exists in the equipment automated manufacturing system.

[0045] S330: If the persistent node does not exist, create a persistent node identified by the current server and child nodes under the persistent node; the child nodes under the persistent node include temporary sequential nodes and temporary nodes.

[0046] Among them, the automated manufacturing program and the daemon program are connected to the registration center cluster, check whether the persistent node identified by the current server exists, and if not, create a persistent node identified by the current server, and create temporary sequential nodes and temporary nodes under the persistent node under the current server persistent node.

[0047] The life cycle of a temporary sequence node is associated with the client session that created the temporary sequence node. When the session between the client that created the temporary sequence node and the registry cluster ends (for example, the client is disconnected or crashes), the temporary sequence node will be automatically deleted. This feature enables the registry to detect state changes of the node creator and is suitable for representing transient states or session-related data, such as identifying the temporary state of the client or the existence of the session in a distributed environment. The registry will add a monotonically increasing number to the temporary sequence node name. This sequential numbering helps to implement more complex synchronization and sorting logic, such as in the implementation of distributed locks, by comparing node numbers to determine the order of acquiring locks and other application scenarios.

[0048] S340: Perform change event detection on the sub-nodes under the current server persistent node in the equipment automated manufacturing system to determine whether the automated manufacturing program is down.

[0049] In the embodiment of the present invention, the change event of the temporary node or temporary sequential node under the persistent node is detected to determine whether the automated manufacturing program is down. The change event may refer to whether the temporary node or temporary sequential node is deleted.

[0050] As an optional but non-limiting implementation, the change event detection of the sub-nodes under the current server persistent node in the equipment automated manufacturing system to determine whether the automated manufacturing program is down includes but is not limited to steps A1-A2:

[0051] Step A1: Perform change event detection on the sub-nodes under the current server persistent node in the equipment automated manufacturing system to determine whether the temporary node under the current server persistent node is deleted.

[0052] Step A2: If the temporary node has been deleted, it is determined that the automated manufacturing program is down.

[0053] In which, since the temporary node will be automatically deleted when the automated manufacturing program is down, the embodiment of the present invention determines whether the automated manufacturing program is down by determining whether the temporary node under the current server persistent node is deleted as an example. If the temporary node is deleted, it is determined that the automated manufacturing program is down.

[0054] It is not difficult to find that in the embodiment of the present application, by detecting whether the temporary node under the persistent node identified by the current server in the equipment automated manufacturing system is deleted, it is determined whether the automated manufacturing program is down, thereby improving the fault detection efficiency of the equipment automated manufacturing system.

[0055] Third embodiment

[0056] The third embodiment of the present application relates to a method for failover of an automated manufacturing system. The third implementation is an improvement on the first embodiment. The specific improvement is that: a detailed description is given to determining whether the automated manufacturing device or the daemon is down, such as Figure 4 As shown, the following steps may be included:

[0057] S410, building a manager cluster, and connecting the automated manufacturing equipment and the daemon to the manager cluster.

[0058] Among them, a manager cluster is built, and the automated manufacturing equipment and the daemon program are connected to the manager cluster. The automated manufacturing equipment and the daemon program are connected to the manager cluster, so that the automated manufacturing equipment and the daemon program can send heartbeat packets to the manager program at regular intervals.

[0059] S420, receiving the heartbeat packet sent regularly by the automated manufacturing equipment and the daemon, and determining whether the received heartbeat packet is abnormal.

[0060] The manager program receives the heartbeat packets sent periodically by the automated manufacturing equipment and the daemon program, and determines whether the heartbeat packets are abnormal.

[0061] As an optional but non-limiting implementation, the step of receiving the heartbeat packet sent periodically by the automated manufacturing equipment and the daemon program and determining whether the received heartbeat packet is abnormal includes but is not limited to steps B1-B3:

[0062] Step B1: Set the preset time interval and number of times of sending the heartbeat packet.

[0063] Step B2: within a preset time, determine whether the received heartbeat packet meets a preset sending time interval or a preset sending number of times.

[0064] Step B3: If it is determined that the received heartbeat packet does not meet the preset sending time interval or the preset sending number of times, it is determined that the received heartbeat packet is abnormal.

[0065] Among them, a preset sending time interval and a preset sending number of times of the heartbeat packet are set, and a heartbeat packet is sent at the preset sending time interval; it is determined whether the received heartbeat packet meets the preset sending time interval or the preset sending number of times within the preset time; if it is determined that the received heartbeat packet does not meet the preset sending time interval or the preset sending number of times, it is determined that the received heartbeat packet is abnormal. For example, the preset sending time interval is 500ms, and the preset sending number of times is 3 times. Under normal circumstances, the manager program receives a heartbeat packet once in 500ms and receives 3 heartbeat packets in 1500ms. If the manager program does not receive any heartbeat packet within 1500ms, it is determined that the heartbeat packet is abnormal. If one or two heartbeat packets are received within 1500ms, it can be considered as an abnormality caused by network delay, and it is not regarded as an abnormal phenomenon of the heartbeat packet at this time.

[0066] Optionally, in order to avoid misjudgment of the heartbeat packet abnormality due to network delay, the limit time for receiving heartbeat packets can be extended; for example, determine whether at least one heartbeat packet is received within 2000ms. If at least one heartbeat packet is received within 2000ms, it can be determined that there is no abnormality in the received heartbeat packet; if no heartbeat packet is received within 2000ms, it is determined that the received heartbeat packet is abnormal.

[0067] S430: If the received heartbeat packet is abnormal, it is determined that the survival status of the automated manufacturing device and the daemon is abnormal, and the automated manufacturing device or the daemon is down.

[0068] Among them, the automated manufacturing equipment and the daemon program periodically send heartbeat packets to the manager program, and the manager program records the status of the automated manufacturing equipment and monitors the survival status of the automated manufacturing equipment and the daemon program; after determining that the heartbeat packet is abnormal, it is determined that the survival status of the automated manufacturing equipment and the daemon program is abnormal, thereby determining that the automated manufacturing equipment or the daemon program is down.

[0069] Optionally, the time limit for determining whether the automated manufacturing equipment or the daemon is down can be determined based on the preset sending time interval of the heartbeat packet; for example, the preset number of sending times is 3 times, and the preset sending time interval of the heartbeat packet can be configured to 100ms, 500ms or 1000ms, then the time limit for determining whether the automated manufacturing equipment or the daemon is down can be configured to 300ms, 1500ms or 3000ms accordingly. In order to avoid misjudging the abnormal scene of the heartbeat packet due to network delay, the time limit for determining whether the automated manufacturing equipment or the daemon is down can be extended, and can be configured to 400ms, 2000ms or 4000ms accordingly. Among them, the time limit for determining whether the automated manufacturing equipment or the daemon is down can refer to the time for the manager program to receive the heartbeat packet within the preset time. In the embodiment of the present invention, there is no specific restriction on the time limit for determining whether the automated manufacturing equipment or the daemon is down, and it can be set according to the actual situation.

[0070] In an optional scheme of an embodiment of the present invention, the preset sending time interval and preset sending number of heartbeat packets are relatively small. When it is determined that the automated manufacturing equipment or the daemon is down, the time for immediate fault migration, restart and data recovery is also relatively short. This is imperceptible to the user and can improve the user's experience.

[0071] It is not difficult to find that in the embodiment of the present application, by setting the time and number of times of receiving heartbeat packets, the survival status of the automated manufacturing equipment and the daemon is detected to determine whether the automated manufacturing equipment or the daemon is down, and timely processing is performed when the automated manufacturing equipment or the daemon is down, thereby reducing service interruption time, improving the reliability and availability of the equipment automated manufacturing system, and improving the user experience.

[0072] Fourth embodiment

[0073] The fourth embodiment of the present application relates to a method for fault transfer of an equipment automation manufacturing system. The fourth implementation is an improvement based on the first embodiment, the second embodiment or the third embodiment, and the specific improvement is: a detailed description of the formulation of different fault transfer strategies may include:

[0074] The automatic transfer is performed by formulating different failover strategies based on the downtime scenario and the number of downtimes, including:

[0075] If at least one automated manufacturing device is down, the server of the automated manufacturing device is not down or the daemon of the automated manufacturing device can send a heartbeat packet, a restart strategy of the automated manufacturing program can be formulated for the at least one automated manufacturing device, and automatic transfer can be performed.

[0076] Among them, see Figure 5 When one or more automated manufacturing devices are detected to be down by the management tool, but the server of the automated manufacturing device is not down or the daemon of the automated manufacturing device can still send heartbeat packets to the management tool, it supports restarting the downed automated manufacturing program on the current server.

[0077] The automatic transfer is performed by formulating different failover strategies according to the downtime scenarios and the number of downtimes, and further includes:

[0078] If there is a backup server in the equipment automation manufacturing system, the backup server will be selected first for automatic transfer and restore the equipment automation manufacturing system to the state before the downtime;

[0079] Or, if the number of downtimes is greater than a preset downtime threshold, the polling server is preferentially selected to automatically transfer and restore the equipment automated manufacturing system to the state before the downtime;

[0080] Or, if there is no backup server in the equipment automated manufacturing system or the server system resources are insufficient, a target main server is selected from the main servers for automatic transfer; the target main server is connected to the least automated manufacturing equipment.

[0081] The failover strategy includes at least one failover strategy of polling standby servers, random primary servers, retry servers, designated servers, polling primary servers, random primary servers and target primary servers.

[0082] Among them, if there is a backup server in the equipment automation manufacturing system, the backup server is preferred for automatic transfer; if the number of downtimes is greater than the preset downtime threshold, the polling server is preferred for automatic transfer. For example, if there is a backup server and the number of downtimes is large, the polling backup server strategy can be selected. By using the polling backup server strategy, the pressure between servers can be reduced and the processing speed can be improved.

[0083] Among them, if there is no backup server, or when the server pointed to by the set failover strategy is unreachable, the server resources are insufficient, or the server resources exceed the threshold, it is necessary to adopt a backup plan using other strategies, such as the minimum number of connections strategy, to reselect an available server for migration; the minimum number of connections strategy can refer to selecting a target main server from the main server for automatic transfer; the target main server has the least number of automated manufacturing devices connected to it.

[0084] If there is no backup server, select another master server for failover. If the number of downtimes is large, you can choose to poll the master server strategy to reduce the pressure on each server. Figure 6 In the absence of a backup server, the automated manufacturing program 1 and the automated manufacturing program 2 in the main server 2 go down, and the automated manufacturing program 1 and the automated manufacturing program 2 in the main server 2 are transferred to the automated manufacturing programs of the main server 1 and the main server 3 respectively.

[0085] If the corresponding server is specified for failover, the corresponding specified server policy is selected for failover. If the server needs to be retried, the corresponding retry server policy is selected for failover.

[0086] In an optional solution of the embodiment of the present invention, when the state machine (E40, E87, E94) changes, when a key event that triggers the process is received, or when the equipment automation manufacturing system sends a remote control command, the automated manufacturing program updates the work-in-progress information, process, status and other related information in real time, and saves the information in the database (Redis\Oracle\Mysql\Guassdb\PostgreSQL). The specific implementation can be represented as follows: the automated manufacturing program assembles the relevant information and sends it to the daemon program or the manager program, and the manager program performs data persistence operations. When the automated manufacturing program crashes and migrates to a new server and starts successfully, it will request the daemon program or the manager program to persist the data and restore to the state of the automated manufacturing program before the crash, so as to ensure that the equipment automation manufacturing system can continue to run from the correct state after the fault migration.

[0087] It is not difficult to find that in the embodiment of the present application, different failover strategies are formulated according to the downtime scenario and the number of downtimes to automatically transfer and restore the equipment automation manufacturing system to the state before the downtime. The problem of downtime of a single automated manufacturing device, multiple automated manufacturing devices and servers in the equipment automation manufacturing system of the semiconductor industry is solved, and more fine-grained control is achieved, which improves the adaptability and robustness of the equipment automation manufacturing system.

[0088] The step division of the above methods is only for the purpose of clear description. When implemented, they can be combined into one step or some steps can be split and decomposed into multiple steps. As long as they include the same logical relationship, they are all within the scope of protection of this patent; adding insignificant modifications to the algorithm or process or introducing insignificant designs without changing the core design of the algorithm and process are all within the scope of protection of this patent.

[0089] Fifth embodiment

[0090] In addition, some embodiments of the present application also provide an electronic device. The electronic device may be a digital computer in various forms, such as a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, etc. The electronic device may also be a mobile device in various forms, such as a personal digital processing, a cellular phone, a smart phone, a wearable device, and other similar computing devices.

[0091] The electronic device includes: one or more processors; and a memory storing computer program instructions, wherein when the computer program instructions are executed, the processor executes the steps of the method provided in any one or more of the above embodiments. Figure 7 An exemplary structural diagram of the electronic device is disclosed. Figure 7 As shown, the electronic device includes: one or more processors 1101, a memory 1102, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. The various components are connected to each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed in the electronic device, including instructions stored in or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some other embodiments, if necessary, multiple processors and / or multiple buses can be used with multiple memories and multiple memories. Similarly, multiple electronic devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Among them, the components shown in this article, their connections and relationships, and their functions are only examples, and are not intended to limit the implementation of the present application described and / or required herein.

[0092] The electronic device may further include: an input device 1103 and an output device 1104. The processor 1101, the memory 1102, the input device 1103 and the output device 1104 may be connected via a bus or other means. Figure 7 The example of connecting through bus is taken in the following.

[0093] The input device 1103 can receive input digital or character information, and generate key signal input related to the user settings and function control of the electronic device, such as a touch screen, a keypad, a mouse, a track pad, a touch pad, an indicator rod, one or more mouse buttons, a trackball, a joystick and other input devices. The output device 1104 may include a display device, an auxiliary lighting device (e.g., an LED) and a tactile feedback device (e.g., a vibration motor), etc. The display device may include, but is not limited to, a liquid crystal display (LCD), a light emitting diode (LED) display and a plasma display. In some embodiments, the display device may be a touch screen.

[0094] To provide interaction with a user, the electronic device may be a computer. The computer has: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball), through which the user can provide input to the computer. Other types of devices may also be used to provide interaction with a user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).

[0095] Sixth embodiment

[0096] In the embodiments of the present application, a computer program / instruction is stored on a computer-readable medium, and when the computer program / instruction is executed by a processor, the steps of the method provided by any one or more of the above embodiments are implemented. The computer-readable medium may be included in the electronic device described in the above embodiments; or it may exist independently without being assembled into the device. The above computer-readable medium carries one or more computer-readable instructions.

[0097] The memory 1102 can be used as a non-transient computer-readable storage medium, which can be used to store non-transient software programs, non-transient computer executable programs and modules. The processor 1101 executes various functional applications and data processing of the server by running the non-transient software programs, instructions and modules stored in the memory 1102, so as to implement the program instructions / modules corresponding to the method provided by any one or more embodiments in the embodiments of the present application.

[0098] The memory 1102 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 1102 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory 1102 may optionally include a memory remotely arranged relative to the processor 1101, and these remote memories may be connected to the electronic device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0099] It should be noted that the computer-readable medium described in this application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, a computer-readable medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device.

[0100] Computer readable media include permanent and non-permanent, removable and non-removable media, and can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, modules of programs or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disk (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission medium that can be used to store information that can be accessed by a computing device.

[0101] Computer program code for performing the operations of the present application may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0102] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware or any combination thereof. For example, an application specific integrated circuit (ASIC), a general-purpose computer or any other similar hardware device may be used to implement the embodiments. In some embodiments, the software program of the present application may be executed by a processor to implement the above steps or functions. Similarly, the software program of the present application (including related data structures) may be stored in a computer-readable recording medium, such as a RAM memory, a magnetic or optical drive or a floppy disk and the like. In addition, some steps or functions of the present application may be implemented by hardware, for example, as a circuit that cooperates with a processor to perform various steps or functions.

[0103] Seventh embodiment

[0104] The computer program product provided in the embodiment of the present application includes one or more computer programs / instructions, which, when executed by the processor, generate in whole or in part the process or function described in the embodiment of the present application. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website site, a computer, a server, or a data center to another website site, a computer, a server, or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or a data center that includes one or more available media integrated. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid state disk (SSD)), etc.

[0105] The flow chart or block diagram in the accompanying drawings shows the possible architecture, function and operation of the equipment, method and computer program product according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented with a dedicated system for hardware that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0106] The scope of the present application is defined by the appended claims rather than the above description, and it is intended that all changes that fall within the meaning and scope of the equivalent elements of the claims are included in the present application. Any figure mark in the claims should not be regarded as limiting the claims involved. In addition, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices stated in the device claim may also be implemented by one unit or device through software or hardware. The words "first", "second", etc. are only used to distinguish the description, and do not indicate any particular order, nor can they be understood as indicating or implying relative importance.

[0107] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily mention changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims, and the above embodiments should be regarded as exemplary and non-restrictive.

Claims

1. A method for failure transfer of an equipment automation manufacturing system, characterized in that: The method comprises: Perform change event detection on the sub-nodes under the current server persistent node in the equipment automated manufacturing system to determine whether the automated manufacturing program is down; Detect the survival status of the automated manufacturing equipment and the daemon program based on the heartbeat packet to determine whether the automated manufacturing equipment or the daemon program is down; Formulate different failover strategies according to the downtime scenarios and downtime quantity to automatically transfer and restore the equipment automated manufacturing system to the state before the downtime; the downtime scenarios include at least one of the automated manufacturing program downtime, the automated manufacturing equipment downtime scenario and the daemon downtime scenario; The detecting of the survival status of the automated manufacturing equipment and the daemon program according to the heartbeat packet to determine whether the automated manufacturing equipment or the daemon program is down includes: Build a manager cluster and connect the automated manufacturing equipment and the daemon to the manager cluster; receive the heartbeat packets sent regularly by the automated manufacturing equipment and the daemon, and determine whether the received heartbeat packets are abnormal; if the received heartbeat packets are abnormal, determine that the survival status of the automated manufacturing equipment and the daemon is abnormal, and the automated manufacturing equipment or the daemon is down; Different failover strategies are formulated based on the downtime scenario and the number of downtimes to automatically transfer and restore the equipment automation manufacturing system to the state before the downtime, including: The automated manufacturing program assembles relevant information and sends it to the daemon or manager program, which then performs data persistence operations. When the automated manufacturing program crashes and is successfully migrated to a new server, it requests persistent data from the daemon or manager program and restores to the automated manufacturing program state before the crash, ensuring that the equipment automated manufacturing system can continue to operate in the correct state after the fault migration.

2. The method according to claim 1, characterized in that Before performing change event detection on a sub-node under a current server persistent node in the equipment automated manufacturing system, the method includes: Building a registry center cluster, and connecting the automated manufacturing program and the daemon program to the registry center cluster; Determine whether a persistent node identified by a current server in the equipment automated manufacturing system exists; If the persistent node does not exist, a persistent node identified by the current server and child nodes under the persistent node are created; the child nodes under the persistent node include temporary sequential nodes and temporary nodes.

3. The method according to claim 1, characterized in that The step of detecting a change event of a sub-node under a persistent node of a current server in the equipment automated manufacturing system to determine whether the automated manufacturing program is down includes: Perform change event detection on the sub-nodes under the current server persistent node in the equipment automation manufacturing system to determine whether the temporary node under the current server persistent node is deleted; If the temporary node has been deleted, it is determined that the automated manufacturing program has crashed.

4. The method according to claim 1, characterized in that: The receiving of the heartbeat packet sent regularly by the automated manufacturing equipment and the daemon program, and determining whether the received heartbeat packet is abnormal, includes: Set the preset time interval and number of times to send the heartbeat packet; Within a preset time, determine whether the received heartbeat packet meets a preset sending time interval or a preset sending number of times; If it is determined that the received heartbeat packet does not meet the preset sending time interval or the preset sending number of times, it is determined that the received heartbeat packet is abnormal.

5. The method according to claim 1, characterized in that The automatic transfer is performed by formulating different failover strategies based on the downtime scenario and the number of downtimes, including: If at least one automated manufacturing device is down, the server of the automated manufacturing device is not down or the daemon of the automated manufacturing device can send a heartbeat packet, a restart strategy of the automated manufacturing program can be formulated for the at least one automated manufacturing device, and automatic transfer can be performed.

6. The method according to claim 1, characterized in that The automatic transfer is performed by formulating different failover strategies according to the downtime scenarios and the number of downtimes, and further includes: If there is a backup server in the equipment automation manufacturing system, the backup server will be selected first for automatic transfer and restore the equipment automation manufacturing system to the state before the downtime; Or, if the number of downtimes is greater than a preset downtime threshold, the polling server is preferentially selected to automatically transfer and restore the equipment automated manufacturing system to the state before the downtime; Or, if there is no backup server in the equipment automated manufacturing system or the server system resources are insufficient, a target main server is selected from the main servers for automatic transfer; the target main server is connected to the least automated manufacturing equipment.

7. An electronic device, characterized in that: The electronic device comprises: one or more processors; and A memory storing computer program instructions, wherein when the computer program instructions are executed, the processor executes the steps of the failover method of the equipment automation manufacturing system according to any one of claims 1 to 6.

8. A computer readable medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the fault transfer method of the equipment automation manufacturing system described in any one of claims 1 to 6 are implemented.

9. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the steps of the fault transfer method of the equipment automation manufacturing system described in any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Method and system for monitoring content delivery network (CDN) equipment status

    CN102111310A

  • Fail-over system of semiconductor equipment server and method thereof

    KR1020150028077A