Computer program product, computer-implemented method, and computing device for quality prediction using process data

By training a supervised learning machine learning model, using the measurement data and status information of the joint operation to predict the quality and abnormality of the joint, the problem of difficulty in effectively predicting the quality of the joint in the prior art is solved, and efficient quality assurance is achieved.

CN117413238BActive Publication Date: 2025-06-27SAS INSTITUTE INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202280038926.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-07-19
Filing Date
2022-01-21
Publication Date
2025-06-27
Estimated Expiration
2042-01-21

AI Technical Summary

Technical Problem

The prior art is difficult to effectively predict the mass of the joint in the engagement operation, especially without breaking the joint.

Method used

By training a supervised learning machine learning model, abnormalities and quality problems in the bonding operation are predicted using process data generated from the measured values ​​of the bonding operation and the state of multiple wires after bonding.

Benefits of technology

It is achieved to accurately predict the quality and potential abnormalities of the joint without destroying the joint, improving the quality assurance and efficiency of the joint operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117413238B_ABST
    Figure CN117413238B_ABST
Patent Text Reader

Abstract

A computing device (2002) accesses a machine learning model (2050) trained on training data (2032) of a first bonding operation (1308, 2040A) (e.g., ball and / or pin bonding). The first bonding operation includes an operation for bonding a first set of wires (1504) to a first set of surfaces (1506, 1508). The machine learning model is trained by supervised learning. The device receives input data (2070) indicating process data (2074) generated from measurements of a second bonding operation (2040B). The second bonding operation includes an operation for bonding a second set of wires to a second set of surfaces. The device weights the input data according to the machine learning model. The device generates an anomaly prediction (2052) indicating the risk of an anomaly occurring in the second bonding operation based on weighting the input data according to the machine learning model. The device outputs the anomaly prediction to control the second bonding operation.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims the benefit and priority of U.S. Provisional Application No. 63 / 223,491, filed Jul. 19, 2021, and U.S. Non-Provisional Application No. 17 / 581,113, filed Jan. 21, 2022, under 35 U.S.C. § 119, the entire disclosures of each of which are incorporated herein by reference. Technical Field

[0003] The present disclosure generally relates to computer-generated predictions for joining operations. Background Art

[0004] Industries such as manufacturing and construction use joining techniques to bond materials together (e.g., welding techniques). Quality assurance testing can be used to determine the quality of the joints. For example, in destructive testing, a subset of the joints is destroyed to make predictions about the quality of the non-destroyed joints. In conventional non-destructive testing, a tester manually inspects the joints to make predictions about the quality of the joints. Summary of the Invention

[0005] In an example embodiment, a computer program product is tangibly embodied in a non-transitory machine-readable storage medium. The computer program product includes instructions operable to cause a computing system to access a machine learning model trained on training data of a first bonding operation. The first bonding operation includes an operation for bonding a plurality of wires of a first group to a first group of surfaces. The machine learning model is trained by supervised learning, including receiving the training data. The training data includes process data generated from measurements of the first bonding operation; and the states of the plurality of wires after being bonded to the first group of surfaces. Each state of the states includes one or more candidate results of a target related to detecting one or more anomalies in the first bonding operation. The machine learning model is trained by supervised learning, including generating one or more weights of the process data such that the process data input to the machine learning model predicts the one or more candidate results of the target. The computer program product includes instructions operable to cause a computing system to receive input data indicating process data generated from measurements of a second bonding operation. The second bonding operation includes an operation for bonding a plurality of wires of a second group to a second group of surfaces. The plurality of wires of the second group is different from the plurality of wires of the first group. The second group of surfaces is different from the first group of surfaces. The computer program product includes instructions operable to cause a computing system to weight the input data according to the machine learning model. The computer program product includes instructions operable to cause a computing system to generate an anomaly predictor indicating a risk of occurrence of an anomaly in the second bonding operation based on weighting the input data according to the machine learning model. The computer program product includes instructions operable to cause a computing system to output the anomaly predictor to control the second bonding operation.

[0006] In one or more embodiments, the first bonding operation bonds wires of a plurality of wires of a first group to corresponding surfaces of a first group of surfaces to form an integrated circuit chip. One or more candidate results include one or more defective chip results of the integrated circuit chip in the first bonding operation. The instructions are operable to cause a computing system to generate an anomaly predictor of a risk of an anomaly in the integrated circuit chip manufacturing process in the second bonding operation.

[0007] In one or more embodiments, the instructions are operable to cause a computing system to receive feedback indicating whether the anomaly predictor correctly or incorrectly predicts an anomaly in a particular chip manufactured in the second bonding operation. The instructions are operable to cause a computing system to update the machine learning model based on the feedback.

[0008] In one or more embodiments, a second bonding operation is performed by a chip manufacturing system. Instructions are operable to cause a computing system to adjust a bonding operation following the second bonding operation performed by the chip manufacturing system based on one or more of: an anomaly prediction measure indicative of a risk of an anomaly occurring in the second bonding operation; and feedback indicative of whether the anomaly prediction measure correctly or incorrectly predicts an anomaly occurring in a particular chip fabricated in the second bonding operation.

[0009] In one or more embodiments, instructions are operable to cause a computing system to receive training data by selectively picking a subset of parameter types observed in a first bonding operation. Measurements of the first bonding operation are measurements for the subset of parameter types.

[0010] In one or more embodiments, the training data includes process data generated by deriving information that takes into account relationships between measurement types in the first bonding operation from a variety of different measurement types. Instructions are operable to receive input data indicative of process data generated from measurements of a second bonding operation by deriving information that takes into account relationships between measurement types in the second bonding operation.

[0011] Embodiments herein also include corresponding computer program products, apparatuses, and methods.

[0012] For example, in one embodiment, a computer program product is tangibly embodied in a non-transitory machine-readable storage medium. The computer program product includes instructions operable to cause a computing system to train a machine learning model by receiving training data. The training data includes process data generated from measurements of a first bonding operation; and the states of a plurality of wires after being bonded to a first set of surfaces. Each state of the states includes one or more candidate results of a target related to detecting one or more anomalies in the first bonding operation. The computer program product includes instructions operable to cause the computing system to train the machine learning model by generating one or more weights of the process data such that the process data input to the machine learning model predicts the one or more candidate results of the target.

[0013] In another example embodiment, a computing device is provided. The computing device includes, but is not limited to, a processor and a memory. The memory contains instructions that, when executed by the processor, control the computing device to access a machine learning model trained on training data of a first bonding operation. The first bonding operation includes an operation for bonding a first set of multiple wires to a first set of surfaces. The machine learning model is trained by supervised learning, including receiving the training data. The training data includes process data generated from measurements of the first bonding operation; and the states of the multiple wires after being bonded to the first set of surfaces. Each state of the states includes one or more candidate results of a target related to detecting one or more anomalies in the first bonding operation. The machine learning model is trained by supervised learning, including generating one or more weights of the process data such that the process data input to the machine learning model predicts the one or more candidate results of the target. The memory contains instructions that, when executed by the processor, control the computing device to receive input data indicating process data generated from measurements of a second bonding operation. The second bonding operation includes an operation for bonding a second set of multiple wires to a second set of surfaces. The second set of multiple wires is different from the first set of multiple wires. The second set of surfaces is different from the first set of surfaces. The computer program product contains instructions operable to cause the computing device to weight the input data according to the machine learning model. The memory contains instructions that, when executed by the processor, control the computing device to generate an anomaly prediction indicating the risk of occurrence of an anomaly in the second bonding operation based on weighting the input data according to the machine learning model. The memory contains instructions that, when executed by the processor, control the computing device to output the anomaly prediction to control the second bonding operation.

[0014] In another example embodiment, a computing device is provided. The computing device includes, but is not limited to, a processor and a memory. The memory contains instructions that, when executed by the processor, control the computing device to train a machine learning model by receiving training data. The training data includes process data generated from measurements of a first bonding operation; and the states of multiple wires after being bonded to a first set of surfaces. Each state of the states includes one or more candidate results of a target related to detecting one or more anomalies in the first bonding operation. The computer program product contains instructions operable to cause the computing device to train the machine learning model by generating one or more weights of the process data such that the process data input to the machine learning model predicts the one or more candidate results of the target.

[0015] In one or more embodiments, the computing device is a computing system or part of a computing system.

[0016] In another example embodiment, a computer-implemented method is provided. The method includes accessing a machine learning model trained on training data of a first bonding operation. The first bonding operation includes an operation for bonding a plurality of wires of a first group to a first set of surfaces. The machine learning model is trained by supervised learning, including receiving the training data. The training data includes process data generated from measurements of the first bonding operation; and the states of the plurality of wires after being bonded to the first set of surfaces. Each state of the states includes one or more candidate results of a target related to detecting one or more anomalies in the first bonding operation. The machine learning model is trained by supervised learning, including generating one or more weights of the process data such that the process data input to the machine learning model predicts the one or more candidate results of the target. The method includes receiving input data indicative of process data generated from measurements of a second bonding operation. The second bonding operation includes an operation for bonding a plurality of wires of a second group to a second set of surfaces. The plurality of wires of the second group are different from the plurality of wires of the first group. The second set of surfaces is different from the first set of surfaces. The method includes weighting the input data according to the machine learning model. The method includes generating an anomaly prediction indicative of a risk of an anomaly occurring in the second bonding operation based on weighting the input data according to the machine learning model. The method includes outputting the anomaly prediction to control the second bonding operation.

[0017] In another example embodiment, a computer-implemented method is provided. The method includes training a machine learning model by receiving training data. The training data includes process data generated from measurements of a first bonding operation; and the states of a plurality of wires after being bonded to a first set of surfaces. Each state of the states includes one or more candidate results of a target related to detecting one or more anomalies in the first bonding operation. The method includes training the machine learning model by generating one or more weights of the process data such that the process data input to the machine learning model predicts the one or more candidate results of the target.

[0018] In any of these computer program products, devices, and methods, one or more candidate results may include one or more destructive quality assurance tests of individual wires or bondings in the first bonding operation. The anomaly prediction may be a prediction of an anomaly in an individual wire or bonding in the second bonding operation without performing a destructive quality assurance test on the individual wire or bonding.

[0019] In any of these computer program products, devices, and methods, the first bonding operation may include one of the following: a ball bonding operation, where the destructive quality assurance test includes a ball shear test for testing the ball bond; and a stitch bonding operation, where the destructive quality assurance test includes a stitch pull test for testing the stitch bond.

[0020] In any of these computer program products, devices, and methods, one or more anomalies may be associated with one or more of the following: a floating condition; lead frame contamination; and die tilt.

[0021] In any of these computer program products, devices, and methods, a second bonding operation may bond the wires of a second set of multiple wires to form an integrated circuit chip. The process data generated from the measurements of the second bonding operation may be related to a specific chip of the integrated circuit chip as a whole and derived from the measurement data of the wires bonded in the second bonding operation and associated with the specific chip.

[0022] In any of these computer program products, devices, and methods, the input data may include real-time sensor measurements received during the second bonding operation. For a given wire of the second set of multiple wires, the sensor measurements may include one or more of the following: heat measurements; power measurements; force measurements; electric flame-off (EFO) measurements; and ultrasonic measurements.

[0023] In any of these computer program products, devices, and methods, the input data may include real-time sensor measurements received during the second bonding operation. The sensor measurements may include measurements of the bonding system involved in the second bonding operation. Anomaly prediction may control the second bonding operation to correct one or more anomalies in the bonding system involved in the second bonding operation and / or reduce the occurrence of one or more anomalies.

[0024] In any of these computer program products, devices, and methods, the input data may include received sensor measurements marked with raw information indicating one or more of the identifier or location of a specific wire, die, or chip involved in the second bonding operation. Anomaly prediction may identify anomalies occurring during the second bonding operation and be associated with the raw information to indicate the location of the anomalies.

[0025] In any of these computer program products, devices, and methods, the input data may include derived data including one or more of the following: a generated value indicating a median or average of measurements related to multiple wires bonded in a particular chip during a second bonding operation; a set of generated deviations including the deviation of each of the multiple wires from the value; and a generated metric of the chip considering the set of deviations.

[0026] In any of these computer program products, devices, and methods, the training data may include process data generated from the derived data, the derived data including generated outlier data values related to multiple different types of measurements and related to the same wires in a first bonding operation. The input data indicating process data generated from measurements of a second bonding operation may include the generated outlier data values. The generated outlier data values may be related to multiple different types of measurements and related to the same wires in the second bonding operation.

[0027] In any of these computer program products, devices, and methods, one or more weights of the process data may be generated for a gradient boosting model of the training data.

[0028] In any of these computer program products, devices, and methods, the machine learning model may further be trained by multiple generated machine learning models and selected based on k-fold cross-validation.

[0029] In any of these computer program products, devices, and methods, the measurements of the first and second bonding operations may include measurements associated with the process of forming ball bonds.

[0030] In any of these computer program products, devices, and methods, the measurements of the first and second bonding operations may include measurements associated with the process of forming pin bonds.

[0031] In any of these computer program products, devices, and methods, the anomaly prediction may be a prediction of defective bonds between the wires of a second set of multiple wires and the lead frame or die of the second set of surfaces.

[0032] Other features and aspects of the example embodiments are presented in the detailed description below when read in conjunction with the accompanying drawings presented with this application. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 A block diagram illustrating the hardware components of a computing system according to at least one embodiment of the present technology is provided.

[0034] Figure 2Describe an example network that includes a set of example devices that communicate with each other via a switching system and via a network, according to at least one embodiment of the present technology.

[0035] Figure 3 Describe a representation of a conceptual model of a communication protocol system, according to at least one embodiment of the present technology.

[0036] Figure 4 Describe a communication grid computing system that includes multiple control and worker nodes, according to at least one embodiment of the present technology.

[0037] Figure 5 Describe a flowchart that shows an example process for adjusting a communication grid or work items in a communication grid after a node failure, according to at least one embodiment of the present technology.

[0038] Figure 6 Describe a part of a communication grid computing system that includes a control node and a worker node, according to at least one embodiment of the present technology.

[0039] Figure 7 Describe a flowchart that shows an example process for performing data analysis or processing items, according to at least one embodiment of the present technology.

[0040] Figure 8 Describe a block diagram of a component that includes an event stream processing engine (ESPE), according to at least one embodiment of the present technology.

[0041] Figure 9 Describe a flowchart that shows an example process that includes operations performed by an event stream processing engine, according to at least one embodiment of the present technology.

[0042] Figure 10 Describe an ESP system that interfaces between a publishing device and multiple event subscription devices, according to at least one embodiment of the present technology.

[0043] Figure 11 Describe a flowchart of an example of a process for generating and using a machine learning model, according to at least one embodiment of the present technology.

[0044] Figure 12 Describe an example of a machine learning model that is a neural network, according to at least one embodiment of the present technology.

[0045] Figures 13 to 14 Describe an example flowchart for manufacturing an integrated circuit chip, according to at least one embodiment of the present technology.

[0046] Figure 15A and 15BDescribe some example components of an integrated circuit chip and corresponding bonding according to at least one embodiment of the present technology.

[0047] Figures 16A to 16D Describe example bonding operations involved in manufacturing an integrated circuit chip according to at least one embodiment of the present technology.

[0048] Figure 17 Describe an example integrated circuit chip according to at least one embodiment of the present technology.

[0049] Figures 18A to 18C Describe the destructive ball shear test process for ball bonds.

[0050] Figures 19A to 19C Describe the destructive wire pull test process for wire bonds.

[0051] Figure 20A Describe an example block diagram of a training system in at least one embodiment of the present technology.

[0052] Figure 20B Describe an example block diagram of a control system in at least one embodiment of the present technology.

[0053] Figure 21A Is a flowchart illustrating an example method for training a machine learning model according to at least one embodiment of the present technology.

[0054] Figure 21B Is a flowchart illustrating an example method for controlling a bonding operation according to at least one embodiment of the present technology.

[0055] Figure 21C Is a flowchart illustrating an example method for updating a machine learning model for controlling a bonding operation according to at least one embodiment of the present technology.

[0056] Figure 22A Is a chart illustrating the relationship between a motion feature pattern and the position of a corresponding chip on a lead frame according to at least one embodiment of the present technology.

[0057] Figure 22B Describe according to at least one embodiment of the present technology corresponding to Figure 22A The example positions of the chips on the lead frame of the chart.

[0058] Figure 23A Describe example quality assurance (QA) data comparing training data and test data of a prediction model for destructive tests of bonds formed during a bonding operation.

[0059] Figure 23B Describe an example method for generating derived processing data according to at least one embodiment of the present technology.

[0060] Figure 24 is a diagram illustrating an example diagram for deriving process data associated with an engagement operating system according to at least one embodiment of the present technology.

[0061] Figure 25 is a diagram illustrating the correspondence between a predicted ball shear value modeled according to an embodiment of the present disclosure and an actual ball shear value obtained by a destructive test.

[0062] Figure 26 is a diagram illustrating the correspondence between a predicted pin pull value modeled according to an embodiment of the present disclosure and an actual pin pull value obtained by a destructive test.

[0063] Figure 27A is a functional block diagram of a stack of an event stream processing (ESP) system according to at least one embodiment of the present technology.

[0064] Figure 27B is a flowchart illustrating an example method for generating a machine learning model according to at least one embodiment of the present technology.

[0065] Figure 28 is a functional block diagram of a computer program product according to at least one embodiment of the present technology. Detailed Description

[0066] In the following description, for purposes of illustration, specific details are set forth to provide a thorough understanding of embodiments of the present technology. However, it will be understood that various embodiments may be practiced without these specific details. The figures and the description are not intended to be restrictive.

[0067] The following description provides only example embodiments and is not intended to limit the scope, applicability, or configuration of the present disclosure. Indeed, the following description of example embodiments will provide those skilled in the art with a thorough description for implementing the example embodiments. It should be understood that various changes may be made in the function and arrangement of elements without departing from the spirit and scope of the technology as set forth in the claims.

[0068] Specific details are given in the following description to provide a thorough understanding of the embodiments. However, those skilled in the art will understand that the embodiments may be practiced without these specific details. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form to avoid obscuring the embodiments with unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail to avoid obscuring the embodiments.

[0069] Moreover, it should be noted that individual embodiments may be described as a processing procedure, which is depicted as a flowchart (flowchart / flow diagram), a data flow diagram, a structural diagram, or a block diagram. Although a flowchart may describe operations as a sequential processing procedure, many operations may be performed in parallel or simultaneously. Additionally, the order of the operations may be rearranged. The processing procedure terminates when its operations are completed, but may have additional operations not included in the figure. The processing procedure may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When the processing procedure corresponds to a function, its termination may correspond to the function returning to the calling function or the main function.

[0070] The systems depicted in some of the figures may be provided in various configurations. In some embodiments, the system may be configured as a distributed system, where one or more components of the system are distributed across one or more networks in a cloud computing system.

[0071] Figure 1 FIG. is a block diagram illustrating the hardware components of a data transmission network 100 according to an embodiment of the present technology. The data transmission network 100 is a dedicated computer system that can be used to process a large amount of data (where a large number of computer processing cycles are required).

[0072] The data transmission network 100 may also include a computing environment 114. The computing environment 114 may be a dedicated computer or other machine for processing data received within the data transmission network 100. The data transmission network 100 also includes one or more network devices 102. The network devices 102 may include client devices that attempt to communicate with the computing environment 114. For example, the network devices 102 may send data to the computing environment 114 for processing, may send signals to the computing environment 114 to control different aspects of the computing environment or the data it is processing, and for other reasons. The network devices 102 may interact with the computing environment 114 in several ways (for example, via one or more networks 108). As Figure 1 shown, the computing environment 114 may include one or more other systems. For example, the computing environment 114 may include a database system 118 and / or a communication grid 120.

[0073] In other embodiments, the network devices may provide a large amount of data to the computing environment 114 via the network 108, either all at once or streamed over a period of time (for example, using event stream processing (ESP), regarding Figures 8 to 10Further description). For example, the network device 102 may include a network computer, a sensor, a database, or other device that can transmit or otherwise provide data to the computing environment 114. For example, the network device may include a local area network device such as a router, a hub, a switch, or other computer network linking device. These devices may provide a variety of stored or generated data, such as network data or data specific to the network device itself. The network device may also include sensors that monitor its environment or other devices to collect data about the environment or the device, and such network devices may provide the data they collect over time. The network device may also include devices within the Internet of Things, such as devices within a home automation network. Some of these devices may be referred to as edge devices and may be involved in edge computing circuitry. Data may be transmitted directly by the network device to the computing environment 114 or to a network-attached data repository (such as the network-attached data repository 110) for storage, such that the data may be retrieved later by the computing environment 114 or other parts of the data transmission network 100.

[0074] The data transmission network 100 may also include one or more network-attached data repositories 110. The network-attached data repositories 110 are used to store data to be processed by the computing environment 114 and any intermediate or final data generated by the computing system in non-volatile memory. However, in a particular embodiment, the computing environment 114 is configured to allow its operations to be performed such that intermediate and final data results can be stored only in volatile memory (e.g., RAM), without the need to store the intermediate or final data results in a non-volatile type of memory (e.g., disk). This may be useful in certain situations, such as when the computing environment 114 receives an ad hoc query from a user and when a response generated by processing a large amount of data needs to be generated in real time. In this non-limiting scenario, the computing environment 114 may be configured to retain the processed information in memory such that responses can be generated for the user at different levels of detail and the user can interactively query this information.

[0075] A network-attached data repository can store multiple different types of data organized in multiple different ways and from multiple different sources. For example, a network-attached data repository can include memory in addition to the main memory located within computing environment 114 that is directly accessible by a processor located therein. The network-attached data repository can include secondary, tertiary, or auxiliary memory such as large drives, servers, virtual memory, and other types. Storage devices can include portable or non-portable storage devices, optical storage devices, and various other media capable of storing and containing data. A machine-readable storage medium or a computer-readable storage medium can include a non-transitory medium in which data can be stored and that does not include a carrier wave and / or transient electronic signals. Examples of non-transitory media can include, for example, magnetic disks or tapes, optical storage media (such as optical discs or digital versatile discs), flash memory, memory, or memory devices. A computer program product can include code and / or machine-executable instructions that can represent any combination of a process, function, subroutine, program, routine, subroutine, module, software package, class, or instructions, data structures, or program statements. A code segment can be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. can be passed, transferred, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, etc. Additionally, the data repository can hold multiple different types of data. For example, network-attached data repository 110 can hold unstructured (e.g., raw) data such as manufacturing data (e.g., a database containing records of products manufactured with parameter data identifying each product with, for example, color and model) or a product sales database (e.g., a database containing individual data records identifying details of individual product sales).

[0076] Unstructured data can be presented to computing environment 114 in different forms (such as flat files or data record aggregations) and can have data values and accompanying timestamps. Computing environment 114 can be used to analyze unstructured data in multiple ways to determine the best way to (e.g., hierarchically) structure the data such that the structured data is customized for the type of further analysis the user wishes to perform on the data. For example, after being processed, unstructured timestamp data can be summarized by time (e.g., into daily time period units) to produce time series data, and / or structured hierarchically according to one or more dimensions (such as parameters, attributes, and / or variables). For example, the data can be stored in a hierarchical data structure (such as a ROLAP or MOLAP database), or can be stored in another tabular form (such as in a flat hierarchical form).

[0077] The data transmission network 100 may also include one or more server farms 106. The computing environment 114 may select communications or route data to one or more server farms 106 or to one or more servers within the server farm. The server farm 106 may be configured to provide information in a predetermined manner. For example, the server farm 106 may access data for transmission in response to a communication. The server farm 106 may be separately housed from every other device (such as the computing environment 114) within the data transmission network 100, and / or may be part of a device or system.

[0078] The server farm 106 may host a variety of different types of data processing as part of the data transmission network 100. The server farm 106 may receive a variety of different data from network devices, the computing environment 114, the cloud network 116, or other sources. The data may have been obtained or collected as input from one or more sensors from a control database, or may have been received as input from an external system or device. The server farm 106 may assist in processing the data by transforming the raw data into processed data based on one or more rules implemented by the server farm. For example, sensor data may be analyzed to determine changes over time or in real time in the environment.

[0079] The data transmission network 100 may also include one or more cloud networks 116. The cloud network 116 may include a cloud infrastructure system that provides cloud services. In certain embodiments, the services provided by the cloud network 116 may include numerous services that are available on demand to users of the cloud infrastructure system. The cloud network 116 is Figure 1 shown as being connected to the computing environment 114 (and thus having the computing environment 114 as its client or user), but the cloud network 116 may be connected to Figure 1 any of the devices in Figure 1 or utilized by any of the devices in

[0080] Although Figure 1 each device, server, and system in

[0081] Each communication within the data transfer network 100 (e.g., between client devices, between a device and a connection management system, between the server 106 and the computing environment 114, or between a server and a device) can occur via one or more networks 108. The network 108 can include one or more of a variety of different types of networks, including wireless networks, wired networks, or a combination of wired and wireless networks. Examples of suitable networks include the Internet, a personal area network, a local area network (LAN), a wide area network (WAN), or a wireless local area network (WLAN). The wireless network can include a wireless interface or a combination of wireless interfaces. As an example, the network in one or more networks 108 can include a short-range communication channel (e.g., a Bluetooth or Bluetooth Low Energy channel). The wired network can include a wired interface. The wired and / or wireless network can be implemented using routers, access points, bridges, gateways, or the like to connect the devices in the network 108, as will be further described with respect to Figure 2 The one or more networks 108 can be entirely incorporated within an internal network, an inter-business network, or a combination thereof, or can include an internal network, an inter-business network, or a combination thereof. In one embodiment, communication between two or more systems and / or devices can be implemented by a secure communication protocol (e.g., Secure Sockets Layer (SSL) or Transport Layer Security (TLS)). Additionally, data and / or transaction details can be encrypted.

[0082] Some aspects can utilize the Internet of Things (IoT), where things (e.g., machines, devices, phones, sensors) can be connected to a network and data can be collected and processed within and / or outside of the things. For example, the IoT can include sensors in many different devices, and high-value analytics can be applied to identify hidden relationships and drive increased efficiency. This can apply to both big data analytics and real-time (e.g., ESP) analytics. The IoT can be implemented in various domains, such as for access (techniques for obtaining and moving data), embedded capabilities (devices with embedded sensors), and services. Industries within the IoT space can include automotive (connected cars), manufacturing (connected factories), smart cities, energy, and retail. This will be further described below with respect to Figure 2 Further description.

[0083] As mentioned, the computing environment 114 can include a communication grid 120 and a transport network database system 118. The communication grid 120 can be a grid-based computing system for processing large amounts of data. The transport network database system 118 can be used to manage, store, and retrieve large amounts of data assigned to and stored in one or more network-attached data repositories 110 or other data repositories located at different locations within the transport network database system 118. The computing nodes in the grid-based computing system 120 and the transport network database system 118 can share the same processor hardware, such as a processor located within the computing environment 114.

[0084] Figure 2 Illustrate an example network including a set of example devices that communicate with each other via a switching system and via a network according to an embodiment of the present technology. As mentioned, each communication within the data transmission network 100 can occur via one or more networks. The system 200 includes a network device 204 configured to communicate with various types of client devices (e.g., client device 230) via various types of communication channels.

[0085] As Figure 2 shown, the network device 204 can transmit communications via a network (e.g., a cellular network via a base station 210). The communication can be routed via the base station 210 to another network device, such as network devices 205 to 209. The communication can also be routed via the base station 210 to the computing environment 214. For example, the network device 204 can collect data from its surrounding environment or from other network devices (e.g., network devices 205 to 209) and transmit the data to the computing environment 214.

[0086] Although network devices 204 to 209 are Figure 2 shown as a mobile phone, a laptop computer, a tablet computer, a temperature sensor, a motion sensor, and an audio sensor, respectively, the network device can be or include a sensor sensitive to aspects of its environment. For example, the network device can particularly include sensors such as water sensors, power sensors, current sensors, chemical sensors, optical sensors, pressure sensors, geographic or position sensors (e.g., GPS), speed sensors, acceleration sensors, flow sensors, etc. Examples of characteristics that can be sensed particularly include force, torque, load, strain, position, temperature, air pressure, fluid flow, chemical properties, resistance, electromagnetic fields, radiation, irradiance, proximity, acoustics, humidity, distance, speed, vibration, acceleration, electric potential, current. The sensors can be mounted on various components that are part of various different types of systems (e.g., drilling operations). The network device can detect and record data related to the environment it monitors and transmit the data to the computing environment 214.

[0087] As mentioned, according to a particular embodiment, one type of system that can include various sensors that collect data to be processed and / or transmitted to a computing environment includes an oil drilling system. For example, one or more drilling operation sensors can include surface sensors that measure hook load, fluid rate, temperature and density inside and outside the wellbore, standpipe pressure, surface torque, rotary speed of the drill pipe, penetration rate, mechanical specific energy, etc., and downhole sensors that measure rotary speed of the bit, fluid density, downhole torque, downhole vibration (axial, tangential, lateral), weight applied at the bit, annulus pressure, differential pressure, azimuth, inclination, dog leg severity, measured depth, vertical depth, downhole temperature, etc. In addition to the raw data directly collected by the sensors, other data can include parameters developed by the sensors or assigned to the system by a client or other control device. For example, one or more drilling operation control parameters can control settings such as mud motor speed to flow rate ratio, bit diameter, predicted formation top, seismic data, weather data, etc. Other data can be generated using physical models such as earth models, weather models, seismic models, bottom hole assembly models, well planning models, annulus friction models, etc. In addition to sensors and control settings, predicted outputs such as penetration rate, mechanical specific energy, hook load, inflow fluid rate, outflow fluid rate, pump pressure, surface torque, rotary speed of the drill pipe, annulus pressure, annulus friction pressure, annulus temperature, circulating equivalent density, etc. can also be stored in a data warehouse.

[0088] In another example, according to certain embodiments, another type of system that can include various sensors that collect data to be processed and / or transmitted to a computing environment includes a home automation or similar automation network in different environments (such as office spaces, schools, public spaces, sports arenas, or a variety of other locations). Network devices in this automation network can include network devices that allow a user to access, control, and / or configure various household appliances located within the user's home (such as, for example, a television, radio, lights, fans, humidifiers, sensors, microwave ovens, irons, and / or the like) or outside the user's home (such as, for example, an external motion sensor, external lighting, a garage door opener, a sprinkler system, or the like). For example, network device 102 can include a home automation switch that can be coupled to household appliances. In another embodiment, the network device can allow a user to access, control, and / or configure devices such as office-related devices (such as, for example, a copier, printer, or fax machine), audio and / or video-related devices (such as, for example, a receiver, speakers, a projector, a DVD player, or a television), media playback devices (such as, for example, a disc player, a CD player, or the like), computing devices (such as, for example, a home computer, a laptop computer, a tablet computer, a personal digital assistant (PDA), a computing device, or a wearable device), lighting devices (such as, for example, lights or recessed lighting), devices associated with a security system, devices associated with an alarm system, devices that can operate in an automobile (such as, for example, a radio device, a navigation device), and / or the like. Data can be collected in its raw form from such various sensors, or the sensors can process the data to establish parameters or other data developed by the sensors based on the raw data or assigned to the system by a client or other control device.

[0089] In another example, according to certain embodiments, another type of system that can include various sensors that collect data to be processed and / or transmitted to a computing environment includes a power or energy grid. A variety of different network devices can be included in the energy grid, particularly, for example, various devices within one or more power plants, energy farms (particularly, for example, wind farms, solar farms), energy storage facilities, factories, consumers' homes, and businesses. One or more of such devices can include one or more sensors that detect energy gain or loss, electrical input or output or loss, and a variety of other efficiencies. These sensors can collect data to inform a user how the energy grid and individual devices within the grid can operate and how they can be made more efficient.

[0090] The network device sensor can also process the data it collects before transmitting the data to the computing environment 114 or before deciding whether to transmit the data to the computing environment 114. For example, the network device can determine whether the collected data meets a specific rule by, for example, comparing the data or values calculated from the data and comparing the data with one or more thresholds. The network device can use this data and / or comparison to determine whether the data should be transmitted to the computing environment 214 for further use or processing.

[0091] The computing environment 214 can include machines 220 and 240. Although the computing environment 214 is shown in Figure 2 as having two machines 220 and 240, the computing environment 214 can have only one machine or can have more than two machines. The machines that make up the computing environment 214 can include dedicated computers, servers, or other machines configured to process large amounts of data individually and / or jointly. The computing environment 214 can also include a storage device containing one or more databases with structured data (e.g., data organized in one or more hierarchies) or unstructured data. The databases can communicate with the processing devices within the computing environment 214 to distribute data to the processing devices. Since the network device can transmit data to the computing environment 214, the data can be received by the computing environment 214 and then stored in the storage device. The data used by the computing environment 214 can also be stored in the data repository 235, which can also be part of or connected to the computing environment 214.

[0092] The computing environment 214 can communicate with various devices via one or more routers 225 or other inter-network or intra-network connection components. For example, the computing environment 214 can communicate with the device 230 via one or more routers 225. The computing environment 214 can collect, analyze, and / or store data stored at one or more data repositories 235 regarding communications, client device operations, client rules, and / or user-associated actions. This data can affect the communication routing within the computing environment 214, how data is stored or processed within the computing environment 214, and other actions.

[0093] Note that various other devices can be further used to affect the communication routing and / or processing between the devices within the computing environment 214 and with devices external to the computing environment 214. For example, as Figure 2 shown, the computing environment 214 can include a network server 240. Thus, the computing environment 214 can retrieve data of interest, such as client information (e.g., product information, client rules, etc.), technical product details, news, current or predicted weather, etc.

[0094] In addition to the computing environment 214 collecting data (e.g., received from network devices such as sensors and client devices or other sources) for processing as part of a big data analytics project, it can also receive data in real time as part of a streaming analytics environment. As mentioned, data can be collected using various sources such as via different types of networks or communicating locally. This data can be received on a real-time streaming basis. For example, a network device can periodically receive data from a network device sensor as the sensor continuously senses, monitors, and tracks changes in its environment. Devices within the computing environment 214 can also perform pre-analysis on the data they receive to determine whether the received data should be processed as part of an ongoing project. The data received and collected by the computing environment 214 (regardless of the source or method or timing of reception) can be processed for a client over a period of time to determine result data based on the client's requirements and rules.

[0095] Figure 3 A representation of a conceptual model of a communication protocol system according to an embodiment of the present technology is illustrated. More specifically, Figure 3 Identify the operations of a computing environment in the Open Systems Interconnection model corresponding to various connected components. For example, model 300 shows how a computing environment such as computing environment 320 (or Figure 2 the computing environment 214 in) can communicate with other devices in its network, and how and under what conditions the communication between the computing environment and other devices is controlled.

[0096] The model can include layers 302 to 314. The layers are arranged in a stack. Each layer in the stack serves the layer one level higher than it (except for the application layer, which is the highest layer), and is served by the layer one level below it (except for the physical layer, which is the lowest layer). The physical layer is the lowest layer because it receives and transmits raw data bytes and is the layer furthest from the user in the communication system. On the other hand, the application layer is the highest layer because it directly interacts with software applications.

[0097] As mentioned, the model includes the physical layer 302. The physical layer 302 represents physical communication and can define the parameters of such physical communication. For example, this physical communication can occur in the form of electrical, optical, or electromagnetic signals. The physical layer 302 also defines the protocols that can control communication within a data transmission network.

[0098] The link layer 304 defines the links and mechanisms for transmitting (i.e., moving) data across a network. The link layer manages communication between nodes within a grid computing environment, for example. The link layer 304 can detect and correct errors (e.g., transmission errors in the physical layer 302). The link layer 304 can also include a Media Access Control (MAC) layer and a Logical Link Control (LLC) layer.

[0099] The network layer 306 defines protocols for routing within a network. In other words, the network layer coordinates the transfer of data across nodes in the same network (for example, in a grid computing environment). The network layer 306 may also define processes for local addressing within the structured network.

[0100] The transport layer 308 may manage the transfer of data and the quality of the transfer and / or reception of the data. The transport layer 308 may provide a protocol for transferring data, for example, the Transmission Control Protocol (TCP). The transport layer 308 may assemble and disassemble data frames for transfer. The transport layer may also detect transfer errors occurring in the layer below it.

[0101] The session layer 310 may establish, maintain, and manage communication connections between devices on the network. In other words, the session layer controls the dialogue or nature of the communication between network devices on the network. The session layer may also establish checkpointing, adjournment, termination, and restart processes.

[0102] The presentation layer 312 may provide translation of the communication between the application and the network layer. In other words, this layer may encrypt, decrypt, and / or format data based on data types known to be accepted by the application or the network layer.

[0103] The application layer 314 interacts directly with software applications and end users and manages the communication between them. The application layer 314 may use the application to identify the destination, the status or availability of local resources, and / or the content or formatting of the communication.

[0104] The intra-network connection components 322 and 324 are shown to operate in lower levels such as the physical layer 302 and the link layer 304, respectively. For example, a hub may operate in the physical layer and a switch may operate in the link layer. The inter-network connection components 326 and 328 are shown to operate at higher levels (such as layers 306 to 314). For example, a router may operate in the network layer and a network device may operate in the transport, session, presentation, and application layers.

[0105] As mentioned, in various embodiments, the computing environment 320 may interact with and / or operate on one, more, all, or any of the various layers. For example, the computing environment 320 may interact with a hub (e.g., via the link layer) to adjust which devices the hub communicates with. The physical layer may be served by the link layer, and thus it may implement this data from the link layer. For example, the computing environment 320 may control which devices it will receive data from. For example, if the computing environment 320 knows that a particular network device has been turned off, damaged, or otherwise become unavailable or unreliable, then the computing environment 320 may instruct the hub to prevent any data from being transmitted from the network device to the computing environment 320. This process may be beneficial in avoiding receiving inaccurate data or data that has been affected by an uncontrolled environment. As another example, the computing environment 320 may communicate with a bridge, switch, router, or gateway and influence which device within the component selection system (e.g., system 200) is used as a destination. In some embodiments, the computing environment 320 may interact with the various layers by exchanging communications with the equipment operating on a particular layer by routing or modifying existing communications. In another embodiment (e.g., in a grid computing environment), nodes may determine how to route data within the environment (e.g., which node should receive particular data) based on specific parameters or information provided by other layers within the model).

[0106] As mentioned, the computing environment 320 may be part of a communication grid environment (the communications of which may be implemented as shown in the Figure 3 protocol). For example, referring back to Figure 2 , one or more of machines 220 and 240 may be part of a communication grid computing environment. A grid computing environment may be employed in a distributed system with a non-interactive workload where data resides in memory on machines or compute nodes. In this environment, the analysis code rather than a database management system controls the processing performed by the nodes. Data is co-located by pre-assigning it to grid nodes, and the analysis code on each node loads the local data into memory. Each node may be assigned a specific task, such as processing a part of a project, or organizing or controlling other nodes within the grid.

[0107] Figure 4 Illustrates a communication grid computing system 400 including multiple control and worker nodes in accordance with an embodiment of the present technology. The communication grid computing system 400 includes three control nodes and one or more worker nodes. The communication grid computing system 400 includes control nodes 402, 404, and 406. The control nodes are communicatively connected via communication paths 451, 453, and 455. Thus, the control nodes may transmit (e.g., information related to the communication grid or notifications) and receive information from each other. Although the communication grid computing system 400 is shown in Figure 4 as including three control nodes, the communication grid may include more or fewer than three control nodes.

[0108] The communication grid computing system (or simply "communication grid") 400 also includes one or more worker nodes. Figure 4 Six worker nodes 410 to 420 are shown therein. Although Figure 4 six worker nodes are shown, a communication grid according to embodiments of the present technology may include more or fewer than six worker nodes. The number of worker nodes included in the communication grid may depend in particular on how large the project or data set that the communication network is processing is, the capacity of each worker node, and the time assigned to the communication grid to complete the project. Each worker node within the communication grid 400 may be (wired or wirelessly, and directly or indirectly) connected to the control nodes 402 to 406. Thus, each worker node may receive information (e.g., instructions for performing work on a project) from the control nodes and may transmit information (e.g., results from the work performed on a project) to the control nodes. Additionally, the worker nodes may communicate with each other (directly or indirectly). For example, the worker nodes may transmit data related to the jobs being performed or individual tasks within the jobs being performed by the worker nodes to each other. However, in certain embodiments, a worker node may not be connected (communicatively or otherwise) to some other worker nodes. In an embodiment, a worker node may only be able to communicate with the control node that controls it and may not be able to communicate with other worker nodes in the communication grid, whether they are other worker nodes controlled by the control node that controls the worker node or worker nodes controlled by other control nodes in the communication grid.

[0109] The control nodes may be connected to external devices and may communicate with the external devices (for example, a grid user such as a server or a computer may be connected to a controller of the grid). For example, a server or a computer may be connected to the control nodes and may transmit a project or a job to the nodes. A project may include a data set. The data set may be of any size. Once a control node receives such a project that includes a large data set, the control node may allocate the data set or the project related to the data set for execution by the worker nodes. Alternatively, for a project that includes a large data set, the data set may be received or stored by a machine other than the control node (e.g., a Hadoop data node).

[0110] The control nodes may maintain knowledge of the states of the nodes in the grid (i.e., grid state information), accept work requests from clients, subdivide work across the worker nodes, coordinate the worker nodes, and other responsibilities. The worker nodes may accept work requests from the control nodes and provide the results of the work performed by the worker nodes to the control nodes. The grid may be started from a single node (e.g., a machine, a computer, a server, etc.). This first node may be assigned or may be started as the master control node that will control any additional nodes entering the grid.

[0111] When a project is submitted for execution (e.g., by a client or a controller of a grid), it can be assigned to a set of nodes. After the nodes are assigned to the project, a data structure (i.e., communicator) can be established. The communicator can be used by the project to share information among the project code running on each node. A communication handle can be established on each node. For example, the handle is a reference to the communicator that is valid within a single process on a single node, and the handle can be used when requesting communication between nodes.

[0112] A control node (e.g., control node 402) can be designated as the primary control node. A server, computer, or other external device can be connected to the primary control node. Once the control node receives a project, the primary control node can allocate portions of the project to its worker nodes for execution. For example, when a project is initiated on communication grid 400, the primary control node 402 controls the work to be performed for the project to complete the project according to a request or instruction. The primary control node can allocate work to the worker nodes based on various factors, such as which subsets or portions of the project can be completed most efficiently and within the correct amount of time. For example, a worker node can perform an analysis on a portion of the data that is already local to the worker node (e.g., stored on the worker node). After each worker node executes and completes its job, the primary control node also coordinates and processes the results of the work performed by each worker node. For example, the primary control node can receive results from one or more worker nodes, and the control node can organize (e.g., collect and assemble) the received results and compile them to produce the complete result of the project received from the end user.

[0113] Any remaining control nodes (e.g., control nodes 404 and 406) can be assigned as backup control nodes for the project. In an embodiment, the backup control nodes may not control any portion of the project. Instead, the backup control nodes can serve as a backup to the primary control node and take over as the primary control node if the primary control node fails. If the communication grid contains only a single control node and the control node fails (e.g., the control node is shut down or interrupted), then the entire communication grid may fail and any project or job running on the communication grid may fail and may not be completed. Although the project can be run again, this failure can cause a delay in the completion of the project (in some cases a severe delay, e.g., an all-night delay). Thus, a grid with multiple control nodes (including backup control nodes) can be beneficial.

[0114] For example, to add another node or machine to the grid, the master control node may open a pair of listening sockets. The sockets can be used to accept work requests from clients, and the second socket can be used to accept connections from other grid nodes. The master control node may be provided with a list of other nodes (e.g., other machines, computers, servers) that will participate in the grid, and the role that each node will play in filling out the grid. When the master control node (e.g., the first node on the grid) starts up, the master control node may use a network protocol to start a server process on each of the other nodes in the grid. For example, command line arguments may tell each node one or more pieces of information, particularly for example: the role the node will have in the grid, the hostname of the master control node, the port number on which the master control node accepts connections from peer nodes. The information may also be provided in a configuration file, transmitted via a secure shell tunnel, retrieved from a configuration server. Although the other machines in the grid may not initially know the configuration of the grid, the information may also be sent by the master control node to each of the other nodes. Updates to the grid information may also be sent to the nodes subsequently.

[0115] For any control node other than the master control node added to the grid, the control node may open three sockets. The first socket may accept work requests from clients, the second socket may accept connections from other grid members, and the third socket may (e.g., permanently) connect to the master control node. When a control node (e.g., the master control node) receives a connection from another control node, it first checks whether the peer node is in the list of configured nodes in the grid. If it is not on the list, then the control node may drop the connection. If it is on the list, then it may attempt to authenticate the connection. If the authentication is successful, then the authenticating node may transmit information to its peer, such as the port number on which the node listens for connections, the hostname of the node, information on how to authenticate the node, and other information. When a node (e.g., a new control node) receives information about another node in operation, it will check whether it already has a connection to the other node. If it does not have a connection to the node, then it may establish a connection to the control node.

[0116] Any worker node added to the grid may establish connections to the master control node and to any other control node on the grid. After establishing the connections, it may authenticate itself to the grid (e.g., any control node, including both master and backup, or the server or user controlling the grid). After successful authentication, the worker node may accept configuration information from the control node.

[0117] When a node joins the communication grid (e.g., when the node powers on or connects to an existing node on the grid, or both), the node is assigned (e.g., by the operating system of the grid) a Universally Unique Identifier (UUID). This unique identifier helps other nodes and external entities (devices, users, etc.) identify the node and distinguish it from other nodes. When a node connects to the grid, the node can share its unique identifier with other nodes in the grid. Since each node can share its unique identifier, each node can know the unique identifier of every other node on the grid. The unique identifier can also specify the hierarchy of each of the nodes (e.g., backup control nodes) within the grid. For example, the unique identifier of each of the backup control nodes can be stored in a list of backup control nodes to indicate the order in which the backup control nodes will take over from a failed primary control node to become the new primary control node. However, methods other than using the unique identifier of the node can also be used to determine the hierarchy of the node. For example, the hierarchy can be predetermined or assigned based on other predetermined factors.

[0118] The grid can add new machines at any time (e.g., starting from any control node). When adding a new node to the grid, the control node can first add the new node to its grid node table. Then, the control node can also notify every other control node about the new node. The nodes that receive the notification can acknowledge that they have updated their configuration information.

[0119] For example, the primary control node 402 can transmit one or more communications to the backup control nodes 404 and 406 (and, for example, to other control or worker nodes within the communication grid). Among other protocols, such communications can be sent periodically at fixed time intervals between known fixed phases of project execution. The communications transmitted by the primary control node 402 can have varying types and can contain various types of information. For example, the primary control node 402 can transmit a snapshot (e.g., status information) of the communication grid so that the backup control node 404 always has the most recent snapshot of the communication grid. For example, the snapshot or grid status can include the structure of the grid (e.g., including the worker nodes in the grid, the unique identifiers of the nodes, or their relationship to the primary control node) and the status of the project (e.g., including the status of the project portion of each worker node). The snapshot can also include analyses or results received from the worker nodes in the communication grid. The backup control node can receive and store the backup data received from the primary control node. The backup control node can transmit a request for this snapshot (or other information) from the primary control node, or the primary control node can send this information periodically to the backup control node.

[0120] As mentioned, if the primary control node fails, then the backup data allows the backup control node to take over as the primary control node without the grid having to start the project from scratch. If the primary control node fails, then the backup control node that will take over as the primary control node can retrieve the most recent version of the snapshot received from the primary control node and use the snapshot to continue the project from the stage of the project indicated by the backup data. This can prevent the overall failure of the project.

[0121] The backup control node can use various methods to determine that the primary control node has failed. In one example of such a method, the primary control node can transmit a communication (e.g., a heartbeat communication) indicating that the primary control node is working and has not failed to the backup control node (e.g., periodically). If the backup control node does not receive a heartbeat communication within a particular predetermined time period, then the backup control node can determine that the primary control node has failed. Alternatively, the backup control node can also receive a communication that the primary control node has failed from the primary control node itself (before it fails) or from a working node (e.g., because the primary control node fails to communicate with the working node).

[0122] Different methods can be implemented to determine which backup control node of a set of backup control nodes (e.g., backup control nodes 404 and 406) will take over the failed primary control node 402 and become the new primary control node. For example, the new primary control node can be selected based on the unique identifier of the backup control node based on its ranking or "hierarchy". In an alternative embodiment, the backup control node can be assigned as the new primary control node by another device in the communication grid or from an external device (e.g., the system infrastructure that controls the communication grid or an end user, such as a server or a computer). In another alternative embodiment, the backup control node that takes over as the new primary control node can be designated based on bandwidth or other statistical data regarding the communication grid.

[0123] Working nodes within the communication grid can also fail. If a working node fails, then the work performed by the failed working node can be redistributed among the operational working nodes. In an alternative embodiment, the primary control node can transmit a communication to each of the operational working nodes still on the communication grid that each of the working nodes should also intentionally fail. After each of the working nodes fails, it can retrieve its most recently saved checkpoint of its state accordingly and restart the project from the checkpoint to minimize the loss of progress of the project being performed.

[0124] Figure 5Figure 500 is a flowchart illustrating an example processing procedure for adjusting a communication grid or work items in a communication grid after a failure of a node. For example, the processing procedure may include receiving grid status information including a project status of a part of an item executed by a node in the communication grid, as described in operation 502. For example, a control node (e.g., a backup control node connected to a primary control node and work nodes on the communication grid) may receive the grid status information, where the grid status information includes the project status of the primary control node or the project status of the work nodes. The project status of the primary control node and the project status of the work nodes may include the status of one or more parts of the items executed by the primary and work nodes in the communication grid. The processing procedure may also include storing the grid status information, as described in operation 504. For example, the control node (e.g., the backup control node) may locally store the received grid status information within the control node. Alternatively, the grid status information may be sent to another device for storage, and the control node may access the information at the other device.

[0125] The processing procedure may also include receiving a failure communication corresponding to a node in the communication grid in operation 506. For example, a node may receive a failure communication including an indication that the primary control node has failed, to prompt the backup control node to take over from the primary control node. In an alternative embodiment, a node may receive a failure that a work node has failed, to prompt the control node to reassign the work performed by the work node. The processing procedure may also include reassigning a node or a part of an item executed by the failed node, as described in operation 508. For example, the control node may designate the backup control node as the new primary control node based on the failure communication upon receiving the failure communication. If the failed node is a work node, the control node may use a snapshot of the communication grid to identify the project status of the failed work node, where the project status of the failed work node includes the status of a part of the item executed by the failed work node at the time of failure.

[0126] The processing procedure may also include receiving updated grid status information based on the reassigning, as described in operation 510, and transmitting an instruction set to one or more nodes in the communication grid based on the updated grid status information, as described in operation 512. The updated grid status information may include the updated project status of the primary control node or the updated project status of the work nodes. The updated information may be transmitted to other nodes in the grid to update their stale stored information.

[0127] Figure 6Illustrates a portion of a communication grid computing system 600 including a control node and worker nodes according to an embodiment of the present technology. For illustrative purposes, the communication grid 600 computing system includes one control node (control node 602) and one worker node (worker node 610), but may include more worker and / or control nodes. The control node 602 is communicatively coupled to the worker node 610 via a communication path 650. Thus, the control node 602 can transmit (e.g., information related to the communication grid or notifications) to and receive information from the worker node 610 via path 650.

[0128] Similar to in Figure 4 the communication grid computing system (or simply "communication grid") 600 includes data processing nodes (control node 602 and worker nodes 610). Nodes 602 and 610 include multi-core data processors. Each node 602 and 610 includes a grid-enabled software component (GESC) 620 that executes on the data processor associated with the node and interfaces with a buffer memory 622 also associated with the node. Each node 602 and 610 includes database management software (DBMS) 628 that executes on a database server (not shown) at the control node 602 and on a database server (not shown) at the worker node 610.

[0129] Each node also includes a data repository 624. Similar to Figure 1 the network-attached data repository 110 in Figure 2 and the data repository 235 in

[0130] the data repository 624 is used to store data to be processed by the nodes in the computing environment. The data repository 624 can also store any intermediate or final data generated by the computing system after processing, for example, in non-volatile memory. However, in certain embodiments, the configuration of the grid computing environment allows its operation to be performed such that intermediate and final data results can be stored separately in volatile memory (e.g., RAM) without the need to store the intermediate or final data results to non-volatile type memory. Storing this data in volatile memory can be useful in certain situations, such as when the grid receives a query (e.g., specific) from a client and when a response generated by processing large amounts of data needs to be produced quickly or in real-time. In this case, the grid can be configured to retain the data in memory such that responses can be generated at different levels of detail and such that the client can interactively query this information.Each node also includes a user-defined function (UDF) 626. The UDF provides a mechanism for the DBMS 628 to transfer data to or receive data from a database stored in the data repository 624 managed by the DBMS. For example, the UDF 626 can be called by the DBMS to provide data to the GESC for processing. The UDF 626 can establish a socket connection (not shown) with the GESC to transfer data. Alternatively, the UDF 626 can transfer data to the GESC by writing the data to a shared memory that can be accessed by both the UDF and the GESC.

[0131] The GESCs 620 at nodes 602 and 610 can be connected via a network (such as Figure 1 the network 108 shown in). Thus, nodes 602 and 610 can communicate with each other via the network using a predetermined communication protocol (for example, such as the Message Passing Interface (MPI)). Each GESC 620 can participate in point-to-point communication with a GESC at another node, or communicate collectively with multiple GESCs via the network. The GESCs 620 at each node can contain the same (or nearly the same) software instructions. Each node may be capable of operating as a control node or a worker node. The GESC at the control node 602 can communicate with the client device 630 via the communication path 652. More specifically, the control node 602 can communicate with the client application 632 hosted by the client device 630 to receive queries and respond to the queries after processing a large amount of data.

[0132] The DBMS 628 can control the establishment, maintenance, and use of databases or data structures (not shown) within nodes 602 or 610. The database can organize the data stored in the data repository 624. The DBMS 628 at the control node 602 can accept requests for data and transfer appropriate data for the requests. Through this process, a data set can be distributed across multiple physical locations. In this example, each of nodes 602 and 610 stores a portion of the total data managed by the management system in its associated data storage device 624.

[0133] In addition, the DBMS can be responsible for protecting against data loss using replication techniques. Replication includes providing a backup copy of the data on one node stored on one or more other nodes. Thus, if a node fails, the data from the failed node can be recovered from the replicated copy residing on another node. However, as described herein with respect to Figure 4 each node's data or status information in the communication grid can also be shared with every node on the grid.

[0134] Figure 7The flowchart 700 illustrates an example method for executing a project within a grid computing system according to an embodiment of the present technology. As described with respect to Figure 6 The GESC at the control node can transmit data with a client device (e.g., client device 630) to receive a query for executing a project and respond to the query after a large amount of data has been processed. The query can be transmitted to the control node, where the query can include a request to execute a project, as described in operation 702. The query can contain instructions regarding the type of data analysis to be performed in the project and whether a grid-based computing environment should be used to execute the project, as shown in operation 704.

[0135] To initiate a project, the control node can determine whether the query requests the use of a grid-based computing environment to execute the project. If the determination is no, then the control node initiates the execution of the project in a separate environment (e.g., at the control node), as described in operation 710. If the determination is yes, then the control node can initiate the execution of the project in a grid-based computing environment, as described in operation 706. In this case, the request can include the requested configuration of the grid. For example, the request can include the number of control nodes and the number of worker nodes to be used in the grid when executing the project. After the project has been completed, the control node can transmit the analysis results generated by the grid, as described in operation 708. Regardless of whether the project is executed in a separate or grid-based environment, the control node provides the results of the project in operation 712.

[0136] As described with respect to Figure 2 The computing environment described herein can collect data (e.g., received from network devices such as sensors (e.g., Figure 2 network devices 204 to 209 in Figure 2 ), client devices, or other sources) as part of a data analysis project and can receive data in real time as part of a streaming analysis environment (e.g., ESP). Data can be collected using multiple sources such as via different types of networks or locally communicated, e.g., on a real-time streaming basis. For example, a network device can periodically receive data from a network device sensor as the sensor continuously senses, monitors, and tracks changes in its environment. More specifically, an increasing number of distributed applications develop or generate continuous flowing data from distributed sources by applying queries to the data before distributing the data to geographically distributed receivers. An event stream processing engine (ESPE) can continuously apply queries to the data when the data is received and determine which entities should receive the data. A client or other device can also subscribe to the ESPE or other devices that process ESP data so that it can receive the data after processing (e.g., based on the entities determined by the processing engine). For example, Figure 10The further described event subscription devices 1024a to 1024c may also subscribe to the ESPE. The ESPE may determine or define how input data or event streams from network devices or other publishers (e.g., Figure 2 the network devices 204 to 209 in Figure 2 are transformed into meaningful output data for consumption by subscribers (e.g., the client device 230 in

[0137] Figure 8 Block diagram showing components of an event stream processing engine (ESPE) according to an embodiment of the present technology. The ESPE 800 may include one or more items 802. An item may be described as a second-level container in an engine model managed by the ESPE 800, where the thread pool size of an item may be user-defined. Each of the one or more items 802 may include one or more continuous queries 804 containing a data stream (which is a data transformation of an incoming event stream). The one or more continuous queries 804 may include one or more source windows 806 and one or more derived windows 808.

[0138] The ESPE may receive streaming data related to a particular event over a period of time, such as an event or other data sensed by one or more network devices. The ESPE may perform operations associated with processing data established by one or more devices. For example, the ESPE may receive data from Figure 2 one or more of the network devices 204 to 209 shown in Figure 2 As mentioned, the network devices may include sensors that sense different aspects of their environment and may collect data over time based on the sensed observations. For example, the ESPE may be implemented within one or more of the machines 220 and 240 shown in

[0139] The engine container is the top-level container in the model that manages the resources of the one or more items 802. In an illustrative embodiment, for example, for each instance of an ESP application, there may be only one ESPE 800, and the ESPE 800 may have a unique engine name. Additionally, each of the one or more items 802 may have a unique item name, and each query may have a unique continuous query name and start with a uniquely named source window of one or more source windows 806. The ESPE 800 may or may not be persistent.

[0140] Continuous query modeling involves manipulating and transforming a defined window directed graph for an event stream. In the context of event stream manipulation and transformation, a window is a processing node in an event stream processing model. Windows in a continuous query can perform aggregation, calculation, pattern matching, and other operations on the data flowing through the window. A continuous query can be described as a directed graph of source, relational, pattern matching, and process windows. One or more source windows 806 and one or more derived windows 808 represent a query that produces updates to a query result set as a continuous execution of a new event block stream through the ESPE 800. For example, a directed graph is a set of nodes connected by edges, where the edges have a direction associated with them.

[0141] An event object can be described as a data grouping that can be accessed as a set of fields, where at least one of the fields is defined as a keyword or a unique identifier (ID). Event objects can be created using a variety of formats including binary, alphanumeric, XML, etc. Each event object can contain one or more fields designated as the primary identifier (ID) of the event, so that the ESPE 800 can support operation codes (opcodes) for events including insert, update, upsert, and delete. If the key field already exists, then the upsert opcode updates the event; otherwise, inserts the event. For illustration, an event object can be a compact binary representation of a set of field values and contain both metadata and field data associated with the event. The metadata can include an opcode indicating whether the event represents an insert, update, delete, or upsert, a set of flags indicating whether the event is a normal, partial update, or a retained event generated from retention policy management, and a set of microsecond timestamps that can be used for latency measurement.

[0142] An event block object can be described as a group or encapsulation of event objects. An event stream can be described as a stream of event block objects. The continuous queries 804 of one or more continuous queries transform a source event stream consisting of streamed event block objects published into the ESPE 800 using one or more source windows 806 and one or more derived windows 808 into one or more output event streams. A continuous query can also be regarded as data stream modeling.

[0143] One or more source windows 806 are at the top of the directed graph and have no windows fed into them. Event streams are published into the one or more source windows 806, and the event streams can be directed from there to the next set of connected windows as defined by the directed graph. One or more sink windows 808 are all instantiated windows that are not source windows and have other windows streaming events into them. The one or more sink windows 808 can perform computations or transformations on the incoming event streams. The one or more sink windows 808 transform the event streams based on the window type (i.e., operators such as, for example, join, filter, compute, aggregate, copy, pattern match, process, merge, etc.) and window settings. When an event stream is published into the ESPE 800, it is continuously queried, and the resulting set of sink windows in these queries is continuously updated.

[0144] Figure 9 The illustration presents a flowchart showing an example processing procedure that includes operations performed by an event stream processing engine according to some embodiments of the present technology. As mentioned, the ESPE 800 (or an associated ESP application) defines how to transform an input event stream into a meaningful output event stream. More specifically, the ESP application can define how to transform an input event stream from a publisher (e.g., a network device providing sensed data) into a meaningful output event stream consumed by a subscriber (e.g., a data analysis project performed by a machine or a group of machines).

[0145] Within the application, a user can interact with one or more user interface windows presented to the user in a display in a user-selectable order, either independently or via a browser application, under the control of the ESPE. For example, the user can execute the ESP application, which causes the presentation of a first user interface window that can include a plurality of menus and selectors associated with the ESP application (e.g., dropdown menus, buttons, text boxes, hyperlinks, etc.), as understood by those skilled in the art. As further understood by those skilled in the art, various operations can be performed in parallel, for example, using multiple threads.

[0146] In operation 900, the ESP application can define and start the ESPE, thereby instantiating the ESPE at the instantiation device (e.g., machines 220 and / or 240). In operation 902, an engine container is established. For illustration, the ESPE 800 can be instantiated using a function call that designates the engine container as a model manager.

[0147] In operation 904, one or more continuous queries 804 are instantiated as models by the ESPE 800. The one or more continuous queries 804 can be instantiated with a set of dedicated threads that produce updates as a new event stream through the ESPE 800. By way of illustration, one or more continuous queries 804 can be established to model business processing logic within the ESPE 800, predict events within the ESPE 800, model physical systems within the ESPE 800, predict the state of physical systems within the ESPE 800, and so on. For example, as mentioned, the ESPE 800 can be used to support sensor data monitoring and management (e.g., sensing can include force, torque, load, strain, position, temperature, air pressure, fluid flow, chemical properties, electrical resistance, electromagnetic fields, radiation, irradiance, proximity, acoustics, humidity, distance, speed, vibration, acceleration, electric potential, or current, etc.).

[0148] The ESPE 800 can analyze and process events or an "event stream" in motion. Instead of storing data and running queries against the stored data, the ESPE 800 can store queries and stream data through it to allow continuous analysis of the data as it is received. One or more source windows 806 and one or more derived windows 808 can be established based on relationships, pattern matching, and process algorithms that transform an input event stream into an output event stream to model, simulate, score, test, predict, etc. based on the defined continuous query model and the application of the streamed data.

[0149] In operation 906, the publish / subscribe (pub / sub) capability is initialized for the ESPE 800. In an illustrative embodiment, the pub / sub capability is initialized for each of the one or more items 802. To initialize and enable the pub / sub capability for the ESPE 800, a port number can be provided. The pub / sub client can use the hostname and port number of the ESP device running the ESPE to establish a pub / sub connection to the ESPE 800.

[0150] Figure 10Describe an ESP system 1000 docked between a publishing device 1022 and event subscription devices 1024a to 1024c according to an embodiment of the present technology. The ESP system 1000 may include an ESP device or subsystem 1001, an event publishing device 1022, an event subscription device A 1024a, an event subscription device B 1024b, and an event subscription device C 1024c. An input event stream is output to the ESP device 1001 through the publishing device 1022. In an alternative embodiment, the input event stream may be established through multiple publishing devices. The multiple publishing devices may further publish the event stream to other ESP devices. One or more continuous queries implemented by the ESPE 800 may analyze and process the input event stream to form an output event stream output to the event subscription device A 1024a, the event subscription device B 1024b, and the event subscription device C 1024c. The ESP system 1000 may include more or fewer event subscription devices of the event subscription devices.

[0151] Publish-subscribe is an indirect-addressing-based message-oriented interaction paradigm. The processed data receiver specifies its interest in receiving information from the ESPE 800 by subscribing to specific categories of events, while the information source publishes events to the ESPE 800 without directly addressing the receiver. The ESPE 800 coordinates the interaction and processes the data. In some cases, the data source receives confirmation that the data receiver has received the published information.

[0152] The publish / subscribe API can be described as a link library that enables an event publisher (e.g., the publishing device 1022) to publish an event stream to the ESPE 800 or an event subscriber (e.g., the event subscription device A 1024a, the event subscription device B 1024b, and the event subscription device C 1024c) to subscribe to an event stream from the ESPE 800. For illustration, one or more publish / subscribe APIs can be defined. Using the publish / subscribe API, an event publishing application can publish an event stream to the running event stream processor project source window of the ESPE 800, and an event subscription application can subscribe to the event stream processor project source window of the ESPE 800.

[0153] The publish / subscribe API provides cross-platform connectivity and endianness compatibility between ESP applications and other network-linked applications (such as an event publishing application implemented at the publishing device 1022 and an event subscription application implemented at one or more of the event subscription devices A 1024a, B 1024b, and C 1024c).

[0154] Return reference Figure 9, operation 906 initializes the publish / subscribe capabilities of the ESPE 800. In operation 908, one or more items 802 are started. The one or more started items may run in the background on the ESP device. In operation 910, an event block object is received from one or more computing devices of the event publishing device 1022.

[0155] The ESP subsystem 1001 may include a publish client 1002, an ESPE 800, a subscribe client A 1004, a subscribe client B 1006, and a subscribe client C 1008. The publish client 1002 may be started by an event publishing application executing at the publishing device 1022 using the publish / subscribe API. The subscribe client A 1004 may be started by an event subscription application A executing at the event subscription device A 1024a using the publish / subscribe API. The subscribe client B 1006 may be started by an event subscription application B executing at the event subscription device B 1024b using the publish / subscribe API. The subscribe client C 1008 may be started by an event subscription application C executing at the event subscription device C 1024c using the publish / subscribe API.

[0156] An event block object containing one or more event objects is injected into the source window of one or more source windows 806 from an example of an event publishing application on the event publishing device 1022. The event block object may be generated, for example, by the event publishing application and received by the publish client 1002. A unique ID may be maintained when passing the event block object between one or more source windows 806 and / or one or more export windows 808 of the ESPE 800, and to the subscribe client A 1004, the subscribe client B 1006, and the subscribe client C 1008, as well as to the event subscription device A 1024a, the event subscription device B 1024b, and the event subscription device C 1024c. The publish client 1002 may further generate a unique embedded transaction ID when the event block object is processed by a continuous query and include it in the event block object, along with the unique ID assigned to the event block object by the publishing device 1022.

[0157] In operation 912, the event block object is processed by one or more continuous queries 804. In operation 914, the processed event block object is output to one or more computing devices of the event subscription devices 1024a to 1024c. For example, the subscribe client A 1004, the subscribe client B 1006, and the subscribe client C 1008 may respectively send the received event block object to the event subscription device A 1024a, the event subscription device B 1024b, and the event subscription device C 1024c.

[0158] The ESPE 800 maintains the event block container relationship (containership) of the received event blocks when the event blocks are published into the source window, and applies various event translations before outputting to subscribers to complete the directed graph defined by one or more continuous queries 804. Subscribers can correlate the subscribed events of a group back to the published events of the group by comparing the unique ID of the event block object attached to the event block object by the publisher (e.g., the publishing device 1022) with the event block ID received by the subscriber.

[0159] In operation 916, a determination is made as to whether to stop processing. If processing is not stopped, then the processing continues in operation 910 to continue receiving one or more event streams containing event block objects from, for example, one or more network devices. If processing is stopped, then the processing continues in operation 918. In operation 918, the started items are stopped. In operation 920, the ESPE is shut down.

[0160] As mentioned, in some embodiments, after receiving and storing data, the big data is processed for an analysis project. In other embodiments, a distributed application processes continuous flowing data from distributed sources in real time by applying queries to the data before distributing the data to geographically distributed receivers. As mentioned, an event stream processing engine (ESPE) can continuously apply queries to data when receiving the data and determine which entities receive the processed data. This allows for the real-time processing and distribution of large amounts of data received and / or collected in a variety of environments. For example, as shown with respect to Figure 2 what is shown, data can be collected from network devices that can include devices within the Internet of Things (such as devices within a home automation network). However, this data can be collected from a variety of different resources in a variety of different environments. In any such case, embodiments of the present technology allow for the real-time processing of this data.

[0161] Aspects of the present disclosure provide technical solutions to technical problems, such as computational problems that occur when an ESP device fails, which results in a complete service interruption and potentially significant data loss. When the streamed data is supporting mission-critical operations (such as operations that support ongoing manufacturing or drilling operations), data loss can be catastrophic. Embodiments of the ESP system implement fast and seamless failover of the ESPE running at multiple ESP devices without service interruption or data loss, thus significantly improving the reliability of operating systems that rely on the live or real-time processing of data streams. The event publishing system, the event subscription system, and each ESPE not executing at the failed ESP device are unaware of or unaffected by the failed ESP device. The ESP system can include thousands of event publishing systems and event subscription systems. The ESP system keeps the failover logic and awareness within the boundaries of the outbound messaging network connectors and outbound messaging network devices.

[0162] In one example embodiment, a system is provided for supporting failover when processing event stream processing (ESP) event chunks. The system includes, but is not limited to, an outbound messaging network device and a computing device. The computing device includes, but is not limited to, a processor and a computer-readable medium operably coupled to the processor. The processor is configured to execute an ESP engine (ESPE). The computer-readable medium has instructions stored thereon that, when executed by the processor, cause the computing device to support failover. An event chunk object is received from the ESPE that includes a unique identifier. A first state of the computing device is determined to be active or standby. When the first state is active, a second state of the computing device is determined to be up-to-date active or not up-to-date active. When the computing device switches from a standby state to an active state, it is determined to be up-to-date active. When the second state is up-to-date active, an identifier of the last published event chunk object that uniquely identifies the last published event chunk object is determined. A next event chunk object is selected from a non-transitory computer-readable medium accessible by the computing device. The next event chunk object has an event chunk object identifier greater than the determined identifier of the last published event chunk object. The selected next event chunk object is published to the outbound messaging network device. When the second state of the computing device is not up-to-date active, the received event chunk object is published to the outbound messaging network device. When the first state of the computing device is standby, the received event chunk object is stored in the non-transitory computer-readable medium.

[0163] Figure 11A flowchart of an example of a process for generating and using a machine learning model according to some aspects. Machine learning is a branch of artificial intelligence related to mathematical models that can learn from data, classify data, and make predictions about data. Such mathematical models, which may be referred to as machine learning models, can classify input data among two or more classes; cluster input data among two or more groups; predict a result based on input data; identify patterns or trends in input data; identify the distribution of input data in space; or any combination of these. Examples of machine learning models can include: (i) neural networks; (ii) decision trees, such as classification trees and regression trees; (iii) classifiers, such as naive bias classifiers, logistic regression classifiers, ridge regression classifiers, random forest classifiers, least absolute shrinkage and selector (LASSO) classifiers, and support vector machines; (iv) clusterers, such as k-means clusterers, mean shift clusterers, and spectral clusterers; (v) factorizers, such as factorization machines, principal component analyzers, and kernel principal component analyzers; and (vi) ensembles or other combinations of machine learning models. In some examples, neural networks can include deep neural networks, feedforward neural networks, recurrent neural networks, convolutional neural networks, radial basis function (RBF) neural networks, echo state neural networks, long short-term memory neural networks, bidirectional recurrent neural networks, gated neural networks, hierarchical recurrent neural networks, stochastic neural networks, modular neural networks, spiking neural networks, dynamic neural networks, cascaded neural networks, neuro-fuzzy neural networks, or any combination of these.

[0164] Different machine learning models can be used interchangeably to perform tasks. Examples of tasks that can be performed using machine learning models at least in part include various types of scoring; bioinformatics; chemoinformatics; software engineering; fraud detection; customer segmentation; generating online recommendations; adaptive websites; determining customer lifetime value; search engines; advertising in real-time or near real-time; classifying DNA sequences; sentiment computing; performing natural language processing and understanding; object recognition and computer vision; robotic motion; playing games; optimization and metaheuristics; detecting network intrusions; medical diagnosis and monitoring; or predicting when an asset (such as a machine) will need maintenance.

[0165] Any number and combination of tools can be used to build a machine learning model. Examples of tools for building and managing machine learning models can include Enterprise Miner, Fast Predictive Modeler, and Model Manager, SAS Cloud Analytics Services SAS All of which are provided by SAS Institute Inc. of Cary, North Carolina, USA.

[0166] Machine learning models can be constructed through a process that is at least partially automated (e.g., with little or no human involvement), known as training. During training, input data can be iteratively supplied to the machine learning model so that the machine learning model can identify patterns related to the input data or identify relationships between the input data and output data. Using training, the machine learning model can be transformed from an untrained state to a trained state. The input data can be split into one or more training sets and one or more validation sets, and the training process can be repeated multiple times. The splitting can follow k-fold cross-validation rules, holdout rules, hold-p rules, or clamping rules. The following flowcharts regarding Figure 11 describe an overview of training and using machine learning models.

[0167] In block 1104, training data is received. In some instances, the training data is received from a remote database or a local database constructed from various data subsets, or is input by a user. The training data can be used in its original form to train the machine learning model, or can be preprocessed into another form that can then be used to train the machine learning model. For example, the original form of the training data can be smoothed, truncated, aggregated, clustered, or otherwise manipulated into another form that can then be used to train the machine learning model.

[0168] In block 1106, the machine learning model is trained using the training data. The machine learning model can be trained in a supervised, unsupervised, or semi-supervised manner. In supervised training, each input in the training data is associated with an expected output. This expected output can be a scalar, a vector, or a different type of data structure, such as text or an image. This enables the machine learning model to learn the mapping between the input and the expected output. In unsupervised training, the training data contains inputs but no expected output, such that the machine learning model must find the structure in the inputs on its own. In semi-supervised training, only some of the inputs in the training data are associated with an expected output.

[0169] In block 1108, a machine learning model is evaluated. For example, an evaluation data set may be obtained, e.g., via user input or from a database. The evaluation data set may contain inputs related to an expected output. The inputs may be provided to the machine learning model and the output from the machine learning model may be compared with the expected output. If the output from the machine learning model closely corresponds to the expected output, then the machine learning model may have high accuracy. For example, if 90% or more of the outputs from the machine learning model are the same as the expected output in the evaluation data set, then the machine learning model may have high accuracy. Otherwise, the machine learning model may have low accuracy. The 90% figure is only an example. The actual and expected accuracy percentages depend on the problem and the data.

[0170] In some instances, if the machine learning model has insufficient accuracy for a particular task, the process may return to block 1106 where the machine learning model may be further trained using additional training data or otherwise modified to improve accuracy. If the machine learning model has sufficient accuracy for a particular task, the process may continue to block 1110.

[0171] In block 1110, new data is received. In some instances, the new data is received from a remote database or a local database constructed from various data subsets, or by user input. The machine learning model may not be aware of the new data. For example, the machine learning model may not have processed or analyzed the new data previously.

[0172] In block 1112, the trained machine learning model is used to analyze the new data and provide a result. For example, the new data may be provided as an input to the trained machine learning model. The trained machine learning model may analyze the new data and provide a result that includes a classification of the new data into a particular category, a clustering of the new data into a particular group, a prediction based on the new data, or any combination of these.

[0173] In block 1114, the result is post - processed. For example, the result may be added to, multiplied by, or otherwise combined with other data as part of a job. As another example, the result may be transformed from a first format (e.g., a time - series format) to another format (e.g., a count - sequence format). Any number and combination of operations may be performed on the result during post - processing.

[0174] A more specific example of a machine learning model is Figure 12The neural network 1200 shown in []. The neural network 1200 is represented as multiple layers of interconnected neurons (e.g., neuron 1208) that can exchange data with each other. The layers include an input layer 1202 for receiving input data, a hidden layer 1204, and an output layer 1206 for providing results. The hidden layer 1204 is called hidden because it may not be directly observable or its inputs may not be directly accessible during the normal operation of the neural network 1200. Although the neural network 1200 is shown for illustrative purposes as having a specific number of layers and neurons, the neural network 1200 can have any number and combination of layers, and each layer can have any number and combination of neurons.

[0175] The neurons and the connections between neurons can have numerical weights that can be tuned during training. For example, training data can be provided to the input layer 1202 of the neural network 1200, and the neural network 1200 can use the training data to tune one or more numerical weights of the neural network 1200. In some instances, backpropagation can be used to train the neural network 1200. Backpropagation can include determining the gradient of a specific numerical weight based on the difference between the actual output of the neural network 1200 and the desired output of the neural network 1200. Based on the gradient, one or more numerical weights of the neural network 1200 can be updated to reduce the difference, thereby increasing the accuracy of the neural network 1200. This process can be repeated multiple times to train the neural network 1200. For example, this process can be repeated hundreds or thousands of times to train the neural network 1200.

[0176] In some instances, the neural network 1200 is a feedforward neural network. In a feedforward neural network, each neuron only propagates the output value to the subsequent layer of the neural network 1200. For example, in a feedforward neural network, data can only move in one direction (forward) from one neuron to the next neuron.

[0177] In other instances, the neural network 1200 is a recurrent neural network. A recurrent neural network can include one or more feedback loops, allowing data to propagate through the neural network 1200 in both the forward and backward directions. This can allow information to persist within the recurrent neural network. For example, a recurrent neural network can determine the output based at least in part on information that the recurrent neural network has previously seen, giving the recurrent neural network the ability to use previous inputs to inform the output.

[0178] In some instances, neural network 1200 operates by: receiving a digital vector from a layer; transforming the digital vector into a new digital vector using a numerical weight matrix, a non-linearity, or both; and providing the new digital vector to a subsequent layer of neural network 1200. Each subsequent layer of neural network 1200 may repeat this process until neural network 1200 outputs a final result at output layer 1206. For example, neural network 1200 may receive a digital vector as an input at input layer 1202. Neural network 1200 may multiply the digital vector by a numerical weight matrix to determine a weighted vector. The numerical weight matrix may be tuned during the training of neural network 1200. Neural network 1200 may use a non-linearity (such as sigmoid tangent or hyperbolic tangent) to transform the weighted vector. In some instances, the non-linearity may include a rectified linear unit that can be expressed using the following equation:

[0179] y = max(x, 0)

[0180] where y is the output and x is the input value from the weighted vector. The transformed output may be supplied to a subsequent layer of neural network 1200, such as hidden layer 1204. A subsequent layer of neural network 1200 may receive the transformed output, multiply the transformed output by a numerical weight matrix and a non-linearity, and provide the result to another layer of neural network 1200. This process continues until neural network 1200 outputs a final result at output layer 1206.

[0181] Other examples of the present disclosure may include any number and combination of machine learning models having any number and combination of features. The machine learning models may be trained in a supervised, semi-supervised, or unsupervised manner or any combination thereof. The machine learning models may be implemented using a single computing device or multiple computing devices (such as the communication grid computing system 400 discussed above).

[0182] Some examples of the present disclosure implemented at least in part by using a machine learning model may reduce the total number of processing iterations, time, memory, power, or any combination thereof consumed by a computing device when analyzing data. For example, a neural network may be more likely to identify patterns in data than other methods. This may enable the neural network to analyze data using fewer processing cycles and less memory than other methods while achieving similar or higher accuracy.

[0183] Some machine learning methods can be executed and processed more efficiently and quickly using machine learning specific processors (e.g., not general-purpose CPUs). Such processors can also provide energy savings when compared to general-purpose CPUs. For example, some of these processors can include graphics processing units (GPUs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), artificial intelligence (AI) accelerators, neural computing cores, neural computing engines, neural processing units, purpose-built chip architectures for deep learning, and / or some other machine learning specific processors that implement machine learning methods or one or more neural networks using semiconductor (e.g., silicon (Si), gallium arsenide (GaAs)) devices. Additionally, these processors can also be used in heterogeneous computing architectures with several and various different types of cores, engines, nodes, and / or layers to achieve various energy efficiency, processing speed improvement, data communication speed improvement, and / or data efficiency goals and improvements at every part of the system when compared to homogeneous computing architectures that use CPUs for general computing.

[0184] One or more embodiments of the present technology are useful for making predictions regarding processes for bonding materials together (e.g., joining operations). There are many different ways to bond materials together, including adhesive bonding, fusion bonding, stitch bonding, and welding bonding. Joining operations are used in many different industries including manufacturing and construction. Joining can be used to manufacture various products and their packages, such as vehicles, consumer equipment, mobile devices, computers, medical devices, clothing, and chemical products. Joining can be used in construction, such as to form buildings, bridges, and electrical systems. In a joining operation, it can be difficult to know the quality of the joined parts and thus the quality of the product being formed without damaging the joined parts themselves. Predictions regarding joining operations can be useful for making determinations regarding joining operations without damaging the manufactured product. For example, one or more embodiments are useful for manufacturing integrated circuit (IC) chips.

[0185] As is known to those skilled in the art, the semiconductor manufacturing process for IC chips is typically performed in two stages often referred to as the "front end" and the "back end". The front end refers to the manufacturing process for fabricating silicon wafers from blank wafers. Once completed, the fabricated silicon wafers have multiple IC chips on them. The "back end" refers to the manufacturing process that occurs after all the IC chips have been built on the silicon wafer. For the purpose of describing example embodiments, the example "back end" process will be described in more detail.

[0186] Figure 13 and 14 illustrate an overview of an example "back end" semiconductor manufacturing process for manufacturing IC chips according to one embodiment. Specifically, Figure 13Describe an example method 1300 for performing the "backend" of a semiconductor manufacturing process. Figure 14 Further illustrate graphically a method 1400 for performing a backend semiconductor manufacturing process and some components associated with the backend of a semiconductor manufacturing process. By way of example, such components include a silicon wafer 1402 having a plurality of dies 1404 thereon. Each die 1404 includes a block of semiconductor material within an IC chip on which functional circuitry is fabricated. Thus, each die 1404 is a fully functional IC chip that has not yet been encapsulated.

[0187] As Figure 13 seen, the example method 1300 begins with a wafer grinding phase (block 1302). In this phase, the thickness of a silicon wafer (e.g., Figure 14 the silicon wafer 1402) is reduced to a desired thickness (e.g., 50 μm to 75 μm) by an abrasive grinding wheel. For example, the grinding can be implemented in multiple steps. For example, in a first step, a grinding wheel (not shown) can first roughly grind the silicon wafer (e.g., Figure 14 the silicon wafer 1402) using coarse grit and remove most of its excess thickness. In a second step, the grinding wheel can use a different grit (e.g., a much finer grit) to finish grinding the silicon wafer. The finer grit is used to grind the silicon wafer precisely to the desired thickness and polish the silicon wafer.

[0188] Once the grinding is complete, Figure 13 method 1300 separates the dies (e.g., Figure 14 the dies 1404) from each other in a wafer sawing phase (block 1304). Wafer sawing, also known as "wafer cutting", is a process by which the dies are separated from each other on a silicon wafer (e.g., Figure 14 the silicon wafer 1402). Generally, wafer cutting is accomplished by mechanically sawing the silicon wafer in the areas located between the dies. As Figure 14 seen, example "areas" for sawing are indicated using vertical and horizontal lines 1406, 1408 respectively and are commonly referred to as "cutting streets" or "scribe lines". The excess areas around the dies can be further removed or reduced by additional sawing or grinding.

[0189] Next, in Figure 13 a die bonding phase (block 1306), the separated dies (e.g., Figure 14 the dies 1404) are attached or "bonded" to a lead frame (e.g., Figure 14 the lead frame 1410). A "lead frame" can be a metallic (e.g., copper) conductive structure within a chip package that is configured to carry signals from the die to the "outside world" or vice versa. In Figure 14In this case, die 1404 is bonded to lead frame 1410 using an epoxy adhesive or solder. Specifically, the epoxy adhesive or solder is placed on lead frame 1410 in a predetermined pattern. Subsequently, individual dies 1404a, 1404b, 1404c that were separated from each other during the wafer sawing stage are picked up and placed on the dispensed epoxy adhesive or solder and cured. Curing hardens the epoxy or solder such that each die 1404a, 1404b, 1404c remains attached to lead frame 1410 and such that each die 1404a, 1404b, 1404c achieves the desired mechanical and electrical properties.

[0190] Those skilled in the art will readily appreciate that the present disclosure is not limited to using an epoxy adhesive or solder for bonding die 1404 to lead frame 1410. In other embodiments, for example, die 1404 is bonded to lead frame 1410 using well-known thermocompression die bonding techniques. Thermocompression die bonding does not use an adhesive material to attach die 1404 to lead frame 1410. In fact, heat and force are applied to die 1404 to form a metal bond with the underlying lead frame 1410.

[0191] Figure 13 The next stage in the back-end process is the wire bonding stage (block 1308). In this stage, for example, a wire bonder (not shown) uses multiple wires to establish electrical connections between a die (e.g., die 1404a) and a lead frame (e.g., lead frame 1410). As described more fully below, each connection involves two types of bonds referred to herein as "ball bonds" and "stitch bonds". A ball bond is formed between the wire and the die, and a stitch bond is formed between the lead frame and the wire. Any number of wires can be used in this stage to bond the die to the lead frame. For example, depending on the manufacturer, different numbers of wires (e.g., >5, >10, >20, or >100) can be used to electrically connect the die to the lead frame.

[0192] Once the wire bonding stage is complete, a molding stage (block 1310) is performed in Figure 13 This stage helps to protect the die (e.g., die 1404) mechanically (e.g., from damage caused by one or more physical impacts) and environmentally (from particles and other contaminants). As Figure 14As seen in, the molding stage encapsulates die 1404 by wrapping it with a predetermined material, e.g., a plastic epoxy material. For example, the molding process can be any of several well-known molding methods. For example, in a first method (i.e., the hermetic method), a ceramic plate or a metal lid is attached to die 1404, thereby forming a seal. In a second method, a plastic epoxy material is melted onto die 1404 and cured to form a seal. However, regardless of the specific method used in the molding stage, the molding compound is cured using heat.

[0193] After Figure 13 the molding stage in, the next stage can include a strengthening stage (block 1312). For example, plating operations can be used to help protect lead frame 1410 from corrosion and wear by applying a thin metal layer over the leads of a lead frame (e.g., lead frame 1410) to mechanically and electrically connect and couple the lead frame to a lower substrate, e.g., a printed circuit board (PCB). As another example, etching operations can be used to enhance the conductivity of the lead frame.

[0194] After the strengthening stage is completed, a marking stage (block 1314) is performed in Figure 13 . In this stage, the manufacturer marks the package of the IC chip with various identifiers and other distinguishing information, e.g., the manufacturer's name and logo, the name of the IC chip, the date code, the lot identifier, and the like. Any number of methods, e.g., ink-based and laser-based methods, can be used to mark the package. However, in at least one embodiment of the present technology, laser-based methods are preferred. However, regardless of the specific method used in the marking stage, the marking can identify the IC chip and its origin and facilitate traceability.

[0195] The next stage in the back-end process is Figure 13 the trimming and forming stage (block 1316) in. In this stage, the metal posts that connect the leads extending from the package of the IC chip, commonly referred to as "dambars", are removed. Additionally, the leads are bent or otherwise shaped into a suitable form for placement on a lower substrate. Finally, in Figure 13 the stage (block 1318), the trimmed and formed IC chips are inspected, tested, and packaged for distribution.

[0196] Figure 13 And 14 show example stages that can be performed as part of a semiconductor manufacturing process. Each stage can have one or more operations. Additionally, different, more, or fewer stages can be performed (e.g., stages can be combined or performed in a different order). Thus, Figure 13 And 14The methods 1300 and 1400 described are merely examples for describing the example bonding operations of an IC chip. In any case, as previously described, a wire bonder may use multiple wires during the wire bonding stage (block 1308) to establish an electrical connection between Figure 14 the die 1404 and the lead frame 1410. Figure 15A and 15B illustrate examples of some components and bonding elements of an IC chip 1520 according to at least one embodiment of the present technology. Specifically, Figure 15A and 15B illustrate an exemplary lead frame 1410 bonded to a lower substrate 1500. The conductive "legs" or "pins" of the lead frame 1410 are also attached to the substrate 1500. In this embodiment, Figure 15A the lead frame 1410 is illustrated as having three (3) conductive legs 1410a, 1410b, 1410c. However, this is for illustrative purposes only. Those skilled in the art will readily understand that more or fewer legs coupled to the lead frame 1410 may be present as needed or desired. However, regardless of the number of conductive legs, as Figure 15B seen in Figure 13 , the legs 1410a, 1410b, 1410c of the lead frame 1410 may be bent during the trimming and forming stage ( Figure 13 block 1316) such that the conductive legs make proper physical contact with the lower substrate (e.g., a printed circuit board (PCB) for example).

[0197] As Figure 15A seen in Figure 15A , metal wires 1504 extend between bonding pads 1506 disposed on the surface of the die 1404a and target pads 1508 disposed on each leg 1410a, 1410b, 1410c of the lead frame 1410. In some embodiments, a conductive film layer 1502 may be disposed between the target pads 1508 and each leg 1410a, 1410b, 1410c of the lead frame 1410. Both the bonding pads 1506 and the target pads 1508 include thin conductive metal layers or films, e.g., silver for example. However, as those skilled in the art will understand, one or both of the bonding pads 1506 and the target pads 1508 may include conductive materials other than silver. However, regardless of their composition, the wires 1504 are attached to the bonding pads 1506 via "ball bonds" 1510 and to the target pads 1508 via "stitch bonds" 1512. The wires 1504, ball bonds 1510, and stitch bonds 1512 provide the physical means for establishing an electrical connection between the die 1404a and the legs of the lead frame 1410.

[0198] Figures 16A to 16DIllustrate an example method 1600 for forming a ball bond 1510 and a wire bond 1512 according to an embodiment of the present disclosure. As seen in these figures, a feed wire 1604 extends through a hole in a capillary tool 1602. Generally, the feed wire 1604 comprises a conductive material, for example, such as gold, copper, or aluminum.

[0199] To create the ball bond 1510, a high voltage charge is applied to the feed wire 1604 ( Figure 16A ). This charge melts the end of the feed wire 1604. However, surface tension rather than dripping or falling causes the melted feed wire 1604 to form a ball 1606 at the tip of the capillary tool 1602. Then, the capillary tool 1602 is lowered toward the die 1404a such that the melted ball 1606 is pressed onto a bond pad 1506 that may be heated in some cases ( Figure 16B ). For example, heat (e.g., thermocompression) and / or ultrasonic energy (e.g., ultrasonic energy) are then applied along with the downward pressure of the capillary tool 1602 to form the ball bond 1510.

[0200] Next, the capillary tool 1602 is lifted away from the ball bond 1510 and moved toward a target pad 1508 of the lead frame 1410. Then, the capillary tool 1602 is lowered to contact the target pad 1508, thereby squeezing the feed wire 1604 between the tip of the capillary tool 1602 and the surface of the target pad 1508 ( Figure 16C ). The pressure applied by the capillary tool 1602 to the feed wire 1604 along with heat and / or ultrasonic energy forms the wire bond 1512. After being formed in this way, the capillary tool 1602 is then lifted away from the target pad 1508, which causes the feed wire 1604 to tear, and moves to the next bond pad 1506 to form the next ball bond 1510 ( Figure 16D ).

[0201] Wire bonding is a very complex process, which becomes even more complex due to the possibility of having multiple dies on an IC chip. Embodiments of the present disclosure read and record multiple process parameters during bonding. According to this embodiment and as described more fully below, the process parameters can be used to generate an anomaly indicator that indicates the risk of an anomaly occurring during such bonding operations.

[0202] Figure 17 is a perspective view of an example IC chip 1700 that has undergone the back end of the semiconductor manufacturing process previously described. As Figure 17As seen, the IC chip 1700 includes a plurality of dies 1404a, 1404b, 1404c, and 1404d. Each die 1404a, 1404b, 1404c, 1404d has its own bonding connection to the lead frame 1410 and is thus physically and electrically connected to the lead frame 1410 via wires 1504 and ball and pin type bonders 1510, 1512 respectively, as previously described. This embodiment illustrates the wire layout of 40 wires 1504. However, those skilled in the art should readily understand that this is for illustrative purposes only. Such a wire layout may include more or fewer wires 1504 as needed or desired.

[0203] In one embodiment, the dies 1404a, 1404b, 1404c, 1404d all have the same number of wires 1504 connecting them to the lead frame 1410. However, this is not necessary. The dies 1404a, 1404b, 1404c, and 1404d need not have the same number of wires 1504. For example, die 1404a may have a first number of wires 1504 connecting it to the lead frame 1410, and die 1404c may have a second, different number of wires 1504 connecting it to the lead frame 1410. One or both of the other dies 1404b, 1404d may also have the same or different number of wires as the other dies on the IC chip 1700.

[0204] As described above, the back end of the semiconductor manufacturing process may include inspection and testing phases (e.g., Figure 13 box 1318 in). During this phase, the quality of the plurality of ball bonders 1510 and pin bonders 1512 generated during the wire bonding phase (e.g., Figure 13 box 1308 in) is tested. This test is not performed on all completed IC chips, but on a representative sample of IC chips from each batch.

[0205] Generally, there are different types of tests (non-destructive and destructive) for determining the quality of wire bonders. Non-destructive tests generally involve visual inspection of the ball bonders 1510 and pin bonders 1512, and their quality is inferred by the corresponding diameters of the bonders. If it passes the visual inspection, the IC chip can be put back into the workflow for distribution. However, visual inspection can be an inaccurate and error-prone way to infer the quality of the bonders.

[0206] In contrast, destructive testing destroys the bond by damaging or deforming the ball and leaded joints 1510, 1512 and / or the wire 1504 such that the IC chip under test cannot be recovered for distribution. Thus, while destructive testing is useful for determining the exact failure point of a given ball and / or leaded joint 1510, 1512, it is necessarily time-consuming and expensive and is typically performed by a separate test machine. In addition, since the bond is destroyed, it can only be used to make inferences about the quality of other bonded chips that have not undergone destructive testing.

[0207] Generally, the semiconductor manufacturing industry relies on two specific destructive tests. These are the "ball shear" test and the "lead pull" test. For details on these tests, the interested reader is referred to the government-regulated standards for microelectronic devices (e.g., the United States government standard in Revision K of MIL-STD-883C, titled "Test Method Standard–Microcircuits," issued on April 25, 2016, for military and aerospace electronic systems, the full text of which is incorporated herein by reference). However, briefly, the "ball shear" test is used to test the quality of the ball joint 1510, while the "lead pull" test is used to test the quality of the leaded joint 1512.

[0208] Figures 18A to 18C is a perspective view illustrating an example method 1800 for performing a ball shear test. As seen in these figures, the ball joint 1510 is bonded to the bond pad 1506 ( Figure 18A ). The tool arm 1802 is positioned above the bond pad 1506 and is laterally moved to press into contact with the ball joint 1510 ( Figure 18B ). The tool arm 1802 continues to move laterally against the ball joint 1510 until the force applied by the tool arm 1802 causes the ball joint 1510 to separate from the bond pad 1506 (i.e., shear) ( Figure 18C ). The shear breaks the bond. The shear force applied by the tool arm 1802 is measured throughout the test (e.g., in grams). Thus, at the end of the test, the exact amount of shear force required to shear the ball joint 1510 off the bond pad 1506 is known.

[0209] Figures 19A to 19C is a perspective view illustrating an example method 1900 for performing a lead pull test. As above, the leaded joint 1512 is attached to the target pad 1508 and the ball joint 1510 is attached to the bond pad 1506.

[0210] In the lead pull test, the hook tool 1902 is lowered and placed under the wire 1504 ( Figure 19A)。Once positioned in this manner, the hook tool 1902 is then raised to contact the wire 1504( Figure 19B )。The hook tool 1902 continues to move upward such that it applies an upward force on the wire 1504. Eventually, this upward force pulls the pin-type joint 1512 away from the target pad 1508, thereby effectively pulling the wire 1504 away from the lead frame 1410( Figure 19C ) and breaking the joint. The upward force applied by the hook tool 1902 is measured throughout the test (e.g., in grams). Thus, at the end of the test, the precise amount of upward force required to lift the pin-type joint off the target pad 1508 can be determined.

[0211] To perform these destructive tests, a selected IC chip is physically moved from any machine performing the wire bonding stage to another machine configured to perform the tests. Typically, these actions are performed manually. Additionally, not all of the ball joints 1510 and pin-type joints 1512 of a given IC chip are tested. Instead, only a representative sample of the joints is tested.

[0212] Since these tests are so destructive, the IC chip under test is destroyed, even if not all of the joints on the IC chip are tested. Thus, chip manufacturers typically take a representative sample of IC chips from a given batch and then perform these destructive tests on the representative sample. The results achieved from testing the (several) samples are then stored and subsequently used to make inferences about the quality of the IC chips in the batch.

[0213] However, such tests are very time-consuming and may slow down the integrated chip manufacturing process. Additionally, since these tests destroy the IC chips, the IC chips must be discarded rather than sold. Thus, destructive tests are also expensive.

[0214] In addition, problems arise regarding which of the wires 1504 on a given IC chip should be selected for testing. Ideally, IC chips with poor quality wires and / or joints should be identified before leaving the factory. However, there is no guarantee that a given random sample of an IC chip and its corresponding wires and joints will include such IC chips. Additionally, a random sample of IC chips and wires may still incorrectly certify that a given batch has acceptable quality. Currently, there is no reliable method for identifying IC chips with poor quality bond connections. Thus, there is currently no way to identify a specific wire or subset of wires for quality testing.

[0215] Similar problems occur in other types of joining scenarios. Thus, the adapter will have to choose between destructive testing (which destroys the joint under study) or non-destructive testing. Previous conventional testing techniques for non-destructive testing are likely to be inaccurate and time-consuming (which requires manual, visual inspection). Additionally, previous testing methods were typically performed only on a small subset of the products in a given batch. Thus, there is usually a problem in determining which products to test.

[0216] Accordingly, embodiments of the present disclosure solve these and other problems by providing a model specifically for predicting likely anomalies without having to destroy the joint (e.g., predicting the force required to shear the ball bond 1510 from the bond pad 1506 and separate the wire bond 1512 from the target pad 1508 without having to perform quality assurance testing). The technique can use feature engineering to select and transform relevant variables derived from raw data measurements to process the data for the model (e.g., a predictive model specified using machine learning or statistical modeling). Feature engineering can involve building features by identifying the raw variables that will be useful predictive variables for the predictive model and / or derived features created by manipulating the raw variables (e.g., addition, subtraction, multiplication, and ratios). Feature engineering can involve transformations performed by further manipulating the established predictor variables to perform model performance (e.g., ensuring the variables are on the same scale, or within an acceptable range for the model). Feature engineering can involve feature extraction by automatically creating new variables from the raw data; and feature selection (e.g., to remove irrelevant or redundant features) by analyzing, judging, and ranking the features using algorithms to determine the features of the model.

[0217] Specifically, the embodiments disclosed herein utilize process data measured during a joining operation (e.g., wire bonding) to develop an analytical model. The process parameters obtained through feature engineering may be referred to herein as "motion features". The analytical model allows the user to accurately estimate the ball shear and wire pull forces (i.e., values) that have heretofore been obtainable only through destructive testing. Thus, this embodiment can accurately predict the quality of wire bonds based on motion feature data.

[0218] Such accurate predictions provide benefits and advantages not offered by conventional testing methods. For example, the analytical model generated according to this embodiment can reduce the amount of inspection performed and in some instances can eliminate it entirely. Even in cases where quantitative inspection or destructive testing is required for shrinkage reasons, the methods herein can still enable the manufacturer to be strategic in sampling within the product population. For example, IC chips predicted by this embodiment to have poorly quality joints can be discarded from the population for shipment and thus from the sampled population. Sampling the preferred population can provide an overall improvement in in-field quality and can improve efficiency (e.g., by reducing the chance of recall or rejection of the shipped and sampled groups).

[0219] As previously stated, one or more embodiments described herein are illustrated only as examples in the context of the bonding of integrated circuits. Those skilled in the art will appreciate that the techniques described herein can be applied to other types of bonding operations.

[0220] Figure 20A An example block diagram of a training system 2000 is illustrated. For example, a training system can be useful for training a computer model to make predictions regarding bonding operations. For example, the bonding operation can be a bonding operation described herein, such as an adhesive bonding operation (e.g., for bonding a windshield to a vehicle or in wafer bonding) or a fusion bonding (e.g., for wire bonding), to name a few. There are many types and subtypes of bonding. For example, regarding fusion bonding, there are subtypes that include: tungsten inert gas (TIG) welding, also known as helium arc and gas tungsten arc welding (GTAW); flux-cored arc welding (FCAW); stick welding, also known as shielded metal arc welding (SMAW); metal inert gas (MIG) welding, also known as gas metal arc welding (GMAW); laser beam welding; electron beam welding; plasma arc welding; atomic hydrogen welding; and electroslag welding. For simplicity, different possible bonding operations or hybrid bonding operations will be referred to herein simply as bonding operations.

[0221] The training system 2000 includes a computing device 2002. The training system 2000 is configured to exchange information (e.g., via wired and / or wireless transmissions) between the devices in the system. For example, a network (not shown) can connect one or more devices of the training system 2000 to one or more other devices of the training system 2000.

[0222] For example, in one or more embodiments, the training system 2000 includes one or more input devices 2004A for receiving information related to a bonding operation 2040A (e.g., for training the computing device 2002) via one or more input interfaces 2005. For example, if the training system is training a system to make predictions regarding a bonding operation, the bonding operation 2040A can include a system for the bonding operation. Bonding operations occur in many different environments (e.g., manufacturing or construction environments). For example, in an integrated circuit manufacturing environment, bonding can be used to bond one or more wires to a surface (e.g., using ball, wedge, pin, and / or flexible bonding). The bonding operation 2040A can include bonding wires of a first set of multiple wires to corresponding surfaces of a first set of surfaces to form an integrated circuit chip. The bonding operation 2040A can have different types of bonding operations. For example, the environment 2040A can include a ball bonding operation 2042A and / or a pin bonding operation 2044A.

[0223] In one or more embodiments, computing device 2002 may receive model information 2030 for building a model to represent bonding operation 2040A or a subsequent environment. As an example, a computing system may receive training data 2032 for building a model.

[0224] Training data 2032 may include measurements 2038 of bonding operation 2040A. Additionally or alternatively, training data 2032 may include process data 2034, which may be generated from measurements 2038. For example, process data may include measurements of measurements 2038 or information generated for one or more measurements or measurement types. For example, if bonding operation 2040A is part of a physical environment, an input device (e.g., sensor 2020A) may capture measurements 2038 about the physical environment, such as heat measurements for heating the bonded surfaces or the bonding agent and / or force measurements for the bonding material. Additionally or alternatively, the environment may be a simulated environment or include additional simulated actions or calculations for a physical environment (e.g., simulated aspects of bonding operation 2040A or inferred information from measurements 2038), and a computing system component (e.g., computing device 2002 or computing system 2024) may generate information about the physical or simulated environment (e.g., process data 2034). Additionally or alternatively, computing device 2002 itself may derive process data 2034 (e.g., using process data application 2012).

[0225] Additionally or alternatively, training data 2032 includes the state of the bond. As an example, each state of the state may include one or more candidate results of a target related to detecting one or more anomalies in bonding operation 2040A. For example, if bonding operation 2040A includes an operation to bond a first set of multiple wires to a first set of surfaces, the state may be the state of the multiple wires after being bonded to the first set of surfaces (e.g., a quality state or a performance state). Candidate options may include one or more defective chip results of an integrated circuit chip in bonding operation 2040A. Candidate options may also be at the individual wire level (e.g., quality assurance test results at the individual wire level).

[0226] In addition, training system 2000 includes one or more output devices 2096 for outputting information based on bonding operation 2040A via one or more output interfaces 2006. For example, computing device 2002 may output a machine learning model 2050 that implements a prediction model based on training data 2032 related to bonding operation 2040A. For example, computing device 2002 may generate one or more weights for process data such that the process data input to the machine learning model predicts one or more candidate results of a target (e.g., a target related to detecting one or more anomalies in bonding operation 2040A).

[0227] As an example, the machine learning model 2050 can be trained through supervised learning based on receiving training data and using it to generate one or more weights such that the machine learning model can predict a candidate result. For example, if a set of measurements is received for a wire and the wire is then tested and given a normal or abnormal status, then the measurements can be weighted to predict whether the status is normal or abnormal for future measurements where the status may be unknown (e.g., where testing will destroy or damage the joint from a joining operation).

[0228] The computing device 2002 has a computer-readable medium 2010 and a processor 2008. The computer-readable medium 2010 is an electronic storage space or memory for information, so that the information can be accessed by the processor 2008. The computer-readable medium 2010 can include, but is not limited to, any type of random access memory (RAM), any type of read-only memory (ROM), any type of flash memory, etc., such as magnetic storage devices (e.g., hard disks, floppy disks, magnetic strips), optical discs (e.g., compact discs (CDs), digital versatile discs (DVDs)), smart cards, flash memory devices, etc.

[0229] The processor 2008 executes instructions (e.g., stored at the computer-readable medium 2010). The instructions can be implemented by a dedicated computer, logic circuit, or hardware circuit. In one or more embodiments, the processor 2008 is implemented in hardware and / or firmware. The processor 2008 executes the instructions, which means that it performs or controls the operations called by the instructions. The term "execute" is the process of running an application program or performing the operations called by the instructions. The instructions can be written using one or more programming languages, scripting languages, assembly languages, etc. In one or more embodiments, the processor 2008 can retrieve an instruction set from a permanent storage device and copy the instructions in an executable form to a temporary storage device such as RAM in a substantially certain form. The processor 2008 is operatively coupled to components of the computing device 2002 (e.g., the input interface 2005, the output interface 2006, and the computer-readable medium 2010) to receive, send, and process information.

[0230] In one or more embodiments, the computer-readable medium 2010 stores instructions for the processor 2008 to execute. For example, in one or more embodiments, the computer-readable medium 2010 includes instructions that cause the process data application 2012 to select or export process data 2034 and cause the machine learning model application 2014 to export the machine learning model 2050, as described herein.

[0231] In one or more embodiments, one or more applications stored on computer-readable medium 2010 are implemented in software (e.g., computer-readable and / or computer-executable instructions) stored in computer-readable medium 2010 and accessible by processor 2008 for executing instructions. One or more applications may be integrated with other analysis tools (e.g., analysis tools provided by SAS Institute Inc., Cary, NC, USA). By way of illustration only, the application is implemented using or integrated with one or more SAS software tools, such as Base SAS, Enterprise Miner TM , Event Stream Processing, High-Performance Analytics Server, Visual Data Mining and Machine Learning, LASR TM In-Database Products, Scalable Performance Data Engine, Cloud Analytics Services, Inventory Optimization, Inventory Optimization Workbench, Visual Analytics, Viya TM , for SAS In-Memory Statistics, Forecasting Server and all of which are developed and provided by SAS Institute Inc., Cary, NC, USA.

[0232] In one or more embodiments, fewer, different, and additional components may be incorporated into computing device 2002 or a system including computing device 2002. For example, in one or more embodiments, there are multiple input devices 2004 (e.g., keyboard 2022, mouse 2026, and computing system 2024). In the same or different embodiments, there are multiple output devices 2096. As another example, the same interface supports both input interface 2005 and output interface 2006. For example, a touch screen provides a mechanism for user input and for presenting output to the user. Alternatively, input interface 2005 has more than one input interface using the same or different interface technologies. Alternatively or additionally, output interface 2006 has more than one output interface using the same or different interface technologies. In one or more embodiments, the functionality of one or more input devices 2004 or one or more output devices 2096 is integrated into computing device 2002. One or more applications may be combined or further decomposed into separate applications.

[0233] The machine learning model 2050 can be used by a control system to make predictions about or control another environment. Figure 20B An example block diagram of a control system 2060 for the bonding operation 2040B is illustrated. In this example, the bonding operation 2040A and the bonding operation 2040B are different bonding operations. For example, the bonding operation 2040A can be a first set of bonding operations including operations for bonding a first set of multiple wires to a first set of surfaces, and the bonding operation 2040B can be a second bonding operation including operations for bonding a second set of multiple wires to a second set of surfaces. The second set of multiple wires can be different from the first set of multiple wires. For example, the first set of wires may have been damaged to determine a state related to the bonding operation 2040A, or the bonding operation 2040A and the bonding operation 2040B (e.g., the ball bonding operation 2042B and / or the stitch bonding operation 2044B) can be performed at different manufacturing locations or at different times. The second set of surfaces can be different from the first set of surfaces (e.g., the second set of surfaces can be related to an integrated chip of a different lot than the first set of surfaces). Embodiments herein provide a framework for making predictions about a bonding operation based on training received in a previous bonding operation. Using this framework, there is no need to damage the bond to make predictions about its quality and / or control the bonding operation to ensure a quality bond.

[0234] The control system 2060 is configured to exchange information (e.g., via wired and / or wireless transmission) between devices in the control system 2060. For example, a network (not shown) can connect one or more devices of the control system 2060 to one or more other devices of the control system 2060. Alternatively or additionally, the control system 2060 is integrated into one device (e.g., the computing device 2002 can be at a manufacturing plant or a construction site and includes equipment for capturing information related to both the bonding operation 2040A and the bonding operation 2040B). In this example, the control system 2060 includes Figure 20A the computing device 2002. In other examples, the control system 2060 includes different computing devices. For example, the computing device of the training system 2000 can be in the bonding operation 2040A and another computing device of the control system 2060 can be in another environment (e.g., a different manufacturing plant or construction site). For example, the computing device 2002 can alternatively be an output device 2096 or receive the machine learning model 2050 from the output device 2096. In any case, the computing device 2002 can access the machine learning model. In this example, the computer-readable medium 2010 has a machine learning model 2050 trained on the training data 2032 of the bonding operation 2040A.

[0235] In one or more embodiments, the control system 2060 includes one or more input devices 2004B for receiving information related to the bonding operation 2040B (e.g., for making predictions regarding the bonding operation 2040B) via one or more input interfaces 2005. For example, the computing device 2002 may receive input data 2070. The input data 2070 may indicate process data 2074 generated from the measurements 2072 of the bonding operation 2040B.

[0236] The input data 2070 may be generated from measurements of the same type as the training data 2032. For example, the process data 2074 may include measurements of the measurements 2038 or information generated for one or more measurements or measurement types. For example, if the bonding operation 2040B is part of a physical environment, input devices (e.g., sensors 2020B) may capture measurements 2072 of the physical environment, such as heat measurements for heating the bonded surface or the bonding agent and / or force measurements for the bonding material. Additionally or alternatively, the environment may be a simulated environment or include additional simulated actions or calculations for the physical environment (e.g., simulated aspects of the bonding operation 2040B or inferred information from the measurements 2072), and a computing system component (e.g., the computing device 2002 or the computing system 2024) may generate information (e.g., process data 2074) regarding the physical or simulated environment. Additionally or alternatively, the computing device 2002 itself may derive the process data 2074.

[0237] Additionally or alternatively, the input data 2070 may contain fewer, more, or different measurement types than or derived from the training data 2032. For example, the manufacturing location may update the machine learning model 2050 to account for fewer, more, or different measurement types.

[0238] In one or more embodiments, the control system 2060 includes one or more output devices 2094 for outputting information based on the bonding operation 2040B via one or more output interfaces 2006. For example, the information may include bonding operation control information 2090 for controlling the bonding operation 2040B (e.g., the control information 2090 may indicate stopping or adjusting the bonding operation due to a predicted quality issue with the bonding operation 2040B). For example, the bonding operation control 2090 may include an anomaly prediction measure 2052 indicating the risk of an anomaly occurring in the bonding operation 2040B. For example, if the bonding operation 2040B includes a bonding operation for bonding a set of multiple wires to form an integrated circuit chip, the anomaly prediction measure may indicate the risk of an anomaly in the integrated circuit chip manufacturing process during the bonding operation 2040B.

[0239] In one or more embodiments, a computing device has a computer-readable medium 2010 and a processor 2008 for generating an anomaly prediction 2052. For example, in one or more embodiments, the computer-readable medium 2010 includes instructions for causing a machine learning model application (e.g., machine learning model application 2014) to access a machine learning model (e.g., machine learning model 2050) trained on training data of a first bonding operation (e.g., machine learning model 2050 trained on bonding operation 2040A), weight input data 2070 according to the machine learning model, and generate an anomaly prediction 2052 based on the weighting of the input data 2070 according to the machine learning model.

[0240] In one or more embodiments, a computing system (e.g., computing device 2002, training system 2000, and / or control system 2060) implements a method as described herein (e.g., Figures 21A to 21C the method shown). For example, the computing system can be part of an Internet of Things (IoT) system having a device with sensors for observing an environment (e.g., a manufacturing or construction environment) and for exchanging data (e.g., feedback from the environment) with other devices or systems via the Internet.

[0241] Figure 21A An example flowchart illustrating a method 2100 for training a machine learning model (e.g., machine learning model 2050). The method 2100 includes an operation 2101 for receiving training data (e.g., training data 2032). The training data includes process data generated from measurements of a first bonding operation. The training data includes the states of multiple wires after bonding to a first set of surfaces. Each state of the states includes one or more candidate results of a target related to detecting one or more anomalies in the first bonding operation. For example, the machine learning model can be trained to predict aspects related to the quality of the bonded parts produced in the bonding operation. In this case, one or more candidate results can include one or more quality assurance tests for individual wires or bonded parts in the first bonding operation. As another example, the machine learning model can be trained to predict aspects related to normal or defective products resulting from the bonding operation. In this case, one or more candidate results can include one or more defective chip results of integrated circuit chips in the first bonding operation. In cases where the bonding operation includes multiple bonded parts (e.g., pin bonds and ball bonds), predictions can be made for individual bonded parts (e.g., pin bonds), components (e.g., wires), or products (e.g., chips) based on measurements of the parameters of the pin bonds and / or ball bonds. For example, in a case where wires are bonded at both ends to form a single path, there can be an advantage in including the parameters of the two bonded parts in the training data.

[0242] Method 2100 includes operation 2102 of generating one or more weights for process data such that the process data input to a machine learning model predicts one or more candidate outcomes of a target. For example, weighting can be used within the model to make a particular variable or observation have a greater impact on the algorithm. For example, one or more weights for the process data can be generated for a gradient boosting model for training data (e.g., assigning weight 0 to excluded variables and the highest weight to the most important variables).

[0243] Figure 21B is a flowchart illustrating an example method 2130 for controlling a bonding operation. Method 2130 includes operation 2131 of accessing a machine learning model trained on training data for a first bonding operation (e.g., a machine learning model generated according to method 2100). The first bonding operation includes an operation of bonding a first set of multiple wires to a first set of surfaces. The machine learning model is trained by supervised learning.

[0244] Method 2130 includes operation 2132 of receiving input data indicative of process data generated from measurements of a second bonding operation. The second bonding operation includes an operation of bonding a second set of multiple wires to a second set of surfaces. The second set of multiple wires is different from the first set of multiple wires. The second set of surfaces is different from the first set of surfaces.

[0245] For example, the input data can include real-time sensor measurements received during the second bonding operation (e.g., regarding a given wire of an integrated computer chip). The sensor measurements can enable supervision of the bonding process quality without human supervision or disrupting the bond. However, the input data can also include operator observations. The sensor measurements can include measurements of the bonding system involved in the second bonding operation. For example, by way of a few examples only, the measurements can include heat measurements, power-related measurements, force measurements, discharge ball bond (EFO) measurements, ultrasonic measurements, ball pin-related measurements, and ball bonding-related measurements.

[0246] Method 2130 includes an operation 2133 for weighting input data according to a machine learning model. For example, a model developed on training data can be used to predict anomalies in subsequent bonding operations. Method 2130 includes an operation 2134 for generating an anomaly prediction quantity indicating the risk of occurrence of an anomaly in a second bonding operation based on weighting the input data according to the machine learning model. For example, the anomaly prediction quantity can be a defect of a component of a bonded part (e.g., a defective bonded part between a wire and a lead frame or a die), or a prediction quantity of a defective product (e.g., due to abnormal process conditions such as floating conditions, material contamination (e.g., lead frame contamination), and die tilt). Method 2130 includes an operation 2135 for outputting the anomaly prediction quantity as needed to control the second bonding operation. For example, the anomaly prediction quantity can control the second bonding operation to correct one or more anomalies in the bonding system involved in the second bonding operation or reduce the occurrence of the one or more anomalies (e.g., indicating the likelihood of defects in a batch of products to alert further investigation of the batch or change the process for future batches).

[0247] In one or more embodiments, the anomaly prediction quantity predicts an anomaly in an individual wire or bonded part in the second bonding operation without performing a destructive quality assurance test on the individual wire or bonded part. Thus, the bonded part does not need to be destroyed until there is an indication of a defect in the bonded part of the product. For example, the bonding operation described herein can include a ball bonding operation, and the destructive quality assurance test includes a ball shear test for testing the ball bond, which will cut off the ball bond. As another example, the bonding operation described herein can include a stitch bonding operation, and wherein the destructive quality assurance test includes a stitch pull test for testing the stitch bond, which will pull out the stitch bond.

[0248] In one or more embodiments, the prediction algorithm can be improved during subsequent bonding operations. Figure 21C is a flowchart illustrating an example method 2160 for updating a machine learning model used to control a bonding operation.

[0249] Method 2160 includes an operation 2161 for receiving feedback indicating whether the anomaly prediction quantity correctly or incorrectly predicts an anomaly in a specific chip manufactured in the second bonding operation. Method 2160 includes an operation 2162 for updating the machine learning model based on the feedback (e.g., weighting the input data differently or selecting different or additional input data types).

[0250] Method 2160 optionally includes an operation 2163 for adjusting a bonding operation. For example, when a second bonding operation is performed by a chip manufacturing system, operation 2163 includes adjusting a bonding operation subsequent to the second bonding operation performed by the chip manufacturing system based on one or more of the following: an anomaly prediction measure indicative of a risk of an anomaly occurring in the second bonding operation; and feedback indicative of whether the anomaly prediction measure correctly or incorrectly predicts the occurrence of an anomaly in a particular chip manufactured in the second bonding operation. For example, the bonding operation may be stopped based on a high percentage indicating a suspected error (e.g., to inspect process components or the manufactured product). Different materials or forces for the bonding material may be used based on consistently indicated problems. Alternatively, a particular product may be positioned for correction of a suspected problematic product. For example, in one or more embodiments, the input data includes received sensor measurements tagged with raw information indicating one or more of an identifier or location of a particular wire, die, or chip involved in the second bonding operation. The anomaly prediction measure identifies an anomaly occurring in the second bonding operation and is correlated with the raw information to indicate the location of the anomaly. This can be useful for correcting individual defective chips or processes.

[0251] For example, Figure 22A is a graph 2200 illustrating the relationship between the motion signature pattern and the position of corresponding chips on a leadframe. The waveforms shown relate to successive process measurements (e.g., force or power measurements) for each wire from 0 to n in a bonding operation. The position of the waveforms in the graph is correlated with the layout of an integrated circuit manufactured in a facility. Figure 22B Illustrates corresponding to Figure 22A An example position of the chips on the leadframe of the graph of. In this example, Figure 22B the leadframe in has 16 dies (four per vertical position identifier). The horizontal position identifier and wire number are correlated with the particular wire positions (e.g., HID 1 to HID 2, W1 to Wn) in FIG. 2250. The vertical position identifier (VID) is correlated with the particular chips (e.g., VID 1 to VID 4) in FIG. 2250. A graphical user interface may be used to display the graph and the figure (e.g., graphs 2200 and 2250) to assist in developing the model (e.g., receiving user indications of important variables for building the model) and observing real-time process behavior for correcting defective chips or processes. A particular wire may have a keyword or identifier corresponding to the leadframe identifier, vertical position identifier, horizontal identifier, and / or wire number. In this example, there is significant variation across wire numbers, but for a given column number, the process measurements exhibit a similar pattern.

[0252] Figure 23AIllustrate an example of quality assurance (QA) data 2301 that compares the training data 2302 and test data 2303 of a predictive model for destructive testing of joints formed during a bonding operation in table 2300.

[0253] In one or more embodiments, process data (subsequently, input data or training data) includes derived data, such as generated values indicating the median or average of measurements related to multiple wires bonded in a particular chip during a second bonding operation, and a set of generated deviations including the deviations of the values of each of the multiple wires.

[0254] Figure 23B Illustrate an example method for generating derived process data. In this example, for each of the considered factors for a machine learning model (e.g., the maximum measured impedance of a wire), data table 2350 shows that one or more derived values are calculated for each selected wire number 2351 and corresponding horizontal position identifier number 2352. All wires of the bonding operation can be considered or selectively considered. For example, the measurement data can be limited to the wires with the largest numbers of target variables to be considered (e.g., the most normal operation or the worst quality). In this example data table 2350, N is the total number of wires on a particular chip.

[0255] In this example, the first derived value is calculated as the median normal value 2353 of the measurements to be considered. The median normal value is calculated by taking the median of the measurements at a given unique wire position for the wires that pass the quality test (e.g., for the measurements of the horizontal position identifier 2352 and wire number 2351 in Figure 22A and 23B . The second derived value (e.g., difference value (normal) 2354) is calculated by subtracting the median from the actual measurement for a particular wire and position identifier. Other derived values, such as the abnormal difference value 2355, can be calculated based on the median including the wires that exhibit abnormal behavior for the quality test.

[0256] Additionally or alternatively, the input data includes derived data for the product considering the deviation group (or at the chip level considering the deviations of the bonded wires of the chip), such as generated metrics. Figure 24 Is an example chart 2400 for deriving process data associated with a bonding operating system at the chip level. In this example, one or more metrics can be calculated. For example, when taking one variable at a time and using its bonding result to indicate the wires that pass the quality, the positive sum, negative sum, and number of zero crossings are used to distinguish normal and abnormal conditions. The positive sum value is calculated by averaging the area above the median, i.e., where Del i > 0. The negative sum value is calculated by averaging the area below the median, i.e., where Del i < 0. N is the total number of wires considered on the chip. The total number of zero crossings (e.g., zero crossing 2402) is 4. A model can be trained on these features to identify abnormal conditions. Fewer, less, or different chip metrics (e.g., maximum, minimum, or average chip manufacturing process measurements) can be considered. In this example, analysis is performed at the chip level to identify the nature of the abnormal condition. The derived process data is generally related to a specific chip of the integrated circuit chip and is derived from measurement data of wires that are associated with the specific chip and are joined in a second bonding operation. The measurement and anomaly detection can be performed at some other granularity (e.g., at the wire or die level).

[0257] In one or more embodiments, process data is generated from the derived data, the derived data including generated singular data values related to multiple different types of measurements and related to the same wire. For example, measurements in a bonding operation can include voltage (V) and current (I) measurements of the wire, but a resistance (R) value can be derived (e.g., according to the equation V = IR). Thus, this one value can consider the relationship between different measurement types.

[0258] Thus, in one or more embodiments, training data including process data can be generated from derived data including generated singular data values that can be related to multiple different types of measurements and related to the same wire in a first bonding operation. Input data indicating process data generated from measurements in a second or subsequent bonding operation can include generated singular data values related to multiple different types of measurements and related to the same wire in the second bonding operation.

[0259] Additionally or alternatively, training data including process data can be generated by deriving information that considers the relationship between measurement types in a first bonding operation from multiple different measurement types. Input data indicating process data generated from measurements in a second bonding operation can include information derived from multiple measurement types that considers the relationship between measurement types in the second bonding operation.

[0260] In one or more embodiments, process data can be used to predict information about quality prediction metrics for a bonding operation. For example, measurements in the first and second operations can include measurements associated with the process of forming a stitch bond and / or the process of forming a ball bond.

[0261] Figure 25 is a graph 2500 illustrating the correspondence between the predicted ball shear value modeled according to this embodiment and the actual ball shear value obtained due to a destructive test.

[0262] Specification limits are used for the final acceptance or rejection of the product. In this example, the dividing line 2510 represents the lower limit of the specification limits for the ball shear test measurements (e.g., in grams) and the dividing line 2512 represents the upper limit. Control limits that are more stringent than the specification limits can be used. Control limits are used for process control purposes to ensure that the product is produced within the specification limits. In this example, the dividing line 2520 represents the lower limit of the control limits for the ball shear test measurements and the dividing line 2522 represents the upper limit. The dividing lines have corresponding values on the predicted ball shear axis 2530 (e.g., on the two axes of the graph 2500, the lower specification limit can be at A, the lower control limit can be at B, the upper control limit can be at C, and the upper specification limit can be at D). The test results of each of the ball shear destructive tests are plotted against the plotted predicted values of the ball shear for different object targets identified by keywords including different normal and abnormal conditions (material contamination (MC), old cover (OC), and floating condition (FC)).

[0263] Evaluate the accuracy of the model. TP represents true positive and indicates a poor quality joint identified as poor by the model. TN represents true negative and indicates a good quality joint identified as good by the model. As shown in Table 2550, the graph shows mostly correct predictions (291 TPs and 11,922 TNs). FN represents false negative and indicates a poor quality joint identified as good by the model. FP represents false positive and indicates a good quality joint identified as poor by the model. There are only a small number of incorrect predictions (27 FPs and 40 FNs).

[0264] Some metrics are more important for different industries. For example, false negatives can be important in integrated circuit construction because these types of joints can lead to an increase in field returns and warranty claims. Therefore, benign false negatives (B_FN) are also considered. These represent joints that have a joint strength between the lower (upper) specification limit and the lower (upper) control limit. These false negatives do not have poor quality with respect to the specification limits. Most false negatives are benign false negatives (32 out of 40). Statistical tools like precision, recall, and f1 can be used to evaluate the model. For example, precision can be calculated as the number of true positives divided by the sum of the number of true positives and the number of false positives:

[0265]

[0266] Recall can be calculated as the number of true positives divided by the sum of the number of true positives and the number of false negatives:

[0267]

[0268] F1 is the harmonic mean of precision and recall. For example, it can be calculated as:

[0269]

[0270] In this example, for a prediction with a precision of 0.91509, a recall of 0.8795, and an F1 of 0.89676, the model performs well. In this case, by using the prediction, the user can potentially avoid performing destructive quality tests and closely understand the integrated circuit results. If 32 benign false negatives are removed from the calculation, a higher model accuracy can be achieved.

[0271] Figure 26 is a graph showing the correspondence between the predicted pin pull values modeled according to this embodiment and the actual pin pull values obtained due to destructive testing. In this case, there is only one target value. The dividing line 2610 represents the lower limit of the specification limit of the pin pull test measurement value (e.g., in grams). The upper limit is not shown on the graph because it is much higher than the plotted values. The dividing line 2620 represents the lower limit of the control limit of the pin pull test measurement value and the dividing line 2622 represents the upper limit. The dividing lines have corresponding values on the predicted pin pull axis 2530 (e.g., on both axes of the graph 2600, the specification lower limit is at E, the control lower limit is at F, the control upper limit is at G, and the specification upper limit is outside the graph). For all accurate predictions as shown in Table 2650 except for 7 outlier leads (which are excluded due to issues with a precision, recall, and F1 of 1 in the pin pull test), this model also performs well. In addition to this model being an accurate model without destructive testing, the predictive ability of the models described herein is not as susceptible to human error as occurs in manual testing.

[0272] In Figure 25 and 26 example, a gradient boosting model consisting of multiple decision trees is used. The predictive model defines the relationship between the input variables and the target variable. The purpose of the predictive model is to predict the target value from the input. The model is built by using training data where the target value is known. Then, the model can be applied to observations where the target is unknown. If the prediction fits the new data well, the model is said to generalize well. Good generalization is the main goal of the prediction task. The predictive model may fit the training data well but not generalize well.

[0273] Decision trees are a type of predictive model that has been independently developed in the statistical and artificial intelligence communities. The gradient boosting process can be used to build a predictive model by fitting a set of additive trees (e.g., using GRADBOOST-related software provided by SAS Institute Inc., Cary, NC).

[0274] The quality of a prediction model can depend on the values of various options that govern the training process; these options are referred to as hyperparameters. The default values of these hyperparameters may not be suitable for all applications, and the only combination of values of these hyperparameters can be selected to minimize an objective function or a certain accuracy metric. Thus, other or additional methods for building models described herein can be used. For example, different models can be generated differently for the same type and one of the trained models can be selected using a validation process (e.g., using k-fold cross-validation). Additionally, other objective values beyond the quality assurance goal can be used.

[0275] For example, there can be various abnormal conditions in a wire bonding machine. These abnormal conditions can shift the distribution of measured features or quality assurance values (e.g., related to ball shear and stitch pull) or they may also result in defective wires. Early identification of abnormal conditions is necessary so that appropriate corrective actions can be taken and the number of poor-quality products can be reduced. Normal target variables can be related to normal operation, such as a stable process, in-control state, conformance and / or exceedance of specification ratios, and an indicator of the absence of any systematic disturbance cause. Abnormal conditions can be predicted based on detected anomalies in process data, such as die tilt (DT), material contamination (MC), old cap (OC), and floating condition (FC).

[0276] For example, in the case of die tilt, the die attach process can be used for the connection between the die, device, and the rest in a system in an electronic package. During this process, the die moves under the action of capillary forces induced by the liquidus solder. This die tilt phenomenon often occurs and severely deteriorates the reliability and performance of the device. Predictive measurements can be used to understand the quality of the joints from the liquidus solder. In the floating condition, there may be non-bonding due to die floating.

[0277] One type of material contamination can be lead frame contamination, where there are surface organics, organic compounds, and / or residues on the lead frame. This can also be related to quality problems in bonding.

[0278] Early action can lead to an increase in process yield without these abnormalities. For example, using the data described Figures 25 to 26 The model accuracy is significant. The following table shows the mean absolute percentage error (MAPE), which is also known as the mean absolute percentage deviation (MAPD), a measure of the prediction accuracy of observations from the production of integrated circuit chips from 5 batches.

[0279] Observation Target Variable MAPE 1 Normal 3.64% 2 Normal 3.74% 3 Abnormal_FC 3.62% 5 Normal 3.72% 6 Abnormal_MC 12.22%

[0280] Table 1

[0281] In the example of Table 1, the overall mapping of the different target variables is 4.39%. Observation 6 has four outliers that result in a large MAPE. If these points are excluded, then the MAPE of Observation 6 is reduced to 5.67% and the overall MAPE drops to 3.84%.

[0282] The model for detecting abnormal conditions uses the gradient boosting method on chip-level features. These features are extracted using the chip-level waveforms of sensor variables (or motion features). The waveform has wire numbers on the x-axis and motion feature values on the y-axis. In this case, the gradient boosting model has a nominal target. In other cases, an interval target value can be used. In this example, the abnormal target variables are related to floating conditions (FC) and material contamination (MC). Other types of abnormal conditions, such as die tilt, can also be predicted.

[0283] Figure 27A is a functional block diagram of a stack illustrating an event stream processing (ESP) system 2700 configured for streaming and edge analysis. The ESP component 2702 may include one or more ESP engines as described herein (e.g., in Figure 8 ). For example, the ESP component can be used to provide one or more of determining the data quality of a manufacturing execution system (MES), preprocessing, AI scoring, prediction, and abnormal prediction amounts, real-time regulatory alerts, in-stream training, federated AI, deploying customer models, and operating systems. Thus, the ESP component 2702 can exchange information with the client service 2704 devices and systems and the data management system 2706. The model manager 2708 can be used to manage, develop, and update the models described herein based on data received (e.g., by the client service 2704 and / or the ESP component 2702). The artificial intelligence and machine learning component 2710 can be used to train and update the models described herein (e.g., as directed by the model manager 2708). Gradient boosting is given as an example method for constructing the model. However, additional or different model methods (e.g., random forests, linear regression, and neural networks) can be employed. The model manager 2708 and the artificial intelligence and machine learning component 2710 can exchange information with the client service 2704 and the data management 2706. Thus, the embodiments described herein (e.g., Figure 27A system) can be used to develop an analytical model to predict wire bond quality, including moving from random testing (which typically can involve destructive testing) to model-recommended risk-based calibration testing. The developed model can provide an early indication of the presence of abnormal conditions (or faults) in the manufacturing process and has a lower rejection rate and increased throughput. Edge processing can provide on-site guidance.

[0284] Figure 27B is a flowchart illustrating an example method 2750 for generating a machine learning model according to an embodiment of the present disclosure. AsFigure 27B As seen in Figure 27B , method 2750 includes operations (block 2752) for developing an extract, transform, and load (ETL) process to read input data and create a data set for generating a machine learning model. In one embodiment, for example, the ETL process includes one or more encoded script files that extract and combine data sets from the input data when executed by the processing circuitry of a computer. The input data read by the ETL process can be in any form desired or expected, but in one embodiment, is formatted as a.csv file and includes data values measured during the wire bonding stage ( Figure 13 block 1308 in Figure 13 ).

[0285] Next, method 2750 performs one or more quality checks on the input data (block 2754) to ensure that the input data and data set have sufficient quality to generate a machine learning model. For example, as stated above, a wire bonder that performs the ball and stitch bonding process ( Figure 13 block 1308 in Figure 13 ) is configured to measure different operating parameters while performing the process. The parameters are measured for each "run" of the subsequent semiconductor manufacturing process, where each run includes a plurality of parameters representing a "lot".

[0286] According to the present disclosure, input data is obtained in batches, and each batch is analyzed to determine the quality of its data. For example, at least one aspect of this embodiment identifies several different variables, variable names, and variable distributions in each batch. The data is then compared to data associated with one or more previously obtained batches (or a baseline batch with known good data) to determine if there are any significant changes in the values of the parameters (e.g., the values are not within a predefined tolerance range or are significantly different from the baseline values). For example, such significant changes in the data may indicate that the input data is not suitable for generating a model.

[0287] However, even in cases where data may not be available, it is still beneficial to determine that there are significant differences between parameter values according to the present disclosure. For example, die tilt and machine reading errors are some common examples of things that can cause unavailable parameter values to be measured and collected. However, knowing that these errors exist can enable the operator to more quickly identify and correct the problems that led to the recorded input data.

[0288] As another example, the quality check may indicate that the range of measured values does not adequately reflect all values that can be generated during actual operation. With this information, the operator can change the monitoring / measurement process for the wire bonding stage to ensure that the appropriate operating range is adequately covered. Ensuring that the monitoring and measurement process adequately covers the operating range actually experienced during production will help ensure that the values of the measured parameters are accurate and suitable for generating the data set and machine learning model.

[0289] According to the present disclosure, any of the parameters related to the operations during the wire bonding stage can be monitored and measured. However, in one embodiment of the present disclosure, one or more of the following parameters are measured to determine an anomaly related to one or more of the abnormal conditions (such as floating conditions and lead frame contamination).

[0290] If the input data is considered to have sufficient quality, then method 2750 transforms and / or creates one or more variables (block 2756). For example, one or more of the coding script files that make up the ETL processing can be configured to transform variables in the input data (e.g., between units, scale values up or down, etc.). Additionally or alternatively, one or more of the coding script files that make up the ETL processing can calculate one or more derived variables from the input data. In one example, the transformed variables and / or the calculated derived variables are used by this embodiment to generate one or more modeling data sets. The training data set and / or the test data set can also be generated by this embodiment using the transformed and / or calculated derived variables.

[0291] Additionally or alternatively, embodiments of the present disclosure perform an analysis on the input data to obtain the distributions of the predicted and target variables, as well as one or more different plots and tables of the data indicating various statistical data. In at least one embodiment, a correlation analysis is performed on the input data to determine whether there is any correlation between the previously recorded "baseline" parameter values and the parameter values generated by one or more sensors that measure the wire bonding operation. Such analysis provides a better understanding of the input data and is very beneficial for generating machine learning models.

[0292] In the case of appropriately processing the input data, method 2750 then generates a machine learning model for predicting wire quality parameters (block 2758). These models include one or more prediction models for estimating the results of a destructive ball shear test (i.e., the force required to shear the spherical bond 1510 from the bond pad 1506) and the results of a stitch pull test (i.e., the force required to separate the wire 1504 from the target pad 1508). Any of a variety of techniques can be utilized to generate the prediction models, including but not limited to linear regression techniques, random forest techniques, gradient boosting techniques, and neural network techniques. In one example embodiment, using the gradient boosting technique to generate the prediction model provides the most accurate results. In this embodiment, automatic tuning using K-fold cross-validation is used to tune the model hyperparameters.

[0293] Method 2750 performs model stability checks to help ensure that the model provides accurate predictions across different data conditions (block 2760). In this step, the stability of the model hyperparameters is checked after the model has been trained. For example, in one embodiment, the tuned hyperparameters are used for multiple random training / test splits (i.e., the number of "runs"). The hyperparameter values can be used to train the training split and then score the test split. The performance of the model is determined on each test split. To accomplish this, one embodiment of the present disclosure evaluates a set of predetermined model performance statistics. Some examples of performance statistics include, but are not limited to, the F1 ratio and mean squared error.

[0294] In addition, in some embodiments of the present disclosure, a performance evaluation summary table is generated. The summary table provides a performance evaluation using the performance statistics across all test splits. This table can have any structure desired or required, but in at least one embodiment, this table has one or more of the following data elements (e.g., table columns): performance statistics; number of runs; minimum value of a particular performance metric; maximum value and / or average value.

[0295] Once the model is complete, it can be used to estimate the presence or absence of an abnormal condition. For example, the presence or absence of an abnormal condition can be determined based on the observation score (block 2762) through the edge device. In some embodiments, the model will provide information indicating the specific abnormal condition that is present. Figure 27B An example method showing stages including one or more operations is presented. Those skilled in the art will understand that the method can have fewer or more stages or be performed in a different order or recursively (e.g., to keep improving the model).

[0296] In one or more embodiments, several different measurements can be observed, and the computing system receives training data and / or input data by selectively selecting a subset of the types of parameters observed in the engagement operation. The measurements used to derive the model can be the measurements for the subset of the types of parameters. Figure 28 is a functional block diagram of a computer program product. In one embodiment, the computer program product includes a control program, for example, executed by the processing circuitry of a computing device.

[0297] More specifically, Figure 28 The processing circuitry 2900 of the computing device is illustrated, as well as the units / modules it executes. The various units / modules can be implemented by hardware and / or by software code executed by a processor or processing circuitry. In this embodiment, the units / modules include a machine learning model access unit / module 2910, an input data receiving unit / module 2920, an input data weighting unit / module 2930, an anomaly indicator generating unit / module 2940, and an anomaly indicator output unit / module 2950.

[0298] In this embodiment, the machine learning model access unit / module 2910 configures the processing circuitry 2900 to access one or more machine learning models trained on training data of a first bonding operation. In at least one embodiment, the first bonding operation is a bonding operation in which a first set of multiple wires is bonded to a first set of surfaces (e.g., bonding a ball and a pin-type bond to a bonding pad and a target pad, respectively). Additionally, the machine learning models are trained by supervised learning. For example, in some embodiments, the supervised training includes receiving training data. The training data includes process data generated from measurements of the first bonding operation and the states of the multiple wires after being bonded to the first set of surfaces. Each state includes one or more candidate results of a target related to detecting one or more anomalies in the first bonding operation. Additionally, one or more weights of the process data are generated such that the process data input to the machine learning model predicts the one or more candidate results of the target.

[0299] The input data receiving unit / module 2920 configures the processing circuitry 2900 to receive input data indicating process data generated from measurements of a second bonding operation in which a second set of multiple wires is bonded to a second set of surfaces (e.g., bonding a ball and a pin-type bond to a bonding pad and a target pad, respectively). The second set of multiple wires is different from the first set of multiple wires, and the second set of surfaces is different from the first set of surfaces.

[0300] The input data weighting unit / module 2930 configures the processing circuitry 2900 to weight the input data according to the machine learning model.

[0301] The anomaly indicator generating unit / module 2940 configures the processing circuitry 2900 to generate an anomaly prediction indicating the risk of occurrence of an anomaly in the second bonding operation based on weighting the input data according to the machine learning model.

[0302] The anomaly indicator output unit / module 2950 configures the processing circuitry 2900 to output the anomaly prediction to control the second bonding operation.

Claims

1. A computer program product tangibly embodied in a non-transitory machine-readable storage medium, the computer program product comprising instructions that are operative to cause a computing system to: Access a machine learning model trained on training data of a first bonding operation, wherein the first bonding operation includes an operation for bonding a plurality of wires of a first group to a first group of surfaces, and wherein the machine learning model is trained by supervised learning, including: Receiving the training data, wherein the training data includes: Process data generated from measurements of the first bonding operation; and The states of the plurality of wires after being bonded to the first group of surfaces, each state of the states including one or more candidate results of a target related to detecting one or more anomalies in the first bonding operation; and Generating one or more weights for the process data such that the process data input to the machine learning model predicts the one or more candidate results of the target; Receiving input data indicating process data generated from measurements of a second bonding operation, wherein the second bonding operation includes an operation for bonding a plurality of wires of a second group to a second group of surfaces, wherein the plurality of wires of the second group are different from the plurality of wires of the first group, and wherein the second group of surfaces is different from the first group of surfaces; Weighting the input data according to the machine learning model; Generating an anomaly prediction indicating a risk of occurrence of an anomaly in the second bonding operation based on weighting the input data according to the machine learning model; and Outputting the anomaly prediction to control the second bonding operation.

2. The computer program product according to claim 1, wherein the one or more candidate results include one or more destructive quality assurance tests for individual wires or bondings in the first bonding operation; and wherein the anomaly prediction is a prediction of an anomaly in an individual wire or bonding in the second bonding operation without performing a destructive quality assurance test on the individual wire or bonding.

3. The computer program product according to claim 2, wherein the first bonding operation includes one of the following: A ball bonding operation, and wherein the destructive quality assurance test includes a ball shear test for testing a ball bonding; and A stitch bonding operation, and wherein the destructive quality assurance test includes a stitch pull test for testing a stitch bonding.

4. The computer program product according to claim 1, wherein the first bonding operation bonds the wires of the plurality of wires of the first group to corresponding surfaces of the first group of surfaces to form an integrated circuit chip; wherein the one or more candidate results include one or more defective chip results of the integrated circuit chip in the first bonding operation; and wherein the instructions are operative to cause the computing system to generate the anomaly prediction of the risk of the anomaly in the integrated circuit chip manufacturing process in the second bonding operation.

5. The computer program product according to claim 1, wherein the one or more anomalies are associated with one or more of the following: Floating condition; Lead frame contamination; and Die tilt.

6. The computer program product according to claim 1, wherein the second bonding operation bonds wires of the second set of wires to form an integrated circuit chip; and wherein the process data generated from the measurements of the second bonding operation is generally related to a specific chip of the integrated circuit chip and is derived from measurement data of the wires bonded in the second bonding operation and associated with the specific chip.

7. The computer program product according to claim 1, wherein the input data includes real-time sensor measurements received during the second bonding operation; and wherein for a given wire of the second set of wires, the real-time sensor measurements include one or more of the following: Heat measurement; Power measurement; Force measurement; Discharge ball bond (EFO) measurement; and Ultrasonic measurement.

8. The computer program product according to claim 1, wherein the input data includes real-time sensor measurements received during the second bonding operation; wherein the real-time sensor measurements include measurements of the bonding system involved in the second bonding operation; and wherein the anomaly prediction controls the second bonding operation to correct the one or more anomalies in the bonding system involved in the second bonding operation or reduce the occurrence of the one or more anomalies.

9. The computer program product according to claim 1, wherein the input data includes received sensor measurements labeled with original information indicating one or more of an identifier or location of a specific wire, die, or chip involved in the second bonding operation; and wherein the anomaly prediction identifies the anomalies occurring in the second bonding operation and is associated with the original information to indicate the location of the anomalies.

10. The computer program product according to claim 1, wherein the instructions are operable to cause the computing system to: Receive feedback indicating whether the anomaly prediction correctly or incorrectly predicts the anomalies in a specific chip manufactured in the second bonding operation; and Update the machine learning model based on the feedback.

11. The computer program product according to claim 1, wherein the second bonding operation is performed by a chip manufacturing system; and wherein the instructions are operable to cause the computing system to adjust a bonding operation subsequent to the second bonding operation performed by the chip manufacturing system based on one or more of the following: The anomaly prediction indicating the risk of occurrence of the anomalies in the second bonding operation; and Feedback indicating whether the anomaly prediction correctly or incorrectly predicts the occurrence of the anomalies in a specific chip manufactured in the second bonding operation.

12. The computer program product according to claim 1, wherein the input data contains derived data including one or more of the following: Generated values indicating the median or average of measurements related to wires bonded in a specific chip in the second bonding operation; A set of generated deviations including the deviation of the value of each of the wires; and Consider the generated metrics of the set of generated-biased chips.

13. The computer program product according to claim 1, wherein the training data includes the process data generated from the derived data, the derived data including generated singular data values related to multiple different types of measurements and related to the same wire in the first bonding operation; and wherein the input data indicating the process data generated from the measurements of the second bonding operation includes generated singular data values, wherein the generated singular data values are related to the multiple different types of measurements and related to the same wire in the second bonding operation.

14. The computer program product according to claim 1, wherein the training data includes the process data generated by deriving information considering the relationships between the measurement types in the first bonding operation from multiple different measurement types; and wherein the instructions are operative to cause the computing system to receive the input data indicating the process data generated from the measurements of the second bonding operation by deriving information considering the relationships between the measurement types in the second bonding operation from multiple measurement types.

15. The computer program product according to claim 1, wherein the instructions are operative to cause the computing system to receive the training data by selectively picking a subset of the types of parameters observed when performing the process of bonding the first set of multiple wires to the first set of surfaces in the first bonding operation; and wherein the measurements of the first bonding operation are measurements for the subset of the types of parameters.

16. The computer program product according to claim 1, wherein the one or more weights of the process data are generated for a gradient boosting model of the training data.

17. The computer program product according to claim 1, wherein the machine learning model is further trained by multiple generated machine learning models and selected based on k-fold cross-validation.

18. The computer program product according to claim 1, wherein the measurements of the first and second bonding operations include measurements captured when performing the process of forming a ball bond.

19. The computer program product according to claim 1, wherein the measurements of the first and second bonding operations include measurements captured when performing the process of forming a stitch bond.

20. The computer program product according to claim 1, wherein the anomaly prediction is a prediction of a defective bond between the wires of the second set of multiple wires and the lead frame or die of the second set of surfaces.

21. The computer program product according to claim 1, wherein the measurements of the first bonding operation include measurements captured when performing the process of bonding the first set of multiple wires to the first set of surfaces; wherein the state includes multiple candidate results of the target; and wherein the measurements of the second bonding operation include measurements captured when performing the process of bonding the second set of multiple wires to the second set of surfaces.

22. A computer-implemented method, comprising: Access a machine learning model trained on training data of a first bonding operation, where the first bonding operation includes an operation for bonding a plurality of wires in a first group to a first group of surfaces, and where the machine learning model is trained by supervised learning, including: Receive the training data, where the training data includes: Process data generated from measurements of the first bonding operation; and The states of the plurality of wires after being bonded to the first group of surfaces, each state of the states including one or more candidate results of a target related to detecting one or more abnormalities in the first bonding operation; and Generate one or more weights for the process data such that the process data input to the machine learning model predicts the one or more candidate results of the target; Receive input data indicating process data generated from measurements of a second bonding operation, where the second bonding operation includes an operation for bonding a plurality of wires in a second group to a second group of surfaces, where the plurality of wires in the second group are different from the plurality of wires in the first group, and where the second group of surfaces is different from the first group of surfaces; Weight the input data according to the machine learning model; Generate an anomaly prediction quantity indicating the risk of an anomaly occurring in the second bonding operation based on weighting the input data according to the machine learning model; and Output the anomaly prediction quantity to control the second bonding operation.

23. The computer-implemented method according to claim 22, where the one or more candidate results include one or more destructive quality assurance tests on individual wires or joints in the first bonding operation; and where the anomaly prediction quantity is a prediction of an anomaly in an individual wire or joint in the second bonding operation, without performing a destructive quality assurance test on the individual wire or joint.

24. The computer-implemented method according to claim 22, where the first bonding operation bonds the wires of the first group of wires to form an integrated circuit chip; where the one or more candidate results include one or more defective chip results of the integrated circuit chip in the first bonding operation; and where generating the anomaly prediction quantity includes generating the anomaly prediction quantity of the risk of the anomaly in the integrated circuit chip manufactured in the second bonding operation.

25. The computer-implemented method according to claim 22, where the second bonding operation bonds the wires of the second group of wires to form an integrated circuit chip; and where the process data generated from the measurements of the second bonding operation is overall related to a specific chip of the integrated circuit chip and is derived from measurement data of the wires associated with the specific chip and bonded in the second bonding operation.

26. The computer-implemented method according to claim 22, further comprising receiving real-time sensor measurements during the second bonding operation; and where for a given wire of the plurality of wires in the second group, the real-time sensor measurements include one or more of the following: Heat measurement values; Power measurement values; Force measurement values; Electrical Discharge Orb (EFO) measurements; and Ultrasonic measurements.

27. The computer-implemented method according to claim 22, further comprising: Receiving feedback indicating whether the anomaly prediction correctly or incorrectly predicts the anomaly in a particular chip manufactured in the second bonding operation; And Updating the machine learning model based on the feedback.

28. The computer-implemented method according to claim 22, wherein the measurements of the first bonding operation include measurements captured during a process of bonding a plurality of wires of the first group to the first group of surfaces; wherein the state includes a plurality of candidate results of the target; and wherein the measurements of the second bonding operation include measurements captured during a process of bonding a plurality of wires of the second group to the second group of surfaces.

29. A computing device comprising a processor and a memory containing instructions executable by the processor, wherein the computing device is configured to: Access a machine learning model trained on training data of a first bonding operation, wherein the first bonding operation includes an operation of bonding a plurality of wires of a first group to a first group of surfaces, and wherein the machine learning model is trained by supervised learning, including: Receiving the training data, wherein the training data includes: Process data generated from measurements of the first bonding operation; and The state of the plurality of wires after being bonded to the first group of surfaces, each state of the state including one or more candidate results of a target related to detecting one or more anomalies in the first bonding operation; and Generating one or more weights for the process data such that the process data input to the machine learning model predicts the one or more candidate results of the target; Receiving input data indicating process data generated from measurements of a second bonding operation, wherein the second bonding operation includes an operation of bonding a plurality of wires of a second group to a second group of surfaces, wherein the plurality of wires of the second group are different from the plurality of wires of the first group, and wherein the second group of surfaces is different from the first group of surfaces; Weighting the input data according to the machine learning model; Generating an anomaly prediction indicating the risk of an anomaly occurring in the second bonding operation based on weighting the input data according to the machine learning model; and Outputting the anomaly prediction to control the second bonding operation.

30. The computing device according to claim 29, wherein the measurements of the first bonding operation include measurements captured during a process of bonding a plurality of wires of the first group to the first group of surfaces; wherein the state includes a plurality of candidate results of the target; and wherein the measurements of the second bonding operation include measurements captured during a process of bonding a plurality of wires of the second group to the second group of surfaces.

Citation Information

Patent Citations

  • Quality prediction using process data

    US11501116B1

  • Methods and Systems for Defects Detection and Classification Using X-rays

    US20210010953A1