Manufacturing Automation Using a Sound Separation Neural Network
By using machine learning technology in complex manufacturing systems, training neural networks to separate task-related signals from acoustic mixed signals, solving the problem of difficulty in accurately detecting abnormal operations in the prior art, and achieving more efficient manufacturing automation.
Patent Information
- Application Number
- CN202080071574.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-10-17
- Filing Date
- 2020-10-14
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2040-10-14
AI Technical Summary
In complex manufacturing systems, especially in systems that are mixed with process manufacturing and discrete manufacturing, it is difficult for the prior art to accurately detect abnormal operations, resulting in degradation of quality, waste of materials and damage to equipment.
Using machine learning technology, the signal generated by the tool performing the task and the signal generated by the actuator of the actuator is separated from the acoustic mixed signal by training a neural network, thereby estimating the task execution status and controlling it.
Accurate detection of abnormal operations in complex manufacturing systems is achieved, the accuracy and efficiency of manufacturing automation is improved, and downtime and material losses are reduced.
Smart Images

Figure CN114556369B_ABST
Abstract
Description
Technical Field
[0001] The present invention generally relates to manufacturing automation using machine learning techniques, and more particularly, to manufacturing automation using a neural network trained to separate signals from acoustic mixtures. Background Art
[0002] In manufacturing where fast and powerful machines can execute complex sequences of operations at high speed, it is very important to monitor and control safety and quality. Deviations from the expected sequence or timing of operations can degrade quality, waste raw materials, cause downtime and equipment damage, and reduce production. The danger to workers is a major concern. For this reason, great care must be taken to carefully design the manufacturing process to minimize unexpected events, and various sensors and emergency switches are also needed to design safety measures into the production line.
[0003] Manufacturing types include process manufacturing and discrete manufacturing. In process manufacturing, the products are usually undifferentiated, such as oil, gas, and salt. Discrete manufacturing produces different items, such as cars, furniture, toys, and airplanes.
[0004] One practical way to increase safety and minimize material and production losses is to detect when the production line is operating abnormally and, if necessary, stop the production line in such cases. One way to implement this method is to use a description of the normal operation of the production line based on the range of measurable variables such as temperature, pressure, etc. to define an allowable operating region and detect operating points outside this region. This method is common in the process manufacturing industry (e.g., oil refining), where the allowable ranges of physical variables are usually well understood, and the quality metrics of product quality are usually directly defined based on these variables.
[0005] However, the nature of the work processes in discrete manufacturing is different from that in process manufacturing, and deviations from the normal work process can have very different characteristics. Discrete manufacturing includes a sequence of operations performed on work cells, such as machining, welding, assembly, etc. Anomalies can include incorrect execution of one or more tasks, or an incorrect order of tasks. Even in abnormal situations, physical variables such as temperature or pressure often do not fall outside the range, so directly monitoring these variables cannot reliably detect these anomalies.
[0006] In addition, complex manufacturing systems can include a combination of process manufacturing and discrete manufacturing. When process manufacturing and discrete manufacturing intermingle on a single production line, the anomaly detection methods designed for different types of manufacturing can be inaccurate. For example, the anomaly detection method for process manufacturing may be designed to detect outliers in the data, while the anomaly detection method for discrete manufacturing may be designed to detect incorrect sequences of operation execution. For this reason, it is natural to design different anomaly detection methods for different classes of manufacturing operations. However, using these separate detection techniques in a complex manufacturing system can become too complex.
[0007] For this reason, systems and methods suitable for anomaly detection in different types of manufacturing systems need to be developed. For example, the method described in U.S. Publication [MERL-3116 15 / 938,411] applies machine learning techniques for anomaly detection in one or a combination of process manufacturing and discrete manufacturing. Using machine learning, the collected data can be used for an automated learning system, where the characteristics of the data can be learned through training. The trained model can detect anomalies in real-time data to achieve predictive maintenance and downtime reduction. However, even with the help of machine learning, the data that needs to be collected to represent some manufacturing operations still makes accurate anomaly detection impractical. SUMMARY OF THE INVENTION
[0008] Some embodiments aim to provide systems and methods for manufacturing automation in complex industrial systems having multiple actuators that actuate one or more tools to perform one or more tasks. Additionally or alternatively, some embodiments aim to use machine learning to estimate the state of execution of these tasks and control the system accordingly.
[0009] Some embodiments are based on the recognition that machine learning can be used for data-driven time series prediction inference of physical systems whose state changes over time, based on data from an unknown underlying dynamic system. For these systems, only the observations related to the state of the system are measured. This can be beneficial for controlling complex industrial systems that are difficult to model.
[0010] However, some embodiments are based on another recognition: that these various machine learning techniques can be used when the observations explicitly represent the state of the system, which can be problematic for some cases. In fact, if the set of observations uniquely corresponds to the state of the system, machine learning methods can be used to design various predictors. However, the sensor data at each time instance may not provide enough information about the actual system state. The number of required observations depends on the dimension d of the system, equal to d for a linear system and 2d + 1 for a non-linear system. If the collected measurements do not include enough observations, the machine learning method will fail.
[0011] Some embodiments are based on the recognition that instead of considering the state of the system that performs the tasks, the state of execution of the tasks themselves can be considered. For example, when the system includes multiple actuators for performing one or more tasks, the state of the system includes the states of all these actuators. However, in some cases, the state of the actuators is not the primary concern for control. In fact, the state of the actuators is needed to guide the execution (e.g., the execution of the task), so the state of execution is the main goal, while the state of the actuators that perform the task is only a secondary goal.
[0012] Some embodiments are based on the understanding that it is natural to equate the state of a system with the state of the execution of a task, because often only the state of the system can be measured or observed, and if enough observations are collected, the state of the system can indeed represent the execution state. However, in some cases, the state of the system is difficult to measure, and the state of the execution of the task is difficult to define.
[0013] For example, consider a computer numerical control (CNC) that uses a cutting tool to machine a workpiece. The state of the system includes the state of the actuator that moves the cutting tool along a tool path. The execution state of the machining is the state of the actual cutting. An industrial CNC system can have many different (sometimes redundant) actuators, and many of the state variables are in complex non-linear relationships with each other, making it difficult to observe the state of the CNC system. However, it can also be difficult to measure the machining state of the workpiece.
[0014] Some embodiments are based on the recognition that the state of the execution of a task can be represented by an acoustic signal generated by such execution. For example, the execution state of the CNC machining of a workpiece can be represented by a vibration signal caused by the deformation of the workpiece during machining. Therefore, if such a vibration signal can be measured, various classification techniques including machine learning methods can be used to analyze such a vibration signal to estimate the task execution state and select appropriate control actions for control execution.
[0015] However, the problem faced by this method is that such a vibration signal does not exist in isolation. For example, in a system including multiple actuators that actuate one or more tools to perform one or more tasks, the signal generated by the tool performing the task is always mixed with the signal generated by the actuator that actuates the tool. For example, the vibration signal generated by the deformation of the workpiece is always mixed with the signal from the motor that moves the cutting tool. If such a vibration signal can be generated syntactically in some way in isolation or captured in some way in isolation, a machine learning system such as a neural network can be trained to extract such a signal. However, in many cases including CNC machining, it is impractical to generate such an isolated signal representing task execution. Similarly, it would also be impractical to separately record different signals with multiple microphones.
[0016] To this end, in order to streamline the manufacturing automation of a system including multiple actuators that actuate one or more tools to perform one or more tasks, it is necessary to separate the source signals from the acoustic mixture of the signals generated by the tool performing the task and the signals generated by the multiple actuators that actuate the tool. Therefore, the aim of some embodiments is to train a neural network for sound separation of the sound sources of a mixed signal in the absence of isolated sound sources. As used herein, at least some of the sound sources in the mixed signal occupy the same time, space, and frequency spectra in the acoustic mixture.
[0017] Some embodiments are based on the recognition that when the sound sources in a sound mixture can be easily isolated and recorded or recognized by a person, for example, supervised learning can be used to provide such training. This type of classification is referred to herein as a strong label. In theory, a person can generate an approximation of this type of label; however, it is largely unrealistic to require a person to precisely label sound activity in both time and frequency to provide a strong label. However, it is realistic to consider some limited labels (weak labels) regarding which sound is active within a certain time range. In general, such weak labels do not require the sound to be active throughout the entire specified range and can occur only at brief instants within that range.
[0018] Some embodiments are based on the recognition that since the structure of the manufacturing process is generally known, the sources that generate the sound signals are also generally known. To this end, weak labels can be provided for the sources of the sound mixtures that represent the operation of the manufacturing system. To this end, some embodiments have developed methods that can learn to separate the sounds in a sound mixture for which only weak-labeled training data is available.
[0019] Accordingly, some embodiments train a neural network to separate from a sound mixture the signal generated by a tool performing a task and the signal generated by an actuator actuating the tool. For example, the neural network is trained to separate different signals from a sound mixture such that each separated signal belongs to a class of signals that exist in the operation of the system, and the separated signals add up to the sound mixture. The weak labels identify the classes of signals that exist in the operation of the system. Each weak label specifies a class of signals that exists at a certain point during the operation.
[0020] In one embodiment, the neural network is jointly trained with a classifier configured to classify the classes of signals identified by the weak labels. For example, the neural network can be jointly trained with the classifier to minimize a loss function that includes a cross-entropy term between the output of the classifier and the output of the neural network. This joint training takes into account the fact that it is difficult to classify signals that do not exist in isolation and allows for end-to-end training of the separation and classification neural network.
[0021] In some implementations, additional constraints are added to the training of the neural network to ensure quality. For example, in one embodiment, the neural network is trained such that when the separated signal of the class identified by the weak label is submitted as input to the neural network, the neural network generates that separated signal as output. Additionally or alternatively, in one embodiment, the neural network is trained such that the separated signals from two or more classes identified by the weak labels are recombined and fed back to the network for re-separation, while an adversarial loss differentiates between the real sound mixture and the synthetic recombined mixture.
[0022] In cases where a neural network has been trained for signal separation, some embodiments use the output of the network to train a state estimator for estimating the task execution state. For example, in one embodiment, the state estimator is trained for signals generated by a tool performing a task and extracted by the neural network from different sound mixtures of different repetitions of system operation. Notably, in some embodiments, individual samples of the extracted signals generated by the tool performing the task define the task execution state, but are not sufficient to define the system operation state. However, the task execution state is sufficient to select an appropriate control action. Thus, embodiments provide dimensionality reduction in manufacturing automation applications.
[0023] For dimensionality reduction, additionally or alternatively, some embodiments allow independent control of different tools performing different operation tasks. For example, when performing an analysis on a signal representing the state of the entire system, such analysis can provide control over the entire system. However, when performing the analysis separately for different tools performing different tasks, independent control of the tools is possible.
[0024] To this end, when multiple tools perform multiple tasks during system operation, some embodiments perform independent control of the tasks. For example, one embodiment controls a system having a first tool performing a first task and a second tool performing a second task. A neural network is trained to separate a first signal generated by the first tool performing the first task and a second signal generated by the second tool performing the second task from the sound mixture. During system operation, the first signal and the second signal are extracted from the sound mixture using the neural network and analyzed independently of each other to estimate a first state of the execution of the first task and a second state of the execution of the second task. The embodiment is configured to perform a first control action selected according to the first state and a second control action selected according to the second state.
[0025] Thus, one embodiment discloses a system for controlling the operation of a machine, the machine including a plurality of actuators assisting one or more tools in performing one or more tasks. The system includes: an input interface configured to receive, during operation of the system, a sound mixture of signals generated by the tools performing the tasks and signals generated by the plurality of actuators actuating the tools; a memory configured to store a neural network trained to separate signals generated by the tools performing the tasks from signals generated by the actuators of the actuating tools from the sound mixture; and a processor configured to submit the sound mixture of signals into the neural network to extract signals generated by the tools performing the tasks from the sound mixture of signals; analyze the extracted signals to generate a state of the execution of the tasks; and perform a control action selected according to the state of the execution of the tasks.
[0026] Another embodiment discloses a method for controlling the operation of a machine that includes a plurality of actuators assisting one or more tools to perform one or more tasks, wherein the method uses a processor coupled with instructions stored for implementing the method, and wherein the instructions, when executed by the processor, perform the steps of the method. The method includes the steps of: receiving an acoustic mixture of signals generated by a tool performing a task and signals generated by the plurality of actuators actuating the tool; submitting the acoustic mixture of signals to a neural network that is trained to separate the signals generated by the tool performing the task from the signals generated by the actuators of the actuating tool to extract the signals generated by the tool performing the task from the acoustic mixture of signals; analyzing the extracted signals to generate a state of execution of the task; and performing a control action selected according to the state of execution of the task.
[0027] Another embodiment discloses a non-transitory computer-readable storage medium embodying a program that is executable by a processor for performing a method. The method includes the steps of: receiving an acoustic mixture of signals generated by a tool performing a task and signals generated by the plurality of actuators actuating the tool; submitting the acoustic mixture of signals to a neural network that is trained to separate the signals generated by the tool performing the task from the signals generated by the actuators of the actuating tool to extract the signals generated by the tool performing the task from the acoustic mixture of signals; analyzing the extracted signals to generate a state of execution of the task; and performing a control action selected according to the state of execution of the task. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1A Figure 1A FIG. shows a block diagram of a system for separating signals corresponding to different actuators that emit vibrations during machine operation and using the separated signals for subsequent monitoring tasks according to some embodiments.
[0029] Figure 1B Figure 1B FIG. shows a block diagram of a system for separating signals corresponding to different actuators that emit vibrations during machine operation and using the separated signals for subsequent monitoring tasks according to some embodiments.
[0030] Figure 1C Figure 1C FIG. shows a spectrogram of an acoustic mixture of signals generated by a tool performing a task and signals generated by a plurality of actuators actuating the tool according to some embodiments.
[0031] Figure 2 Figure 2 FIG. shows a flowchart illustrating the training of an acoustic signal processing system for source separation according to some embodiments.
[0032] Figure 3A Figure 3A A block diagram illustrating a single-channel mask inference source separation network architecture according to some embodiments.
[0033] Figure 3B Figure 3B A block diagram illustrating a convolutional recurrent network architecture for sound event classification according to some embodiments.
[0034] Figure 4 Figure 4 A flowchart illustrating some method steps for training a source separation network with weak labels according to some embodiments.
[0035] Figure 5 Figure 5 A schematic diagram showing the classification loss function and target required for forced separation of signals according to some embodiments.
[0036] Figure 6 Figure 6 A graph illustrating the execution state based on the analysis of isolated signals according to some embodiments.
[0037] Figure 7 Figure 7 A schematic diagram showing the control of a machining operation according to some embodiments.
[0038] Figure 8 Figure 8 A schematic diagram illustrating an actuator of a manufacturing anomaly detection system according to some embodiments. DETAILED DESCRIPTION
[0039] Figure 1A and Figure 1B A block diagram showing a system 100 for analyzing, performing, and controlling the operation of a machine 102 according to some embodiments. The machine 102 may include one or more actuators (components) 103, each of which performs a unique task and is connected to a coordination device 104. Examples of tasks performed by the actuators 103 may be machining, welding, or assembly. In some embodiments, the machine actuators 103 may operate simultaneously, but the coordination device 104 may need to control each actuator individually. An example of the coordination device 104 is a tool that performs a task. For example, in some embodiments, the system 100 controls the operation of a machine that includes multiple actuators that assist one or more tools in performing one or more tasks.
[0040] In some embodiments, sensor 101 can be a microphone or an array of multiple microphones, and sensor 101 captures vibrations generated by respective actuators 103 during the operation of machine 102. Additionally, some machine actuators 103 may be co-located in the same spatial region such that even if sensor 101 is a multi-microphone array, they cannot be captured individually by sensor 101. Thus, the vibration signal captured by sensor 101 is an acoustic mixture signal 195, which consists of the sum of vibration signals generated by respective machine actuators 103. In some embodiments, at least some sound sources in the spectrogram of the acoustic mixture signal occupy the same time, space, and frequency in the acoustic mixture.
[0041] In some embodiments, since quieter actuators may be masked by louder actuators, it may not be possible to estimate the execution states of respective machine actuators 103 from the acoustic mixture signal 195. To this end, system 100 includes a separation neural network 131, which can isolate the vibrations generated by respective machine actuators 103 from the acoustic mixture signal 195. Once the signals of respective machine actuators 103 are isolated from the acoustic mixture signal 195, they can be used for execution state estimation 135 and task execution 137.
[0042] To this end, system 100 includes various modules executed by processor 120 to control the operation of the machine. The processor submits the acoustic mixture of the signals into neural network 131 to extract the signals generated by the tools performing the tasks from the acoustic mixture of the signals, analyzes the extracted signals using estimator 135 to generate the task execution state, and executes control actions selected by controller 137 based on the execution state of the task and communicated through control interface 170 for applications such as avoiding failures and maintaining smooth operation in machine 102.
[0043] System 100 can have many input 108 and output 116 interfaces that connect system 100 to other systems and devices. For example, network interface controller 150 is adapted to connect system 100 to network 190 via bus 106. Through network 190 (wireless or wired), system 100 can receive acoustic mixture input signal 195. In some implementations, the human-machine interface 110 within system 100 connects the system to a keyboard 111 and a pointing device 112, where the pointing device 112 can include a mouse, trackball, touchpad, joystick, pointing stick, stylus, or touch screen, etc. Through interface 110 or NIC 150, system 100 can receive data such as the acoustic mixture signal 195 generated during the operation of machine 102.
[0044] System 100 includes an output interface configured to output separate acoustic signals corresponding to vibrations of respective actuators 103 generated during operation of machine 102, or an output of an execution state estimation system 135 that operates on the separate acoustic signals. For example, the output interface may include a memory to render the separate acoustic signals or state estimation results. For example, system 100 may be linked via bus 106 to a display interface 180 adapted to connect system 100 to a display device 185 (such as, for example, a speaker, headphones, computer monitor, camera, television, projector, or mobile device, etc.). System 100 may also be connected to an application interface 160 adapted to connect the system to a device 165 for performing various operations.
[0045] System 100 includes a processor 120 configured to execute stored instructions and a memory 140 storing instructions executable by the processor. Processor 120 may be a single-core processor, multi-core processor, computing cluster, or any number of other configurations. Memory 140 may include random access memory (RAM), read-only memory (ROM), flash memory, or any other suitable memory system. Processor 120 is connected via bus 106 to one or more input devices and output devices. These instructions implement a method of separating vibration signals generated during operation of machine 102 for performing estimation and future control.
[0046] Some embodiments are based on the recognition that instead of considering the state of the system performing a task, the execution state of the task itself may be considered. For example, when a system includes multiple actuators for performing one or more tasks, the state of the system includes the states of all of these actuators. However, in some cases, the state of the actuators is not the primary concern for control. In fact, the state of the actuators is needed to guide execution (e.g., execution of a task), and thus the execution state is the primary goal, while the state of the actuators performing the task is only a secondary goal.
[0047] However, in some cases, the state of the system is difficult to measure, and the state of execution of the task is difficult to define. For example, consider a computer numerical control (CNC) that machines a workpiece using a cutting tool. The state of the system includes the state of the actuator that moves the cutting tool along a tool path. The execution state of the machining is the state of the actual cutting. An industrial CNC system may have many different (sometimes redundant) actuators, where many state variables are in complex non-linear relationships with each other, making it difficult to observe the state of the CNC system. However, it may also be difficult to measure the workpiece machining state.
[0048] Some embodiments are based on the recognition that the state of task execution can be represented by an acoustic signal generated by such execution. For example, the execution state of CNC machining of a workpiece can be represented by a vibration signal caused by deformation of the workpiece during machining. Thus, if such a vibration signal can be measured, various classification techniques including machine learning methods can be used to analyze such a vibration signal to estimate the task execution state and select appropriate control actions for control execution.
[0049] However, the problem faced by this method is that such a vibration signal does not exist in isolation. For example, in a system including multiple actuators that actuate one or more tools to perform one or more tasks, the signal generated by the tool performing the task is always mixed with the signal generated by the actuator that actuates the tool. For example, the vibration signal generated by deformation of the workpiece is always mixed with the signal from the motor that moves the cutting tool. If such a vibration signal could be generated syntactically in some way in isolation or captured in some way in isolation, a neural network could be trained to extract such a signal. However, in many cases including CNC machining, it is impractical to generate such an isolated signal representing task execution. Similarly, it would also be impractical to separately record different signals with multiple microphones.
[0050] Figure 1C A spectrogram showing the acoustic mixing of the signal 196 generated by the tool performing task 102 and multiple actuators actuating tool 103 according to some embodiments is shown. In the case where the sensor 101 is a single microphone, the signals generated by all machine actuators 103 can overlap in both time and frequency 196. In this case, various acoustic signal processing techniques such as frequency or time selective filtering are ineffective for isolating the individual acoustic signals generated by the machine actuators. In the case where the sensor 101 is a microphone array, the sources generating sound signals at different spatial positions can be isolated by considering the delay differences between different microphones in the array. For example, in Figure 1C where the machine 102 includes actuators 103 indicated by C 1 ,..C 5 the microphone array can isolate the signal generated by actuator C 3 since it is located at a spatially unique position. However, since the pairs of signals (C 1 ,C 2 ) and (C 4 ,C 5 ) overlap spatially based on their physical positions in the machine 102 and overlap in both time and frequency based on the acoustic signal spectrogram 196 that cannot be separated by conventional techniques.
[0051] To this end, in order to streamline the manufacturing automation of a system including multiple actuators that actuates one or more tools to perform one or more tasks, it is necessary to separate the source signals from the acoustic mixture of the signals generated by the tools performing the tasks and the signals generated by the multiple actuators that actuate the tools. Therefore, an object of some embodiments is to train a neural network for sound separation of the sound source of the mixed signal in the absence of an isolated sound source.
[0052] In some cases, at least some sound sources occupy the same time and / or frequency spectrum in the spectrogram of the acoustic mixture. For example, in one embodiment, in a room where the use of microphone array technology is impracticable, the sound sources in the mixture occupy the same area. Additionally or alternatively, the acoustic mixture of some embodiments is a single channel from the output of only a single microphone.
[0053] Some embodiments are based on the recognition that when the sound sources in the acoustic mixture can be easily isolated and recorded or recognized by a person, for example, supervised learning can be utilized to provide such training. This type of classification is referred to herein as a strong label. In theory, a person can produce an approximation of this type of label. However, it is largely unrealistic to require a person to precisely label the sound activity in both time and frequency to provide a strong label. However, it is realistic to consider some limited labels (weak labels) regarding which sound is active within a certain time range. In general, such weak labels do not require the sound to be active throughout the entire specified range and can only appear at brief instants within that range.
[0054] Some embodiments are based on the recognition that since the structure of the manufacturing process is usually known, the sources that generate the acoustic signals are usually also known. To this end, weak labels can be provided for the sources of the acoustic mixture representing the operation of the manufacturing system. To this end, some embodiments have developed methods that can learn to separate the sounds in an acoustic mixture for which training data with only weak labels is available.
[0055] Therefore, some embodiments train a neural network 131 to separate the signal generated by the tool performing the task from the signal generated by the actuator that actuates the tool from the acoustic mixture. For example, the neural network is trained to separate different signals from the acoustic mixture such that each separated signal belongs to a class of signals present in the system operation, and the separated signals add up to the acoustic mixture. The weak labels identify the classes of signals present in the system operation. Each weak label specifies a class of signals present at a certain point during the operation.
[0056] Figure 2 is a flowchart showing the training of an acoustic signal processing system 200 for separating acoustic mixture signals according to some embodiments of the present disclosure. Figure 2Shows a general source separation scenario where the system estimates multiple target acoustic signals from a mixture of a target acoustic signal and potentially other non-target sources (e.g., noise). In an example where the target acoustic signal is generated by vibrations from a machine and these vibrations cannot exist in isolation, the training objective of the source separation system is identified by weak labels, i.e., the training only requires whether a source exists in a specific time block, rather than an isolated source. The acoustic mixture input signal 204 includes the sum of multiple overlapping sources and is sampled from a training set that includes the acoustic mixture signal and the corresponding weak labels 222 recorded during machine operation 202.
[0057] The mixture input signal 204 is processed by a spectrogram estimator 206 to compute the time-frequency representation of the acoustic mixture. Then, the spectrogram is input to a mask inference network 230 using stored network parameters 215. The mask inference network 230 makes decisions regarding the presence of each source class in each time-frequency bin of the spectrogram and estimates a set of amplitude masks 232. There is one amplitude mask for each source, and a set of enhanced spectrograms 234 is computed by multiplying each mask with the complex time-frequency representation of the acoustic mixture. A set of estimated acoustic signal waveforms 216 is obtained by passing each enhanced spectrogram 234 through a signal reconstruction process 236 that reverses the time-frequency representation computed by the spectrogram estimator 206.
[0058] The enhanced spectrograms 234 can pass through a classifier network 214 using stored network parameters 215. The classifier network provides a probability regarding the presence of a given source class for each time frame of each enhanced spectrogram. The classifier operates once for each time frame of the spectrogram, however the weak labels 222 can have a much lower time resolution than the spectrogram, so the output of the classifier network 214 passes through a temporal pooling 217 module which, in some embodiments, can take the maximum of all frame-level decisions corresponding to one weak label time frame, average the frame-level decisions, or use some other pooling operation to combine the frame-level decisions. The network training module 220 can update the network parameters 215 using an objective function.
[0059] Figure 3AIt is a block diagram showing a single-channel mask inference network architecture 300A according to an embodiment of the present disclosure. A sequence of feature vectors obtained from an input mixture (e.g., the log magnitude of the short-time Fourier transform of the input mixture) is used as the input to the mixture encoder 310. For example, the dimension of the input vectors in the sequence can be F. The mixture encoder 310 consists of a plurality of bidirectional long short-term memory (BLSTM) neural network layers, from the first BLSTM layer 330 to the last BLSTM layer 335. Each BLSTM layer consists of a forward long short-term memory (LSTM) layer and a backward LSTM layer, and the outputs of which are combined by the next layer and used as the input. For example, the dimension of the output of each LSTM in the first BLSTM layer 330 can be N, and both the input dimension and the output dimension of each LSTM in all other BLSTM layers including the last BLSTM layer 335 can be N. The output of the last BLSTM layer 335 is used as the input to the mask inference module 312 (including the linear neural network layer 340 and the non-linearity 345). For each time frame and each frequency in the time-frequency domain (e.g., the short-time Fourier transform domain), the linear layer 340 uses the output of the last BLSTM layer 335 to output C numbers, where C is the number of target sources to be separated. The non-linearity 345 is applied to this set of C numbers for each time frame and each frequency, thereby obtaining mask values indicating that the target source dominates in the input mixture at that time frame and that frequency for each time frame, each frequency, and each target source. The separated coding estimate from the mask module 313 uses these masks together with the representation of the input mixture in the time-frequency domain (e.g., the magnitude short-time Fourier transform domain) of the estimated mask to output separated coding for each target source. For example, the separated coding estimate from the mask module 313 can multiply the mask of the target source by the magnitude short-time Fourier transform of the input mixture to obtain an estimate of the magnitude short-time Fourier transform of the separated signal of the target source (which is used as the separated coding of the target source if observed in isolation).
[0060] Figure 3BIt is a block diagram showing a single-channel convolutional recurrent network classification architecture 300B according to an embodiment of the present disclosure. A sequence of feature vectors obtained from an input mixture (e.g., the log magnitude of the short-time Fourier transform of the input mixture) is used as the input to the mixture encoder 320. For example, the dimension of the input vectors in the sequence can be F. The mixture encoder 320 consists of a plurality of convolutional blocks from the first convolutional block 301 to the last convolutional block 302, followed by a recurrent BLSTM layer 303. Each convolutional block consists of a convolutional layer with learned weights and biases, followed by a pooling operation in both the time dimension and the frequency dimension, which reduces the input dimension to the subsequent layer. The BLSTM layer 303 at the end of the mixture encoder 320 consists of a forward long short-term memory (LSTM) layer and a backward LSTM layer, and their outputs are combined and used as the input to the classifier module 322. The classifier module 322 includes a linear neural network layer 305 and a module that implements a sigmoid non-linearity 307. The classifier module 322 outputs C numbers for each time frame of the input signal, which represent the probability that a given type of source is active in the current time frame.
[0061] Figure 4 It is a flowchart of a method for collecting data and training a neural network to separate acoustic mixed signals that do not need to exist in isolation or are spatially separated from each other. A sequence of acoustic recordings containing the signals to be separated is collected from a machine or a sound environment to form a set of training data records 400. These records are annotated to generate weak labels for all the training data signals 410. This annotation process can include annotations taken for the periods of source activity and inactivity during the time of collecting the training data records 400. Additionally or alternatively, the weak labels for all the training data signals 410 can be collected forensically by listening to the training data records 400 and annotating the periods of source activity and inactivity to be separated. Additionally or alternatively, the weak labels are determined based on the operating specifications of a controlled machine.
[0062] The next step is to train a classifier to predict the weak label 420 of each training data signal 400 using the weak labels of all the training data signals 410 as the target. This classifier can also be referred to as a sound event detection system. Subsequently, the separator network is trained using the classifier as supervision 430, which is described in more detail below. The separator network then extracts isolated sound sources 440 as a set of signals, where each separated signal belongs to only a single class or signal type. These separated signals can then be used to further train the classifier 420 to only predict the class of the separated signals, while all other class outputs are zero. In addition, the training of the separator network 430 can also use the previously separated signals as input, and the weights of the network can be updated such that the correctly separated signals pass through the separator network unchanged.
[0063] Figure 5A schematic diagram illustrating the relationship between weak labels and classifier outputs that facilitate signal separation according to some embodiments is shown. In these embodiments, the neural network is jointly trained with the classifier to minimize a loss function that includes a cross-entropy term between the weak labels and a separation neural network that runs through the classifier. Figure 5 The illustration is for a single time frame and is repeated at the time resolution at which the weak labels are available.
[0064] The classifier output 510 is obtained by running each of the C signals extracted by the separation network through the classifier. That is, each row of the classifier output matrix 510 corresponds to the classifier output for one separated signal. The weak label target matrix 530 arranges the provided weak labels such that they can be used to train the separation system. For each of the C signal classes, if the signal of class i does not exist, then we set t i = 0, and if the signal of class i exists, then t i = 1. For classifier training, the cross-entropy loss function 520 only needs to have the diagonal elements of the weak label target matrix 530 with the classifier output from the acoustic mixture, and can be mathematically represented as
[0065]
[0066] where p i (i = 1, …, C) is the classifier output when operating on the acoustic mixture signal for class i.
[0067] However, since we need to separate the signals, the embodiments enforce that the off-diagonal terms in the classifier output matrix 510 are equal to zero in the weak label target matrix 530. This helps to enforce that each separated signal belongs to only a single source class. Thus, the cross-entropy loss function 520 for separation can be mathematically represented as:
[0068]
[0069] where H(t i , p ii ) is the cross-entropy loss defined above. One problem with using the classification loss to train the separation system is that the classifier often can make its decision based on only a small subset of the available frequency spectrum, and if the separator only learns to separate a part of the spectrum, the goal of extracting the isolated signals will not be achieved. To avoid this, another loss term is added that enforces that the extracted signals of all active sources sum to the acoustic mixture signal, and penalizes any energy belonging to inactive sources. This is mathematically represented as:
[0070]
[0071] where f is the frequency index, X(f) is the acoustic mixture amplitude, is the amplitude of the separated source. Finally, the two actuators are combined and the overall loss function is performed as follows:
[0072] L overall = L Class + αL Mag
[0073] where α is a term that allows weighting the relative importance of the individual loss actuators.
[0074] In this way, the present disclosure proposes a system and method for manufacturing automation using a sound separation neural network. The machine learning algorithm proposed in the present disclosure uses a neural network to isolate individual signals caused by vibrations of different parts in a manufacturing scenario. In addition, since the vibration signals that make up the sound mixture may not exist in isolation or be spatially separated from each other, the proposed system can be trained using weak labels. In this case, a weak label refers to only having access to the time periods of different source activities in the sound mixture, rather than the isolated signals that make up the mixture. Once the signals that make up the sound mixture are separated, they can be used to monitor the execution status of individual machine parts or as control signals for independently controlling different actuators.
[0075] Figure 6 A graph is shown that illustrates the execution status based on the analysis of signals isolated according to some embodiments. In the case where the neural network 131 is trained for signal separation, some embodiments use the output of this network to train a state estimator 135 for estimating the task execution status. For example, in one embodiment, the state estimator is trained on signals 601 and 602 generated by a tool performing a task and extracted from different sound mixtures of different repetitions of the system operation by the neural network.
[0076] For example, in some embodiments, the state estimator 135 is configured to detect a prediction pattern 615 indicating the execution status. For example, the real-valued time series of the isolated signals collected during period 617 may include a normal region 618 and an abnormal region T 619 that leads to a fault point 621. The state estimator 135 may be configured to detect the abnormal region 619 to prevent the fault 621. For example, in some implementations, the state estimator 135 uses a Shapelet discovery method to search for a prediction pattern until the best prediction pattern is found. At least one advantage of using the Shapelet discovery algorithm is an efficient search for prediction patterns of different lengths. Internally, Shapelet discovery optimizes the prediction pattern according to a predetermined measurement criterion. For example, the prediction pattern should be as similar as possible to one pattern in the abnormal region and as different as possible from all patterns in the normal region.
[0077] Additionally or alternatively, in different embodiments, the state estimator 135 is implemented as a neural network that is trained to estimate the task execution state from the output of the extraction neural network 131. Advantageously, the state estimator of this embodiment can be co-trained with both the neural network 131 and the classifier used to train the neural network 131. In this way, this embodiment provides an end-to-end training solution for controlling complex machines.
[0078] Notably, in some embodiments, individual samples of the extraction signal generated by the tool performing the task define the task execution state, but are not sufficient to define the system operation state. However, the task execution state is sufficient to select an appropriate control action. In this way, the embodiment provides dimensionality reduction in manufacturing automation applications.
[0079] For example, in some embodiments, the system is configured to execute an operation sequence for machining a workpiece. The tool in these embodiments is a machining tool, the processor executes computer numerical control (CNC) for actuating the machining tool along a tool path, and the signal generated by the machining tool is a vibration signal generated by the deformation of the workpiece during machining of the machining tool. Many machining systems have multiple actuators for positioning the tool. Additionally, many machining systems may have redundant actuators for positioning the tool along various degrees of freedom. Additionally, the type of tool can also affect the performance of the system. However, all of these variables can be captured by the embodiment through isolation and classification of the signal indicating task execution.
[0080] Figure 7 A schematic diagram showing the control of a machining operation according to some embodiments is shown. A set of machining instructions 701 is provided to an NC machining controller 702 (e.g., as a file via a network). The controller 702 includes a processor 703, a memory 704, and a display 705 for displaying the operation of the machine. According to some embodiments, the processor runs the extraction neural network 131, the state estimation 135 of the state machining, and the controller operation 137. In some implementations, the neural network 131, the state estimator 135, and the controller 137 are adapted to perform different machining tools 702, 704, 706, and 708 for different types of machining 712, 714, 716, and 718 of the workpiece 710. For example, the controlled machine can execute an operation sequence for manufacturing a workpiece including one or a combination of machining, welding, and assembling the workpiece, such that the signal generated by the tool is a vibration signal generated by the modification of the workpiece during its manufacture.
[0081] In fact, state estimation based on the extraction signal enables control to be adapted to different types of complex manufacturing. However, some embodiments are not limited to factory automation. For example, in one embodiment, the controlled machine is a gearbox to be monitored for potential anomalies, and the gearbox can be recorded only in the presence of vibrations from the motor, the coupling, or other vibrations from the moving parts.
[0082] For dimensionality reduction, additionally or alternatively, some embodiments allow independent control of different tools performing different operation tasks. For example, when analyzing a signal representing the state of an entire system, such analysis can provide control over the entire system. However, when performing analysis separately for different tools performing different tasks, independent control of the tools is possible.
[0083] To this end, when multiple tools perform multiple tasks during system operation, some embodiments perform independent control of the tasks. For example, in one embodiment, a system with a first tool performing a first task and a second tool performing a second task is controlled. A neural network is trained to separate a first signal generated by the first tool performing the first task and a second signal generated by the second tool performing the second task from an acoustic mixture. During system operation, the first signal and the second signal are extracted from the acoustic mixture using the neural network and analyzed independently of each other to estimate a first state of the execution of the first task and a second state of the execution of the second task. The embodiment is configured to perform a first control action selected according to the first state and a second control action selected according to the second state.
[0084] Figure 8 FIG. shows a schematic diagram of an actuator of a manufacturing anomaly detection system 800 according to some embodiments. The system 800 includes a manufacturing production line 810, a training data pool 820, a machine learning model 830, and an anomaly detection model 840. The production line 810 uses sensors to collect data. The sensors can be digital sensors, analog sensors, and combinations thereof. The data collected is used for two purposes. Some of the data is stored in the training data pool 820 and used as training data to train the machine learning model 830, and some of the data is used by the anomaly detection model 840 as operational time data to detect anomalies. Both the machine learning model 830 and the anomaly detection model 840 can use the same data.
[0085] To detect anomalies in the manufacturing production line 810, training data is first collected. The machine learning model 830 uses the training data in the training data pool 820 to train an extraction neural network 131. The training data pool 820 can include labeled data or unlabeled data. Labeled data has been labeled with labels (e.g., anomaly or normal). Unlabeled data has no labels. Based on the type of training data, the machine learning model 830 applies different training methods. For labeled training data, supervised learning is typically used, and for unlabeled training data, unsupervised learning is typically applied. In this way, different embodiments can handle different types of data.
[0086] The machine learning model 830 learns the features and patterns of the training data, including normal data patterns and abnormal data patterns. The anomaly detection model 840 uses the trained machine learning model 850 and the collected operation time data 860 to perform anomaly detection. The operation time data 860 can be identified as normal or abnormal. For example, using the normal data patterns 855 and 858, the trained machine learning model 850 can classify the operation time data into normal data 870 and abnormal data 880. For example, the operation time data X1 863 and X2 866 are classified as normal, and the operation time data X3 869 is classified as abnormal. Once an anomaly is detected, necessary actions 890 are taken.
[0087] In some embodiments, neural networks are trained for extraction for each monitored process X1 863, X2 866, and X3 869. The controller can take actions 890 to control one process (e.g., process X1 863) independently of other processes (e.g., processes X2 866 and X3 869). This separation of process control based on signal extraction simplifies the control of complex manufacturing processes and makes such control more accurate and practical.
[0088] The above-described embodiments of the present invention can be implemented in any of numerous ways. For example, the embodiments can be implemented using hardware, software, or a combination thereof. When implemented in software, the software code can be executed on any suitable processor or set of processors, whether disposed in a single computer or distributed among multiple computers. These processors can be implemented as integrated circuits having one or more processors in the integrated circuit actuator. However, the processors can be implemented using any suitable format of circuitry.
[0089] Additionally, embodiments of the present invention can be specifically implemented as a method, and examples thereof are provided. The actions performed as part of the method can be sequenced in any suitable manner. Accordingly, embodiments can be constructed that execute the actions in an order different from that shown, which can include performing some actions simultaneously, although shown as sequential actions in the illustrative embodiments.
[0090] The use of ordinal terms such as "first" and "second" in the claims to modify the claim elements themselves does not imply any precedence, anteriority, or order of one claim element over another or the temporal order of method act executions, but is only used as a label to distinguish one claim element having a particular name from another element having the same name (but using an ordinal term) to distinguish the claim elements.
[0091] Although the present invention has been described by way of example of the preferred embodiments, it will be understood that various other adaptations and modifications can be made within the spirit and scope of the present invention.
[0092] Accordingly, the aim of the appended claims is to cover all such variations and modifications as fall within the true spirit and scope of the present invention.
Claims
1. A system for controlling the operation of a machine, the machine including a plurality of actuators assisting one or more tools to perform one or more tasks, the system comprises: An input interface configured to receive, during operation of the system, an acoustic mixture of signals generated by a tool performing a task and signals generated by the plurality of actuators actuating the tool; A memory configured to store a neural network trained to separate, from the acoustic mixture, signals generated by the tool performing the task from signals generated by the actuators actuating the tool; And A processor configured to submit the acoustic mixture of signals to the neural network to extract, from the acoustic mixture of signals, signals generated by the tool performing the task; analyze the extracted signals to generate a state of execution of the task; And execute a control action selected according to the state of execution of the task, wherein, during operation of the machine, a plurality of tools perform a plurality of tasks, the plurality of tools including a first tool performing a first task and a second tool performing a second task, wherein the neural network is trained to separate a first signal generated by the first tool performing the first task and a second signal generated by the second tool performing the second task from the acoustic mixture, and wherein, during operation of the system, the processor is configured to use the neural network to extract the first signal and the second signal from the acoustic mixture, analyze the first signal independently of the second signal to estimate a first state of execution of the first task and a second state of execution of the second task, and execute a first control action selected according to the first state and execute a second control action selected according to the second state.
2. The system according to claim 1, wherein, The machine is configured to perform an operation sequence for machining a workpiece, wherein the tool is a machining tool, the processor executes computer numerical control (CNC) for actuating the machining tool along a tool path, and wherein the signal generated by the machining tool is a vibration signal generated by deformation of the workpiece during machining of the workpiece by the machining tool.
3. The system according to claim 1, wherein, The machine is configured to perform an operation sequence for manufacturing a workpiece, the operation sequence including one or a combination of machining, welding, and assembling the workpiece, such that the signal generated by the tool is a vibration signal generated by modification of the workpiece during manufacture of the workpiece.
4. The system according to claim 1, wherein, The machine is a gearbox to monitor potential anomalies, and the gearbox can be recorded only in the presence of vibrations from a motor, a coupling, or other vibrations from moving parts.
5. The system according to claim 1, wherein, The neural network is trained to separate different signals from the acoustic mixture such that each separated signal belongs only to one class of signals present in the operation of the system, and the separated signals together total the acoustic mixture.
6. The system according to claim 5, wherein, The classes of signals present during operation of the system are identified by weak labels, each weak label specifying the class of signal present at a point during the operation.
7. The system according to claim 6, wherein, the neural network is jointly trained with a classifier configured to classify the classes of signals identified by the weak labels.
8. The system according to claim 7, wherein, the neural network is jointly trained with the classifier to minimize a loss function including a cross-entropy term between the weak labels and a separate neural network passing through the classifier.
9. The system according to claim 6, wherein, the neural network is trained such that when separate signals of the classes identified by the weak labels are submitted as input to the neural network, the neural network generates the separate signals as output.
10. The system according to claim 1, wherein, the processor executes a state estimator to estimate the state of execution of the task, wherein the state estimator is trained on signals generated by the tool performing the task and separately extracted by the neural network from different sound mixtures of different repetitions of the operation of the system.
11. The system according to claim 1, wherein, each sample of the extracted signals generated by the tool performing the task defines the state of execution of the task, but is insufficient to define the state of operation of the machine.
12. The system according to claim 1, wherein, the signals generated by the tool performing the task are mixed with the signals generated by the actuator actuating the tool to occupy the same time and frequency spectrum in the sound mixture.
13. The system according to claim 1, wherein, the sound mixture is from a single channel of the output of a single microphone.
14. The system according to claim 1, wherein, at least some of the actuators spatially overlap based on their physical positions in the machine.
15. A method of controlling the operation of a machine including a plurality of actuators assisting one or more tools in performing one or more tasks, wherein, the method uses a processor coupled with stored instructions implementing the method, wherein the instructions, when executed by the processor, perform the steps of the method, the method including the steps of: receiving a sound mixture of signals generated by a tool performing a task and signals generated by the plurality of actuators actuating the tool; submitting the sound mixture of signals into a neural network trained to separate the signals generated by the tool performing the task from the signals generated by the actuators actuating the tool from the sound mixture to extract the signals generated by the tool performing the task from the sound mixture of signals; analyzing the extracted signals to generate the state of execution of the task; and performing a control action selected according to the state of execution of the task, During operation of the machine, a plurality of tools perform a plurality of tasks, the plurality of tools including a first tool that performs a first task and a second tool that performs a second task, wherein the neural network is trained to separate from the acoustic mixture a first signal generated by the first tool performing the first task and a second signal generated by the second tool performing the second task, and wherein during operation of the machine, the processor is configured to use the neural network to extract the first signal and the second signal from the acoustic mixture, analyze the first signal independently of the second signal to estimate a first state of execution of the first task and a second state of execution of the second task, and perform a first control action selected based on the first state and a second control action selected based on the second state.
16. The method according to claim 15, wherein, signals generated by the tool performing the task are mixed with signals generated by the actuator actuating the tool to occupy the same time and frequency spectrum in the acoustic mixture.
17. The method according to claim 15, wherein, the acoustic mixture is a single channel from the output of a single microphone.
18. The method according to claim 15, wherein, at least some of the actuators overlap spatially based on their physical positions in the machine.
19. A non-transitory computer-readable storage medium having embodied thereon a program that can be executed by a processor to perform a method that comprises the steps of: receiving an acoustic mixture of signals generated by a tool performing a task and signals generated by a plurality of actuators actuating the tool; submitting the acoustic mixture of signals into a neural network that is trained to separate from the acoustic mixture signals generated by the tool performing the task and signals generated by the actuator actuating the tool to extract from the acoustic mixture of signals signals generated by the tool performing the task; analyzing the extracted signals to generate a state of execution of the task; and performing a control action selected based on the state of execution of the task, wherein a plurality of tools perform a plurality of tasks, the plurality of tools including a first tool that performs a first task and a second tool that performs a second task, wherein the neural network is trained to separate from the acoustic mixture a first signal generated by the first tool performing the first task and a second signal generated by the second tool performing the second task, and wherein during operation of the method, the processor is configured to use the neural network to extract the first signal and the second signal from the acoustic mixture, analyze the first signal independently of the second signal to estimate a first state of execution of the first task and a second state of execution of the second task, and perform a first control action selected based on the first state and a second control action selected based on the second state.
Citation Information
Patent Citations
Neural network model training method and device for weak annotation data
CN110070183A
Device and method of detecting abnormality of cutting tool
JP2014213412A