Acoustic Detection of Glass Breakage Events
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- AMGEN INC
- Filing Date
- 2023-04-25
- Publication Date
- 2026-04-24
AI Technical Summary
Conventional manufacturing processes, particularly biomanufacturing, face challenges in detecting glass breakage events due to high production speeds and the difficulty of human monitoring, leading to undetected events and associated risks and costs.
The use of machine learning techniques to identify glass breakage events by training a model with audio data representing ambient, glass, and glass breakage sounds, allowing for the processing of audio data to detect breakage events and automatically notify operators.
This approach enables timely detection of glass breakage events, reducing labor and scrap costs, ensuring product quality, and improving safety by quickly stopping operations upon detection.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications The specification of U.S. Patent Application No. 63 / 409,498 filed on September 23, 2022, and the specification of U.S. Patent Application No. 63 / 336,847 filed on April 29, 2022 are hereby incorporated by reference in their entirety.
[0002] This application generally relates to the use of a predictive model to identify glass breakage events. More particularly, this application may relate to the identification of glass breakage events in a biomanufacturing process using audio data.
Background Art
[0003] Many manufacturing processes process glass or otherwise utilize glass. For example, glass may be included in components of machines or equipment used in the manufacturing process, may be used to contain or package products produced by the manufacturing process, or may be included in the product itself, among other things. A biomanufacturing process is one type of manufacturing process that routinely requires glass. In a biomanufacturing process, a filling and finishing process relates to preparing a biological formulation (e.g., a pharmaceutical) in its delivery container. In many biological formulations, the delivery container is a glass vial. Other biomanufacturing processes / biological formulations that include glass are pre-filled glass syringes, auto-injectors, or intravenous (IV) bags. Glass is an exemplary material for use in biomanufacturing processes due to its chemical durability, gas tightness, strength, cleanliness, and transparency. However, glass containers are not without risk. Problems with glass containers include breakage, delamination, and the inclusion of glass particles or shards, all of which can affect the safety and efficacy of biological formulations. In fact, due to product recalls over the past decade due to glass problems including vial flaking, breakage, and particulate contamination, hundreds of millions of units of biological formulations packaged in vials, syringes, or glass auto-injectors have been removed from the market as a result. Conventionally, glass breakage events detected during manufacturing stop operations, record, and then trigger an immediate predefined process to investigate the root cause of the glass breakage event, which can often take several days. Thus, glass breakage events not only result in lost labor hours during downtime, but also in the form of scrap costs because all pharmaceuticals produced during operation usually have to be discarded, which can cost hundreds of thousands of dollars. Ultimately, glass breakage poses a human safety issue for operators because glass shards pose a cut hazard during cleaning.
[0004] While the risks and costs of glass breakage events in manufacturing processes (e.g., biomanufacturing processes) remain significant, the progress of manufacturing processes due to automation, robotics, and disposable technologies, as well as continuous growth, has resulted in production speeds that are too high to be feasibly monitored or inspected by humans alone from start to finish. Without human monitoring, certain events such as glass breakage events may not be detected. Many of the adverse effects considered earlier for glass breakage events only worsen if the glass breakage event goes undetected for a long period. Moreover, even in manufacturing processes with production speeds low enough such that humans may monitor the process, glass breakage events are difficult to detect by human vision and may sometimes simply result in small particles or micro-damage.
[0005] Accordingly, in conventional manufacturing processes (e.g., biomanufacturing processes), there is an increased likelihood that glass breakage events will go undetected over a long period, thus resulting in significant risks and costs.
SUMMARY OF THE INVENTION
MEANS FOR SOLVING THE PROBLEM
[0006] Aspects of the present disclosure are methods for using machine learning to identify glass breakage events, the method comprising: (a) accessing or obtaining a machine learning model trained using training audio data representing (i) training ambient sound, (ii) training glass sound, and (iii) training glass breakage sound; (b) obtaining audio data over a period of interest by one or more processors; (c) processing the audio data using the machine learning model to identify whether a glass breakage event occurred during the period of interest; and (d) indicating that a glass breakage event occurred by one or more processors when it is identified that a glass breakage event occurred.
[0007] In some aspects, the training audio data and the audio data each include one or more spectrograms, and obtaining the audio data over an interest period includes generating a spectrogram of the audio data from the raw audio data using a Fourier transform.
[0008] In some aspects, (i) the training ambient sound includes sounds caused by the operation of the machine, (ii) the training glass sound includes sounds caused by a first glass surface contacting either a second glass surface or a non-glass surface, and (iii) the training glass break sound includes sounds caused by either a glass crack or a glass break. In some aspects, the machine performs a biomanufacturing process, and the training glass sound and the training glass break sound are generated by one or more containers of one or more pharmaceuticals.
[0009] In some aspects, the method further includes automatically stopping the operation of the machine by one or more processors after identifying a glass break event.
[0010] In some aspects, the method further includes repeatedly adjusting, by one or more processors, the gain of amplification applied to either or both of the training audio signal or the audio signal until a performance threshold is met, wherein the training audio signal and the audio signal respectively correspond to the training audio data and the audio data.
[0011] In some aspects, the machine learning model is a convolutional neural network.
[0012] Another aspect of the present disclosure provides a computer system comprising: (a) one or more processors; and (b) a program memory coupled to the one or more processors and storing executable instructions that, when executed by the one or more processors, cause the computer system to perform any one of the methods of the previous aspects.
[0013] A further aspect of the present disclosure is a method and executable instructions for training a machine learning model for identifying glass breakage events, the method and executable instructions including: (a) obtaining training audio data; (b) classifying the training audio data into a plurality of subsets each corresponding to different actual result data, the subsets including: (i) at least one subset representing training ambient sound, (ii) at least one subset representing training glass sound, and (iii) at least one subset representing training glass breakage sound; and (c) generating a machine learning model for identifying glass breakage events using the classified subsets of the training audio data.
[0014] Those skilled in the art will understand that the figures included herein are for illustrative purposes and do not limit the present disclosure. The drawings are not necessarily to scale and instead focus on illustrating the principles of the present disclosure. In some examples, various aspects of the illustrated embodiments may be shown exaggerated or enlarged to facilitate understanding of the illustrated embodiments. In the drawings, like reference numerals generally refer to components that are functionally or structurally similar throughout the various figures.
Brief Description of the Drawings
[0015]
Figure 1A
Figure 1B
Figure 2
Figure 3
Figure 4A
Figure 4B
Figure 5A
Figure 5B
Figure 6
Figure 7
[0016] As the pace of the biomanufacturing process quickens, there is an increasing need to monitor and detect glass breakage events as quickly as possible. This disclosure aims to reduce problems with the prior art (e.g., as described in the background art section) by providing techniques that use machine learning techniques to identify glass breakage events. The techniques may train a machine learning model using training audio data representing ambient sound, glass sound, and glass breakage sound that may be labeled in the training audio data accordingly. The techniques may use the machine learning model to identify whether a glass breakage event has occurred over a period of interest represented by audio data, and if so, notify or indicate as such. By identifying when a glass breakage event has occurred and then notifying / indicating, the techniques provide insight to the operator of the biomanufacturing process system and aim to reduce adverse effects associated with glass breakage events (e.g., as described in the background art section).
[0017] This technique detects when a glass breakage event occurs and immediately stops the manufacturing operation to reduce operating costs and scrap, thereby ensuring the highest product quality for patients receiving biologics (e.g., pharmaceuticals) manufactured by the biomanufacturing process. Audio is selected as the most viable medium for detecting glass breakage events due to the following four main factors: namely, cost, data volume, visibility, and adaptability. A basic audio capture device is relatively inexpensive compared to other possible detection means (e.g., video, laser curtain systems, pressure sensors, accelerometers, etc.) while requiring little calibration other than gain setting. Video detection, on the other hand, may require expensive equipment because glass breakage events occur very quickly and sometimes within a narrow area. Since glass breakage events must always be monitored, it is important to minimize the computational power required both to manipulate and store the relevant data. Finally, since a large number of biomanufacturing processes can theoretically result in glass breakage, audio detection is well-suited in that it is adaptable to various biomanufacturing processes (unlike video detection) as it does not require a line of sight. This point is particularly important since many processes for vial, autoinjector, and syringe filling and finishing are performed without external visibility to the glass components.
[0018] Advantageously, by providing improved insight into glass breakage events, the present technique reduces the adverse effects associated with conventional approaches to dealing with glass breakage events. One advantage of these insights is that fewer resources (e.g., fewer biological agents) are wasted during glass breakage events that are detected more quickly using the present technique, and thus resource efficiency is increased and the sustainability of the biomanufacturing process system is improved. By making the biomanufacturing process system more sustainable with respect to resource use, it may also improve the energy efficiency of the biomanufacturing process system and reduce the monetary or economic costs of producing each biological agent. Another advantage of the improved insight is that more biological agents can be produced in a given time and the production throughput may increase because the reset time is shortened in the event of glass breakage. The techniques of the present disclosure provide insight into the utility of biological agents produced under glass breakage events (e.g., by providing insight into the likelihood that broken glass has contaminated the biological agent), and thus resource, energy, and cost efficiency may also be improved when dealing with glass breakage events in the biomanufacturing process. Further, by making glass breakage events easier to identify, data on the biomanufacturing process before, during, and after a glass breakage event (e.g., audio data, operational data, etc.) can be analyzed to provide insight into the causes and effects of the glass breakage event, thereby potentially improving the technical field of glass breakage event diagnosis.
[0019] Additional advantages of the techniques of the present disclosure over conventional approaches to identifying glass breakage events will be appreciated by those skilled in the art through the present disclosure. The various concepts and techniques introduced above and discussed in more detail below may be implemented in any number of ways and the concepts described are not limited to any particular implementation. Examples of embodiments are provided for illustrative purposes.
[0020] Exemplary System Figure 1A is a block schematic diagram of an exemplary system 100A for identifying glass breakage events in a biomanufacturing process machine 160 that may produce, for example, pharmaceuticals. In some embodiments, system 100A includes a stand-alone device, while in other embodiments, system 100A is incorporated into other devices. At a high level, system 100A includes components of computing devices 110, one or more training audio data sources 150, a biomanufacturing process machine 160, and one or more audio sensors 162. In FIG. 1A, computing device 110, biomanufacturing process machine 160, and training audio data source 150 are communicatively coupled via a network 170 that may be a proprietary network, a secure public internet, a virtual private network, and / or any other suitable wired or wireless network (e.g., dedicated access lines, satellite links, cellular data networks, combinations thereof, etc.), or that may include them. In embodiments where network 170 comprises the internet, data communication may occur on network 170 via internet communication protocols. In some aspects, system 100A may include more or fewer instances of various components than shown in FIG. 1A (e.g., one instance of computing device 110, ten instances of biomanufacturing process machine 160, ten instances of audio sensor 162, two instances of training audio data source 150, etc.).
[0021] Although system 100A is shown as including biomanufacturing process machine 160, it is worth noting that those skilled in the art will understand that the present technology and components of system 100A may be applicable to detecting glass breakage events in other processes or fields. For example, instead of biomanufacturing process machine 160, the present techniques and components of system 100A may be applicable to manufacturing in food / beverage, automotive, electronics, chemical, and / or other industries.
[0022] Biomanufacturing process machine 160 may include a single biomanufacturing process machine, or multiple biomanufacturing process machines that are either located in the same place or separated from each other and are suitable for producing biological agents such as pharmaceuticals. Biomanufacturing process machine 160 may generally include physical devices configured for use in producing (e.g., manufacturing) biological agents (e.g., pharmaceuticals), such as filling devices, agitation devices, star wheels, or other container handling devices.
[0023] In some embodiments, biomanufacturing process machine 160 may be connected to computing device 110 either via network 170 or directly, enabling at least some of the functionality of biomanufacturing process machine 160 to be controlled by computing device 110. In some embodiments, biomanufacturing process machine 160 may be capable of receiving commands directly from a user (e.g., biomanufacturing process machine 160 may be manually configurable). For example, in some embodiments, biomanufacturing process machine 160 may receive commands directly from a user to control its operation (e.g., start or stop its operation).
[0024] The audio sensor 162 may be included in the biomanufacturing process machine 160 (e.g., integrated into the biomanufacturing process machine 160), or may be an external sensor connected to the biomanufacturing process machine 160. The audio sensor 162 may be used to collect (e.g., directly or indirectly) audio data inside, outside, or around the biomanufacturing process machine 160. The audio sensor 162 may provide audio data to, for example, the computing device 110 (e.g., via the network 170). The audio data may be any suitable data type, such as nominal data, ordinal data, discrete data, or continuous data. The audio data may be in the form of a suitable data structure, which may be stored in a suitable format such as one or more of the following: M4A, FLAC, MP3, MP4, WAV, WMA, AAC, JSON, XML, CSV, etc. The audio data may be collected or provided automatically or in response to a request. For example, a user of the computing device 110 may wish to monitor for glass breakage events in the biomanufacturing process machine 160 over a period of time. Accordingly, one or more audio sensors 162 may collect audio data over that period and provide it to the computing device 110. In some aspects, the audio sensor 162 may collect audio data in response to the operation of the biomanufacturing process machine 160. For example, the audio sensor 162 may start collecting audio data when the biomanufacturing process machine 160 is powered on / operation begins, and may continue collecting audio data until the biomanufacturing process machine 160 is powered off / operation ends. In some embodiments, one or more audio sensors 162 may include a database of data / information regarding product quality, or may be configured to receive data / information regarding product quality, such as via user input.
[0025] The biomanufacturing process machine 160 may include one or more devices (not shown) used in the manufacture of biological agents (e.g., pharmaceuticals as discussed in the Background section). The biomanufacturing process machine 160 may be configured to be controllable via manual or automatic input. In some embodiments, the biomanufacturing process machine 160 may be configured to receive such control input locally, such as via a user input device local to the biomanufacturing process machine 160. In some embodiments, the biomanufacturing process machine 160 is configured to receive control input remotely, such as from the computing device 110 (e.g., via the network 170). The control input may include operation commands such as commands to power on / start operation of the biomanufacturing process machine 160. In some aspects, the biomanufacturing process machine 160 may end operation in response to one or more of the following. That is, (i) the biomanufacturing process machine 160 completes the production of a biological agent (e.g., all batches of a pharmaceutical are complete), (ii) a glass breakage event is detected, or (iii) a manual command to end operation is received.
[0026] The training audio data source 150 generally may correspond to one or more biomanufacturing processes for producing one or more biological agents using the biomanufacturing process machine 160 (e.g., it may be collected during its execution) and includes training audio data. The training audio data may represent (i) training ambient sound, (ii) training glass sound, and (iii) training glass breakage sound, and may be collected (by the computing device 110 or another device / system) using the audio sensor 162 or other similar sensors. More specifically, in some aspects, (i) the training ambient sound includes sounds caused by the operation of a machine (e.g., the biomanufacturing process machine 160), (ii) the training glass sound includes sounds caused by a first glass surface (e.g., a first pharmaceutical container) contacting either a second glass surface (e.g., a second pharmaceutical container) or a non - glass surface, and / or (iii) the training glass breakage sound includes sounds caused by either a glass crack or a glass breakage. Exemplary training labels may include "glass crack", "glass breakage", "glass grinding", "glass contacting metal", "glass contacting plastic", "glass contacting glass", "glass contacting others", "ambient machine", "others in the surroundings", etc. In some aspects, at least a portion of the training audio data includes sounds from the biomanufacturing process machine 160. However, in some aspects, all of the training audio data is collected using different biomanufacturing process systems. The training audio data may include data from a biomanufacturing process having the same scale / size, settings / parameters, equipment model, etc. as the biomanufacturing process machine 160, or data from a biomanufacturing process having a different scale / size, settings / parameters, equipment model, etc. from the biomanufacturing process machine 160.In some embodiments, system 100A omits training audio data source 150 and instead may receive training audio data locally, such as via user input in computing device 110 (e.g., a user providing training audio data on a portable memory drive). In some examples, the training audio data includes training glass sounds or training glass breakage sounds that are combined (or mixed) with second audio data that includes ambient sounds that do not include any training glass sounds or training glass breakage sounds. In some aspects, training audio data source 150 or computing device 110 combines (or mixes) the first audio sound and the second audio sound.
[0027] Computing device 110 may include a single computing device or multiple computing devices that are either located in the same place or separated from each other. Computing device 110 is generally configured to apply audio data over an interest period as an input to a model trained with training audio data to determine whether a glass breakage event occurred during the interest period by processing the audio data using a machine learning model. The components of computing device 110 may be interconnected via an address / data bus or other means. The components included in computing device 110 may include processing unit 120, network interface 122, display 124, user input device 126, and memory 128, which are discussed in more detail below.
[0028] The processing unit 120 includes one or more processors, each of which may be a programmable microprocessor that executes software instructions stored in the memory 128 to perform some or all of the functions of the computing device 110 as described herein. Alternatively, one or more of the processors within the processing unit 120 may be other types of processors (e.g., application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), etc.).
[0029] The network interface 122 may include any suitable hardware (e.g., front-end transceiver hardware), firmware, or software configured to communicate with external devices or systems (e.g., audio sensor 162, biomanufacturing process machine 160, training audio data source 150, etc.) via the network 170 using one or more communication protocols. For example, the network interface 122 may be or include an Ethernet interface.
[0030] The display 124 may use any suitable display technology (e.g., LED, OLED, LCD, etc.) to present information to the user, and the user input device 126 may be a keyboard or other suitable input device. In some embodiments, the display 124 and the user input device 126 are integrated within a single device (e.g., a touch screen display). Generally, the display 124 and the user input device 126 may be combined such that the user can interact with the graphical user interface (GUI) or other (e.g., text) user interface provided by the computing device 110 (e.g., for the purpose of notifying the user of glass breakage, etc.).
[0031] Memory 128 includes one or more physical memory devices or units including volatile or non-volatile memory, and may or may not include memory located in various computing devices of computing device 110. Any suitable one or more memory types such as read-only memory (ROM), solid-state drive (SSD), hard disk drive (HDD), etc. may be used. Memory 128 may store instructions for one or more software applications included in a glass breakage (GB) application 130 that can be executed by processing unit 120. In exemplary system 100A, GB application 130 includes a data collection unit 132, a model training unit 134, a user interface unit 136, a glass breakage detection unit 138, and a notification unit 140. Units 132-140 may be separate software components or modules of GB application 130, or may simply represent the functions of GB application 130 that are not necessarily divided among different components / modules. For example, in some embodiments, data collection unit 132 and user interface unit 136 are included in a single software module. Moreover, in some embodiments, units 132-140 may be distributed among multiple copies of GB application 130 (e.g., executed in different components in computing device 110), or among different types of applications stored and executed in one or more devices of computing device 110.
[0032] The data collection unit 132 is generally configured to receive data (e.g., audio data, operator instructions, etc.). In some embodiments, the data collection unit 132 receives training audio data for a biomanufacturing process for producing a biological agent (e.g., including historical audio data of multiple instances of the biomanufacturing process and corresponding historical audio data). The data collection unit 132 may receive the training audio data via, for example, a training audio data source 150, user input received via a user interface unit 136 by a user input device 126, or other suitable means. In some embodiments, the data collection unit 132 may receive audio data via, for example, an audio sensor 162, user input received via a user interface unit 136 by a user input device 126, or other suitable means. In some embodiments, the computing device 110 may, for example, receive an indication in the data collection unit 132 that a biomanufacturing process has been initiated, and one or more components of the computing device 110 may, for example, begin monitoring audio data provided by an audio sensor 162. In some aspects, the data collection unit 132 may apply preprocessing, e.g., an amplification gain applied to one or both of a training audio signal or an audio signal, to the received audio data, where the training audio signal and the audio signal correspond to the training audio data and the audio data, respectively.
[0033] The model training unit 134 is generally configured to generate, train, or apply a model. The model may be any suitable model for detecting glass breakage events in audio data. In some embodiments, and as further described below, the model may be trained using at least some of the system 100A, or in some embodiments, the model may be pre-trained (i.e., trained before being acquired by the computing device 110). The model may be trained using training audio data representing (i) training ambient sounds, (ii) training glass sounds, and (iii) training glass breakage sounds. In some embodiments, the model may include a statistical model that may be parametric, non-parametric, or semi-parametric. One suitable example of a statistical model that may be included in the model is a linear regression model. In other embodiments, the model includes a machine learning model. For example, the model may employ a neural network such as a convolutional neural network or a deep learning neural network. Other examples of machine learning models in the model are models that use support vector machine (SVM) analysis, k-nearest neighbor analysis, naive Bayes analysis, clustering, reinforcement learning, or other machine learning algorithms or techniques. The machine learning model included in the model may identify and recognize patterns in the training data to facilitate making predictions for new data. The model training unit 134 may train the model using training audio data that may be received from the training audio data source 150.
[0034] The user interface unit 136 is generally configured to receive user input. In one embodiment, the user interface unit 136 generates a user interface for presentation via the display 124 and may receive user input training audio data to be used by the model training unit 134 when training the model via the user interface and the user input device 126. In another embodiment, the user interface unit 136 may receive an input to initiate the operation of the biomachining process machine 160 or the audio sensor 162 via the user interface and the user input device 126. The user interface unit 136 may also be used to display information. For example, the user interface unit 136 may be used to display an indication of whether a glass breakage event has been detected.
[0035] The glass breakage detection unit 138 may also apply or access a model trained by the model training unit 134 (or otherwise obtained by the computing device 110 as a pre-trained model) when identifying whether a glass breakage event has occurred. In some embodiments, the glass breakage detection unit 138 begins monitoring for glass breakage events in response to the data collection unit 132 receiving audio data. The glass breakage detection unit 138 may monitor the audio data as it is collected by the data collection unit 132 in real time, substantially in real time (i.e., with some buffering), or asynchronously (i.e., after the audio data has been fully collected over a period of interest). It should be understood that if the glass breakage detection unit 138 is considered to detect whether a glass breakage event has occurred, this also includes detecting whether a glass breakage event is occurring (since the glass breakage detection unit 138 may monitor in real time).
[0036] The notification unit 140 is generally configured to notify the user of either (i) whether a glass breakage event has occurred or (ii) that a glass breakage event has occurred. The notification unit 140 may display the notification in cooperation with the user interface unit 136. The notification unit 140 may send an electronic message (e.g., email, text, etc.) along with the notification to the user of the computing device 110 or an external computing device. In some aspects, the notification unit 140 may send a control signal to stop the operation of the biomanufacturing process machine 160 when a glass breakage event is detected by the glass breakage detection unit 138. In some aspects, the instructions may be stored (e.g., in the memory 128) along with other data (such as operating data, etc.) related to the biomanufacturing process machine 160 that may be useful in diagnosing the cause of the glass breakage event, in some cases.
[0037] In some embodiments, some or all of the functions of the GB application 130 may be provided by a third party (i.e., not on the computing device 110). For example, machine learning may be hosted by a third party, and the GB application 130 may remotely access a machine learning model by sending data (e.g., audio data) and receiving data (e.g., an indication of whether a glass breakage event has been detected). In such an embodiment, the function of the glass breakage detection unit 138 may be hosted by a third party. Looking at different embodiments, the machine learning model may be trained by a third party, and the GB application 130 may receive the machine learning model remotely from the third party (e.g., by the computing device 110 receiving one or more elements of the machine learning model such as weights or architecture). In such an embodiment, the function of the model training unit 134 may be hosted by a third party. In other embodiments, one or more instances of the functions of any of the units 132 - 140 may be hosted by a third party, for example, on a remote server accessible via the network 170.
[0038] FIG. 1B shows an exemplary system 100B representing an embodiment in which the computing device 110 of FIG. 1A may be an Internet of Things (IoT) device and may be communicatively coupled to one or more other devices (e.g., audio sensor 162). For example, the computing device 110 may be the same as or similar to the architecture shown in FIG. 1B, in which the computing device 110 is a Raspberry Pi that includes a USB microphone (which may be one or more of the audio sensors 162), any coral USB TPU accelerator, and an interface for hosting Amazon Web Services (AWS) services including SnS, IoT Core, or an S3 bucket. The Raspberry Pi may be selected for the computing device 110 for its technical specifications that enable the recording and storage of live audio data while providing remote access and a simple user interface. As correctly recognized, the Raspberry Pi also presents certain cost advantages and is well supported by cloud-sourced software and technical support. In addition, the Raspberry Pi can be easily customized and additional features can be retrofitted. As further illustrated, the coral USB TPU accelerator may be available to host and process deep learning models, enabling real-time audio data processing.
[0039] Exemplary glass breakage event Figure 2 shows an exemplary scenario 200 in which a glass breakage event occurs. As shown, scenario 200 occurs in a biomanufacturing process machine 260 (which may be the same as or similar to the biomanufacturing process machine 160 of FIG. 1A). As shown, the events of scenario 200 occur in proximity to audio sensors 262A and 262B (which may be the same as or similar to the audio sensors 162 of FIG. 1A) communicatively coupled to an IoT device 210 (which may be the same as or similar to the computing device 110 of FIG. 1A or as shown in system 100B).
[0040] The biomanufacturing process machine 260 is shown as performing the filling of containers (such as glass vials as shown) with pharmaceuticals (such as liquid pharmaceuticals as shown). More specifically, some containers (including container 270A) are shown as not filled with pharmaceuticals, some containers (including container 270B) are shown as filled with pharmaceuticals, and container 270C is shown as partially filled with pharmaceuticals. Further, a damaged unfilled container 272A is shown together with corresponding fragments 272B. As shown, the damaged container 272A or the fragments 272B may pose a safety hazard to, for example, an operator of the biomanufacturing process machine 260 and / or an end user / consumer of the pharmaceuticals.
[0041] The two audio sensors 262A and 262B are shown as being included in proximity to the events of scenario 200. The standard microphone 262A may be farther from the dispenser 280 than the surface microphone 262B. The standard microphone 262A may be an omnidirectional microphone. The surface microphone 262B may be a "contact" microphone that senses mechanical audio vibrations through direct contact with an object. The surface microphone 262B may be seated on a disk (e.g., a metal disk).
[0042] Audio data corresponding to the events of scenario 200 where container 272B is damaged may be captured by two audio sensors 262A and 262B. Additionally, the two audio sensors 262A and 262B may capture ambient audio data corresponding to the operating sounds of the biomanufacturing process machine 260. Further, the two audio sensors 262A and 262B may capture the sound of glass in contact with glass or other surfaces (such as metal, wood, plastic, etc.) that the glass contacts.
[0043] The two audio sensors 262A and 262B may be communicatively coupled to the IoT device 210 (e.g., via a wired connection or WLAN, etc.). As shown, the IoT device 210 may be external to the biomanufacturing process machine 260, while in some embodiments, the IoT device 210 may be integrated into the biomanufacturing process machine 260. The IoT device 210 may detect a glass breakage event of scenario 200 involving container 272A and debris 272B, and in response, may instruct the biomanufacturing process machine to stop operating. As shown, the dispenser 280 that dispenses pharmaceuticals may stop dispensing pharmaceuticals in response to the detection of a glass breakage event of scenario 200 (e.g., via the GB application 130). By quickly stopping the dispensing of pharmaceuticals from the dispenser 280 and the operation of the biomanufacturing process machine 260 as a whole, it may be possible to reduce the adverse effects caused by the events of scenario 200 (such as those described in the background art section).
[0044] Exemplary audio data processing Figure 3 shows a data representation example 300 that processes raw audio data 310 into a mel spectrogram 320 and a patch 330. The data representation 300 may correspond to the same or similar devices / apparatus as those discussed above in relation to the system 100A / B. For example, the computing device 110 may implement / generate part or all of the data representation 300 (e.g., the GB application 130 using the audio data collected from the biomachining process machine 160 by the audio sensor 162). Further, the data representation 300 may correspond to a glass breakage event detection process that may be similar to the glass breakage event of scenario 200 of FIG. 2.
[0045] The raw audio data 310 may be collected via one or more audio sensors (e.g., audio sensor 162 or two audio sensors 262A and 262B) that may be a standard microphone, a surface microphone, or other suitable microphone for collecting audio data. As shown, the raw audio data 310 may be sampled at 16 kHz mono as input, while in other embodiments, other suitable sampling may be performed.
[0046] As shown, the raw audio data 310 may be processed (e.g., by the computing device 110 using the data collection unit 132) to calculate a spectrogram. In some embodiments, a Fourier transform (e.g., a short-time Fourier transform) is performed to calculate the mel spectrogram 320 from the raw data 310. The spectrogram is mapped to 64 mel bins covering the range of 125 - 7500 Hz, and then the logarithm of the mel spectrum 320 is taken (a small buffer is added so as not to take the logarithm of zero), thereby converting it into a stabilized logarithmic mel spectrogram (i.e., the mel spectrogram 320).
[0047] The Mel spectrogram 320 may be processed by the patch 330 (e.g., by the computing device 110 using the data collection unit 132). The patch 330 spans 0.96 seconds with a 50% overlap as shown in the figure. The patch 330 may correspond to a plurality of scores that may be generated by a model (e.g., using the GB application 130), and each score corresponds to a possible classification for a single patch of the patch 330. The patch 330 may have various classification possibilities such as, for example, "glass crack", "glass breakage", "glass in contact with metal", "glass in contact with plastic", "surrounding", etc. It is worth noting that the scores do not necessarily have to be calibrated and may ultimately be dimensionless. For example, a score of 0.5 for a particular classifier does not necessarily mean a 50% probability for each classifier detected by the model, and not all of the output scores across each class necessarily sum to 1.
[0048] Overlaid on the data representation 300 are the time window 340A and the time window 340B. The time window 340A shows the first 0.96 - second patch of the audio, which is temporally coincident in both the raw waveform and the Mel spectrogram. Note that the scores included in the patch 330 for the time window 340B are significantly different from the scores of the time window 340A despite the 50% overlap between the time window 340A and the time window 340B.
[0049] Exemplary machine - learning model Figures 4A and 4B show exemplary models 400A and 400B that may be generated, trained, and / or used (e.g., executed or accessed) by the computing device 110 of FIG. 1A to predict that a glass breakage event has occurred. Models 400A and 400B may be used to predict glass breakage events, such as the glass breakage event of scenario 200 shown in FIG. 2, using audio data that may be processed in a similar manner as described with respect to the data representation 300 of FIG. 3. Models 400A and 400B may identify whether a glass breakage event occurred during a period of interest by processing the audio data. As previously explained, models 400A and 400B may be statistical or machine learning models, but as shown, models 400A and 400B correspond to machine learning models.
[0050] In embodiments where models 400A and 400B are machine learning models, models 400A and 400B may be universal (i.e., applicable to all situations) or more specialized (i.e., different models for different situations). The machine learning model may be trained using a supervised or unsupervised machine learning program or algorithm. The machine learning program or algorithm may employ a neural network, which may be a convolutional neural network (CNN), a deep learning neural network, or a composite learning model or program that learns in two or more features or feature datasets in a particular area of interest. In one embodiment, an adversarial generative neural network may be used. The machine learning program or algorithm may also include regression analysis, support vector machine (SVM) analysis, decision tree analysis, random forest analysis, k-nearest neighbor analysis, naive Bayes analysis, clustering, reinforcement learning, or other machine learning algorithms or techniques. In some embodiments, due to the processing power requirements for training the machine learning model, the selected model may be trained using additional computing resources (e.g., cloud computing resources) based on data provided by an external source (e.g., training audio data source 150). The training data may not be labeled or may be labeled by a person or the like. The training of the machine learning model may continue until at least one of the machine learning models is verified and meets the selection criteria to be used as a prediction model for determining whether a glass breakage event has occurred. In one embodiment, the machine learning model may be verified using a second subset of the training data to identify the accuracy and robustness of the algorithm. Such verification may include applying the machine learning model to the training data of the second subset of the training data to predict whether a glass breakage event has occurred in the second subset of the training data. The machine learning model may then be evaluated to determine whether the machine learning model performance is sufficient based on the verification stage prediction.The sufficiency criteria applied may vary depending on the size of the training data available for training, the performance of previous iterations of the machine learning model, or user-specified performance requirements.
[0051] For maximum effectiveness, models 400A and 400B should be computationally inexpensive to enable real-time or near-real-time detection of glass breakage events (e.g., at the edge, i.e., by the device itself, to process and classify live audio data, or to send audio data to the cloud for real-time processing). It is generally preferred that models 400A and 400B maximize predictive power within the computational constraints driven by the device specifications. Since glass breakage events are extremely rare and the impact of detection can be significant (e.g., halting a biomanufacturing process), models 400A and 400B should preferably have an exceptionally strong predictive power to avoid false positive events (classifying nominal environmental noise as a glass breakage event). In some aspects, models 400A and 400B should achieve a false positive rate close to 0% while still maintaining a reasonable true positive rate (the rate at which the model correctly classifies actual glass breakage events). These basic design criteria motivate choosing the open-source CNN called "YAMNet" as a suitable option for a feasibility study.
[0052] In some embodiments, CNNs are well-suited for machine vision applications due to their pattern recognition capabilities. As correctly recognized, CNNs generally differ from standard multi-layer perceptrons (MLPs) by using convolutional layers in which a number of matrices, generally referred to as filters, are convolved with an input image to produce a tensor representing a new image with any number of channels. This new tensor can then be convolved with a new set of filters in another convolutional layer to produce yet another tensor. This process is repeated for each layer defined in the CNN. In a typical classification task, the final output of the CNN is a set of vectors representing the predicted likelihoods for each class. The filters of the CNN can be trained and selected based on recognizing distinct patterns such as edges, corners, or shapes.
[0053] YAMNet is a convolutional neural network (CNN) that utilizes the MoblieNet depthwise separable convolutional architecture, an efficient convolutional network architecture designed for mobile and embedded vision applications, using two hyperparameters. More specifically, two alternative filters replace the standard convolutional filter, namely, the depthwise convolution that applies a single filter to each input, and the 1×1 pointwise convolution. YAMNet is an audio classification model that incorporates the MobileNet (depthwise separable CNN) architecture pre-trained on audio data to predict different audio events. In some embodiments, YAMNet may not require any feature extraction before passing the audio data to the model, since the model has a feature extraction layer incorporated into the model. The feature extraction layer may convert the audio data into a spectrogram, which is then passed to the MobileNet. As discussed previously, it may be necessary to preprocess the audio data in order to train YAMNet using the audio signal. One way to preprocess the audio data is to apply the STFT to the audio data and convert the output from the STFT into a spectrogram or mel spectrogram using a logarithmic scale (as shown, for example, in Data Representation 300).
[0054] According to some of the basic design criteria, by splitting the filtering and combining steps of a typical convolution operation, it is possible to significantly reduce the computational cost. MobileNet produces roughly the same results as the general neural nets AlexNet, VGG16, and GoogleNet, but is a fraction of the size and computationally intensive. Thus, as would be correctly recognized by those skilled in the art, the YAMNet model architecture that employs MobileNet may have certain technical advantages in IoT application examples where computational efficiency is often limited by the device footprint (e.g., by IoT device 210 of FIG. 2). More specifically, YAMNet is pre-trained on AudioSet, an ontology of 521 audio event classes, and a collection of over 2 million human-labeled 10-second audio clips from YouTube videos. The event classes range from musical instruments to natural sounds to a wide variety of animal noises. The categories "glass", "clinking", "crushing", and "cracking" are particularly important for detecting glass breakage events.
[0055] There are two variations of YAMNet. These are the standard YAMNet model and a TensorFlow Lite model called YAMNet-lite, which is further tuned for mobile applications. The YAMNet model takes raw audio data that is processed according to the data representation 300 in Figure 3 (sampled at 16kHz mono as input, taking the STFT to compute the spectrogram, and mapping the spectrogram to 64 mel bins covering the range 125 - 7500Hz to transform the spectrogram into a stabilized log mel spectrogram, and then taking the logarithm of the mel spectrum (a small buffer is added to avoid taking the logarithm of zero)). A 0.96-second time window with 50% overlap (as also shown in the data representation 300 in Figure 3) is fed to MobileNet, returning 521 scores in the range from 0 to 1 for each class in the AudioSet ontology. Importantly, the output scores are not calibrated and ultimately have no units. For example, a score of 0.5 for a particular classifier is not interpretable as a 50% probability of each classifier detected by the model, nor do all the output scores across all classes sum to 1.
[0056] The YAMNet-lite model operates similarly to the standard YAMNet model but has two main differences. The model is quantized and thus retrained with the Relu6 non-linearity instead of the Relu activation, and the input must be a fixed 0.975-second frame of 16kHz mono audio. Logically, the output of the YAMNet-lite model is a single vector of 521 classification scores. As would be correctly recognized by those skilled in the art, the YAMNet-lite model may have additional technical advantages for use in optimizing the computational cost by IoT devices such as the IoT device 210 in Figure 2.
[0057] Referring to FIG. 1B, the Raspberry Pi of FIG. 1B hosts a collection of software modules (e.g., Python code) that interface with peripheral devices such as modules for continuously recording a time window (e.g., 0.975 seconds) of raw 16-bit PCM audio to be processed by the YAMNet model, makes calls to various AWS clients to send text messages to subscribers, saves audio recordings to an s3 bucket in the AWS cloud, and may deliver MQTT messages to AWS IoT Core. One or more modules of the Raspberry Pi may also interface with the audio stream to locally save the audio as a backup on the Raspberry Pi. One or more modules of the Raspberry Pi may also interface with a USB microphone to initialize continuous audio streaming and recording capabilities. One or more modules of the Raspberry Pi may also initialize the YAMNet-lite model as well as the model's hyperparameters, and the model can be hosted locally on the Raspberry Pi or on a Coral USB accelerator for increased processing power. One or more modules of the Raspberry Pi may also initialize connections to various AWS clients. The various modules of the Raspberry Pi may also cooperate to operate so as to continuously record and classify time windows (e.g., 0.975 seconds) of live audio using the YAMANet model architecture. The communication protocol via AWS is established to enable the device to store audio in the cloud. Additionally, when conditional heuristics are established (e.g., when the output of the YAMANet model is >0.2 for the "glass" classifier), the device can be configured to automatically send an alert to all relevant parties that a potential glass breakage event has occurred.
[0058] Models 400A and 400B provide a visual representation of the YAMNet architecture optimized for mobile audio detection applications. Assuming the predictive power of YAMNet's pre-trained audio classifier for nominal audio data, models 400A and 400B may be trained via transfer learning (i.e., improvement of learning in a new task by transfer of knowledge from a previously learned related task). In connection with models 400A and 400B, transfer learning may be employed by the YAMANet model architecture, including additional custom-trained classifiers. Model 400A is a visual representation of the YAMNet-lite model architecture and receives two new classifiers as shown in model 400B. Models 400A and 4000B return a 2×1 vector of classification scores. The "Process A" classifier of model 400B is trained on Process A audio data, while the "Breakage Event" classifier of model 400B is trained on glass breakage event sounds. Process A refers to audio data from an integrated downstream filling and finishing process (an example of a biomanufacturing process) as a representative environment where glass containers are handled and thus glass breakage may occur.
[0059] Exemplary Glass Breakage Event Detection Performance Figures 5A and 5B show exemplary experimental data 500A and 500B representative of the performance of one embodiment of the techniques of the present disclosure in identifying whether a glass breakage event occurred during a period of interest by processing audio data using a machine learning model. The machine learning model corresponding to data 500A may be trained or used via system 100A / B to detect glass breakage events (e.g., the glass breakage event of scenario 200) using a model that may be the same as or similar to models 400A or 400B that process audio data that may be pre-processed according to data representation 300.
[0060] In some aspects, the audio signal of the audio data used by the machine learning model corresponding to data 500A or 500B may have a gain added (e.g., by IoT device 210 or computing device 110). Adding a gain to the audio signal may affect the performance of the machine learning model when predicting whether a glass break event occurred over the period that the audio data contains. Adding a gain may create a separation between the glass break sound and other sounds and may improve the ability of the machine learning model to distinguish the sounds. Identifying a preferred amount of gain may be done using an application-specific iterative approach that has sufficient criteria for performance.
[0061] As shown in data 500A, a gain level of +16 decibels was selected (e.g., by an operator) for the audio signal of the audio data input into the machine learning model. Data 500A corresponds to the performance of a machine learning model that classifies a closed data set when trained with the AudioSet data (training audio data) described above. A receiver operating characteristic curve (ROC) 510A for the machine learning model, as well as a binary confusion matrix 520A for a threshold score of 0.003, are included in data 500A. As demonstrated in data 500A, the machine learning model approaches a perfect classifier. At a score threshold of 0.003 or higher, the true positive rate is 90% and the false positive rate is 1%. This is a strong result for demonstrating the prediction performance of the machine learning model. The main limitation in data 500A is the simulated environment of the audio data. The glass containers in the actual biomanufacturing process are used to generate the audio data, but the containers are empty (e.g., not filled with biopharmaceuticals). In addition, since the audio recordings were not obtained from the production line, the ambient sound of the environment where a glass break event may occur is not accurately simulated.
[0062] Data 500B addresses some of the limitations of Data 500A by predicting the occurrence of glass breakage events for audio data representing (i) ambient sound (e.g., the sound of a laboratory, equipment, machinery, etc.), (ii) glass sound (e.g., glass coming into contact with another material which may or may not be glass), and (iii) training glass breakage sound (e.g., cracking, breaking, shattering, etc.). Data 500B corresponds to a machine learning model custom-trained using the architecture of Model 400B with two classifiers (i.e., the "Process A" classifier is trained on Process A audio data and the "Breakage Event" classifier is trained on glass breakage event noise). Data 500B demonstrates that the machine learning model performs almost perfectly when classifying breakage events and Process A audio data with a false positive rate reaching 0% for a score threshold > 0.02. Additionally, variations in microphone gain can be ignored. Thus, Data 500B demonstrates the success of this technique in distinguishing the sound of glass breaking from other sounds (i.e., ambient sound and glass sound).
[0063] Exemplary Flow Diagram FIG. 6 is a flowchart showing an exemplary method 600 for identifying a glass breakage event (e.g., the glass breakage event of scenario 200). Method 600 may be implemented by one or more components of system 100A / B, such as processing unit 120, when also implementing GB application 130 and possibly biomachining process machine 160 (which may be operating the biomachining process). Method 600 may be executed as part of a process that is the same as or similar to data representation 300. Method 600 may receive audio data (e.g., from biomachining process machine 160 and audio sensor 162) over a period of interest. Exemplary method 600 may include the following elements: (1) accessing or obtaining a machine learning model trained using training audio data (block 602), (2) obtaining audio data (block 604), (3) determining whether a glass breakage event has occurred by processing the audio data using the machine learning model (block 606), and (4) indicating that a glass breakage event has occurred if it is determined that a glass breakage event has occurred.
[0064] The trained machine learning model obtained (e.g., downloaded, generated, trained, etc.) or accessed (e.g., accessing a third-party remote server hosting the machine learning model) in block 602 may be trained using training audio data such as the training audio data included in the training audio data source 150 (e.g., as described above). In some aspects, the training audio data may include one or more spectrograms that may be the same as or similar to the mel spectrogram 320. The training audio data may represent (i) training ambient sound, (ii) training glass sound, and (iii) training glass breakage sound. More specifically, (i) the training ambient sound may include sounds caused by the operation of a machine, (ii) the training glass sound may include sounds caused by a first glass surface contacting either a second glass surface or a non-glass surface, and (iii) the training glass breakage sound may include sounds caused by either a glass crack or glass breakage. In some aspects, the machine learning model is a convolutional neural network or is trained using a supervised learning technique. The training audio data may include (i) ambient sound labels, (ii) glass sound labels, and (iii) breakage sound labels. In some aspects, the machine implements a biomanufacturing process (e.g., via the biomanufacturing process machine 160), and the training glass sound and the training glass breakage sound are generated by one or more containers of one or more pharmaceuticals. In some embodiments, obtaining the machine learning model in block 602 includes receiving a pre-trained model (i.e., trained prior to being obtained by, e.g., the system 100) or generating / training the machine learning model (e.g., by the system 100).The machine learning model may be obtained internally (e.g., by accessing files / programs / data / information stored locally in a computing system such as computing device 110) or externally (e.g., by receiving the model from an external source such as receiving the machine learning model at computing device 110 via network 170).
[0065] Block 604 may include obtaining audio data over an interest period. The audio data may be input or provided by a user via a user interface (e.g., using user input device 126) or by collecting the audio data as data (e.g., via data collection unit 136). The audio data may be collected by an audio sensor (e.g., audio sensor 162) that may be inside, outside, or in the vicinity of a biomanufacturing process machine (e.g., biomanufacturing process machine 160). In some aspects, similar to the training audio data, the audio data may include one or more spectrograms that may be the same as or similar to mel spectrogram 320, and obtaining the audio data over the interest period includes generating a spectrogram of the audio data from the raw audio data using a Fourier transform (e.g., SFTF). In some aspects, the audio data may be preprocessed by (i) as described in data representation 300 or (ii) by one or more processors, such as by iteratively adjusting the gain of an amplification applied to one or both of the training audio signal or the audio signal until a performance threshold is met, where the training audio signal and the audio signal respectively correspond to the training audio data and the audio data.
[0066] Block 606 may include determining whether a glass breakage event occurred during the period of interest (which may include a glass breakage event that has occurred or is occurring) by processing audio data using a machine learning model. Determining that a glass breakage event has occurred may be a single binary value (e.g., 0 indicating that no glass breakage event has occurred, 1 indicating that a glass breakage event has occurred), or a plurality of values representing the probability that a glass breakage event has occurred (e.g., a score representing a range from 0% to 100% probability). In some aspects, the machine learning model may be a CNN such as YAMNet or YAMNet-lite (as described herein), a linear regressor, a random forest model, a support vector machine (SVM) analysis, a k-nearest neighbor analysis, a naive Bayes analysis, a clustering, or a model using reinforcement learning, or another suitable machine learning model. In some embodiments, a statistical model such as a linear regression model or some other suitable statistical model may be used in addition to or as an alternative to the machine learning model for identifying glass breakage events.
[0067] In some embodiments, method 600 may end at block 608, which may include indicating that a glass breakage event has occurred (e.g., via a computing device such as computing device 110). In some embodiments, the indication may be visual (e.g., displayed on display 124), auditory (e.g., played on a speaker), tactile, or any other suitable notification method. In some embodiments, the notification may include electronic messaging such as sending a text message or an email message to a user (e.g., an operator of a biomanufacturing process machine). In some embodiments, the indication may optionally be stored with other data (such as operational data) related to the biomanufacturing process that may be useful in diagnosing the cause of the glass breakage event. In some embodiments, after identifying the glass breakage event, block 607 may include automatically stopping the operation of the biomanufacturing process machine.
[0068] In some embodiments, method 600 may be performed either fully automated, for example, by one or more processors (such as a CPU or GPU) executing instructions stored on one or more non-transitory computer-readable storage media (such as volatile or non-volatile memory, read-only memory, random access memory, flash memory, electronically erasable programmable read-only memory, or one or more other types of memory). Method 600 may use any one or more of the components, processes, or techniques of FIGS. 1-5.
[0069] Exemplary flowchart FIG. 7 is a flowchart showing an exemplary method 700 for training a machine learning model for identifying a glass breakage event (e.g., the glass breakage event of scenario 200). The method 700 may be implemented by one or more components of the system 100A / B, such as the processing unit 120, or by different devices or systems (e.g., a third-party server that develops and hosts the machine learning model), when the GB application 130 and optionally the biomachining process machine 160 (which may be operating the biomachining process) are also implemented. The method 700 may be executed as part of a process that is the same as or similar to the data representation 300. The method 700 may receive training audio data (e.g., from the training audio data source 150) over a period of interest. The method example 700 may include the following elements: (1) obtaining training audio data (block 702), (2) classifying the training audio data into a plurality of subsets (block 704), and (3) generating a machine learning model for identifying glass breakage events using the classified subsets (block 706).
[0070] Block 702 may include obtaining training audio data (e.g., in a computing device from a training audio data source 150 via a data collection unit 132, etc.). In some aspects, the training audio data may include one or more spectrograms that are the same as or similar to the mel spectrogram 320. The training audio data may represent (i) training ambient sound, (ii) training glass sound, and (iii) training glass breakage sound. More specifically, (i) the training ambient sound may include sounds caused by the operation of a machine, (ii) the training glass sound may include sounds caused by a first glass surface contacting either a second glass surface or a non-glass surface, and (iii) the training glass breakage sound may include sounds caused by either a glass crack or glass breakage. In some aspects, the training audio data includes (i) an ambient sound label, (ii) a glass sound label, and (iii) a breakage sound label. In some aspects, the machine implements a biomachining process (e.g., via a biomachining process machine 160), and the training glass sound and the training glass breakage sound are generated by one or more containers of one or more pharmaceuticals. In some aspects, the training audio data may be preprocessed by repeatedly adjusting the gain of amplification applied to the training audio signal (e.g., in a manner the same as or similar to the data representation 300) until a performance threshold is met, where the training audio signal corresponds to the training audio data.
[0071] Block 704 may include classifying the training audio data into a plurality of subsets, each corresponding to a different actual result data (meaning a specific result or a range of results), the subsets including (i) at least one subset representing training ambient sound, (ii) at least one subset representing training glass sound, and (iii) at least one subset representing training glass breakage sound. In some aspects, the classification of the training audio data into the plurality of subsets may be performed by a computing device (e.g., computing device 110 using GB application 130 and model training unit 134) based on, for example, labels associated with the training audio data.
[0072] In some aspects, method 700 may end in block 706 by generating a machine learning model for identifying glass breakage events using the classified subsets of the training audio data. In some aspects, the machine learning model may be a CNN such as YAMNet or YAMNet-lite (as described herein), a linear regressor, a random forest model, a support vector machine (SVM) analysis, a k-nearest neighbor analysis, a naive Bayes analysis, a clustering, or a model using reinforcement learning, or another suitable machine learning model. In some embodiments, a statistical model such as a linear regression model or some other suitable statistical model may be used in addition to, or as an alternative to, the machine learning model for identifying glass breakage events. The machine learning model may be stored using a computing device such as computing device 110 (e.g., specifically using memory 128).
[0073] In some aspects, method 700 may be performed either by fully automating, e.g., by one or more processors (e.g., a CPU or GPU) that execute instructions stored in one or more non-transitory computer-readable storage media (e.g., volatile memory or non-volatile memory, read-only memory, random access memory, flash memory, electronically erasable programmable read-only memory, or one or more other types of memory). Method 600 may use any one or more of the components, processes, or techniques of FIGS. 1-6.
[0074] Additional Considerations Some of the drawings described herein show exemplary block diagrams having one or more functional components. It will be understood that such block diagrams are for illustrative purposes only and that the devices described and shown may have more, fewer, or alternative components than those illustrated. Also, in various aspects, components (and the functions provided by each component) may be associated with any suitable component or alternatively integrated as part thereof.
[0075] Some aspects of the present disclosure relate to a non-transitory computer-readable storage medium having instructions / computer-readable code for performing various computer-implemented operations. As used herein, the term "instructions / computer-readable storage medium" includes any medium that can store or encode a series of instructions or computer code for performing the operations, techniques, and techniques described herein. The medium and computer code may be specially designed and constructed for the purposes of the aspects of the present disclosure, or may be of the kind known and available to those skilled in the art of computer software technology. Examples of computer-readable storage media include magnetic media such as hard disks, floppy disks, magnetic tapes, etc., optical media such as CD-ROMs, holographic devices, etc., magneto-optical media such as optical disks, ASICs, programmable logic devices ("PLDs"), and hardware devices specially configured for storing and executing program code, such as ROMs and RAM devices, but are not limited thereto.
[0076] Examples of computer code include files containing machine code generated by a compiler and high-level code that is executed by a computer using an interpreter or compiler. For example, one aspect of the present disclosure may be implemented using Java, C++, or other object-oriented programming languages and development tools. Additional examples of computer code include encryption code and compression code. Further, aspects of the present disclosure may be downloaded as a computer program product and transferred from a remote computer (e.g., a server computer) to a requesting computer (e.g., a computer or a different server computer) via a transmission channel. Another aspect of the present disclosure may be implemented in hardwired circuitry as an alternative to, or in combination with, machine-executable software instructions.
[0077] As used herein, the singular terms "a", "an", and "the" may include plural referents unless the context clearly dictates otherwise. This specification and the claims that follow are to be read to include one or at least one, and the singular forms also include the plural unless explicitly stated otherwise or it is clear from the context that the contrary is meant. The terms "comprise", "comprising", "include", "including", "have", "having", or any other variation thereof used herein are intended to cover non-exclusive inclusion. For example, a process, method, article, or apparatus that comprises a list of elements is not necessarily limited to only those elements, and may include other elements not expressly listed or inherent to such process, method, article, or apparatus. Further, unless explicitly stated to the contrary, "or" refers to an inclusive logical disjunction rather than an exclusive logical disjunction. For example, the condition A or B is satisfied by any one of the following: A is true (or present) and B is false (or absent), A is false (or absent) and B is true (or present), or both A and B are true (or present).
[0078] The terms "substantially", "substantially the same", "substantially identical", "substantially similar", "generally", and "about" as used herein are used to describe and account for minor differences. When used with an event or situation, these terms may refer to not only when the event or situation occurs exactly, but also when the event or situation occurs approximately. For example, when used in connection with a numerical value, these terms may mean a variation range of ±10% or less of that numerical value, such as ±5% or less, ±4% or less, ±3% or less, ±2% or less, ±1% or less, ±0.5% or less, ±0.1% or less, or ±0.05% or less. For example, two numerical values can be considered "substantially" the same if the difference between them is ±10% or less of the average of the numerical values, such as ±5% or less, ±4% or less, ±3% or less, ±2% or less, ±1% or less, ±0.5% or less, ±0.1% or less, or ±0.05% or less.
[0079] In addition, amounts, ratios, and other numerical values may be presented in a range format in this specification. It should be understood that such range formats are used for convenience and brevity and that they include not only the numerical values explicitly specified as the limits of a range but also, as if each numerical value and subrange were explicitly specified, all individual numerical values or subranges included within that range.
[0080] Regarding the technology disclosed in this specification, although it has mainly been described using specific operations performed in a specific order, it should be understood that these operations may be combined, divided into parts, or the order may be changed to form equivalent technologies without departing from the teachings of this disclosure. Therefore, unless otherwise specifically indicated in this specification, the order and grouping of operations are not limitations of this disclosure.
Claims
1. A computer implementation method that uses machine learning to identify glass breakage events, Accessing or obtaining a machine learning model trained using training audio data representing (i) training ambient sounds, (ii) training glass sounds, and (iii) training glass breaking sounds, by one or more processors, The acquisition of audio data over a period of interest by one or more processors, One or more processors that process the audio data using the machine learning model determine whether a glass breakage event occurred during the period of interest. When it is determined that the aforementioned glass breakage event has occurred, one or more processors shall indicate that the glass breakage event has occurred. Computer implementation methods, including those mentioned above.
2. The training audio data and the audio data each include one or more spectrograms. Acquiring the audio data over the aforementioned period of interest includes generating the spectrogram of the audio data from the raw audio data using a Fourier transform. The computer implementation method according to claim 1.
3. (i) The training ambient noise includes sounds caused by the operation of the machine, (ii) The training glass sound includes a sound produced by the first glass surface in contact with either the second glass surface or the non-glass surface, (iii) The training glass breakage sound includes a sound caused by either a glass crack or a glass break, The computer implementation method according to claim 1, wherein the machine performs a biomanufacturing process, and the training glass sound and the training glass break sound are generated by one or more containers of one or more pharmaceuticals.
4. The computer implementation method according to claim 3, further comprising, after identifying the glass breakage event, causing one or more processors to automatically stop the operation of the machine.
5. The computer implementation method according to claim 1, further comprising: using one or more processors iteratively adjusting the gain of an amplification applied to one or both of the training audio signal and / or the audio signal until a performance threshold is met, wherein the training audio signal and / or the audio signal correspond to the training audio data and / or the audio data, respectively.
6. The computer implementation method according to claim 1, wherein the machine learning model is a convolutional neural network.
7. The computer implementation method according to claim 1, wherein the machine learning model is trained using supervised learning techniques, and the audio data includes (i) ambient sound labels, (ii) glass sound labels, and (iii) breakage sound labels.
8. A computer system that uses machine learning to identify glass breakage events, One or more processors, A program memory connected to one or more processors, which stores executable instructions that, when executed by the one or more processors, cause the computer system to perform the method described in any one of claims 1 to 7, Computer system.
9. A computer implementation method for training a machine learning model to identify glass breakage events, Acquiring training audio data using one or more processors, The classification of the training audio data by one or more processors into a plurality of subsets, each corresponding to different actual result data, wherein the subsets include (i) at least one subset representing training ambient sounds, (ii) at least one subset representing training glass sounds, and (iii) at least one subset representing training glass breaking sounds. The process includes generating the machine learning model for identifying glass breakage events using the classified subset of the training audio data by one or more processors, Computer implementation method.
10. The computer implementation method according to claim 9, wherein the training audio data includes one or more spectrograms.
11. (i) The training ambient noise includes sounds caused by the operation of the machine, (ii) The training glass sound includes a sound produced by the first glass surface in contact with either the second glass surface or the non-glass surface, (iii) The training glass breakage sound includes sounds caused by either glass cracking or glass breakage, The computer implementation method according to claim 9.
12. The computer implementation method according to claim 11, wherein the machine performs a biomanufacturing process, and the training glass sound and the training glass break sound are generated by one or more containers of one or more pharmaceuticals.
13. The computer implementation method according to claim 9, further comprising: one or more processors iteratively adjusting the gain of amplification applied to a training audio signal until a performance threshold is met, wherein the training audio signal corresponds to the training audio data.
14. The computer implementation method according to claim 9, wherein the machine learning model is a convolutional neural network.
15. One or more tangible, non-transient, computer-readable media for storing executable instructions for training a machine learning model to identify glass breakage events, which, when executed by one or more processors of a computer system, stores executable instructions causing the computer system to perform the method according to any one of claims 9 to 14. One or more tangible, non-transient, computer-readable media.