Sound and / or vibration predictive maintenance

A sensor-based predictive maintenance system using AI analysis for sound and vibration monitoring addresses the challenge of undetected component failures in apparatuses, enhancing detection and reducing repair costs and downtime.

WO2025233124A1PCT designated stage Publication Date: 2025-11-13KONINKLIJKE PHILIPS NV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/061070
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-10
Filing Date
2025-04-23
Publication Date
2025-11-13

AI Technical Summary

Technical Problem

Existing apparatuses, such as medical imaging scanners, lack effective predictive maintenance methods to detect component failures through sound and vibration analysis, which are often unnoticed by personnel and can lead to costly repairs and downtime.

Method used

Implementing a system with sensors to monitor sound and vibration patterns, utilizing artificial intelligence for analysis, to predict component health and notify of abnormalities, enabling continuous automated monitoring.

Benefits of technology

The system effectively detects early signs of component failure, reducing repair costs and downtime by providing timely notifications and reports on apparatus health.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000024_0000
    Figure 00000024_0000
  • Figure 00000024_0001
    Figure 00000024_0001
  • Figure 00000024_0002
    Figure 00000024_0002
Patent Text Reader

Abstract

An apparatus (102) includes at least one component (108) that produces a characteristic sound during normal, healthy functioning and a different sound outside of the normal, healthy functioning, at least one sensor (110) disposed in spatial proximity to the at least one component to sense the characteristic sound and configured to sense the characteristic sound and generate a signal indicative of the characteristic sound, a memory (112) storing instructions (114) for a trained artificial intelligence-based analyzer module (1202) trained to identify the characteristic sound in the signal using audio analysis, and a controller (118) with a processor configured to execute the instructions to process signals from the at least one sensor (110) to predict a health state of the apparatus based on whether the characteristic sound is identified. The controller determines the health state is not normal, healthy functioning and invokes transmission of a notification in response to not identifying the characteristic sound in the signal.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] SOUND AND / OR VIBRATION PREDICTIVE MAINTENANCE

[0002] TECHNICAL FIELD

[0003] The following generally relates to predictive maintenance of an apparatus, and, more particularly, to sound and / or vibration based predictive maintenance, including sound and / or vibration based predictive maintenance of a healthcare apparatus, such as a medical imaging scanner, a personal care apparatus, and / or other apparatus.

[0004] BACKGROUND

[0005] Electrical and / or mechanical based apparatuses (e.g., a medical imaging scanner, an electric toothbrush, etc.) include electrical and / or mechanical components that can fail. The failure of individual components in some of these apparatuses often is not directly noticeable, even where the failure leads to results other than the results expected from the apparatus, e.g., degraded image quality with a medical imaging scanner. Some components of these apparatuses may produce sounds and / or vibrations during operation, including characteristic sounds and / or vibrations during normal, healthy operation. However, when individual components begin to function abnormally or malfunction, the sound and / or vibrational behavior of certain components is expected to change from the characteristic sounds and / or vibrations of normal, healthy operation to other sounds and / or vibrations.

[0006] With a medical imaging scanner, personnel, e.g., a technologist performing an imaging procedure, a radiologist, etc., generally are not trained to detect such failures. Moreover, job function and / or work shift rotation may prevent personnel from observing a medical imaging scanner with the requisite detail to detect such failures. Moreover, the source of a problem, e.g., which component changed its functioning, is usually unknown until service personnel have investigated and possibly disassembled the apparatus. Often, impactful failures (e.g., dangerous and / or costly) are preceded by smaller parts and functions, e.g., bearings or pumps, beginning to fail, which lead to subsequent more significant failure that can increase repair cost, system damage, and / or system down time. Unfortunately, sounds and / or vibrations of such apparatuses are not monitored with sufficient detail for useful predictive maintenance.

[0007] As such, there is an unresolved need for an improved sound and / or vibration predictive maintenance approach(s).

[0008] SUMMARY

[0009] Aspects described herein address the above-referenced problems and / or others. The following describes a sound and / or vibration predictive maintenance approach(s). In one instance, this approach(s) allows for continual automated monitoring of a health of an apparatus during regular operation. As such, this approach(s) may mitigate at least a shortcoming of existing maintenance approaches. For example, the approach(s) described herein can monitor sounds and / or vibrations with sufficient detail for successful predictive maintenance, including the beginning of a failure, which can mitigate subsequent more significant failure and, hence, increased repair cost, system damage, and / or system down time.

[0010] In one aspect, an apparatus includes at least one component that produces a characteristic sound during normal, healthy functioning and a different sound outside of the normal, healthy functioning. The apparatus further includes at least one sensor disposed in spatial proximity to the at least one component to sense the characteristic sound and configured to sense the characteristic sound and generate a signal indicative of the characteristic sound. The apparatus further includes a memory storing instructions for a trained artificial intelligence-based analyzer module trained to identify the characteristic sound in the signal using audio analysis. The apparatus further includes a controller with a processor configured to execute the instructions to process signals from the at least one sensor to predict a health state of the apparatus based on whether the characteristic sound is identified in the signal. The controller is configured to determine the health state is not normal, healthy functioning and at least invokes transmission of a notification in not identifying the characteristic sound in the signal.

[0011] In another aspect, a computer-implemented method includes training an analyzer module to identify a characteristic sound produced by a component, a sub-system, or a system of an apparatus during normal, healthy functioning, where the component, the sub-system, or the system produces a different sound outside of the normal, healthy functioning. The computer- implemented method further includes sensing, with a sensor, sound produced by the component, the sub-system, or the system during operation of the apparatus. The computer-implemented method further includes generating, with the sensor, a signal indicative of the sensed sound. The computer-implemented method further includes performing audio analysis, with the trained analyzer module, on the signal to determine whether the signal includes the characteristic sound. The computer-implemented method further includes transmitting a notification in response to a result of the audio analysis indicating absence of the characteristic sound in the signal.

[0012] In another aspect, a computer readable medium is encoded with computer executable instructions that cause a processor to: train an analyzer module to identify a characteristic sound produced by a component, a sub-system, or a system of an apparatus during normal, healthy functioning, where the component, the sub-system, or the system produces a different sound outside of the normal, healthy functioning, receive a signal from a sensor of an apparatus, wherein the apparatus includes a component, a sub-system, or a system that produces a characteristic sound during operation of the apparatus, and the signal sensor is configured to sense sounds of the component, the sub-system, or the system, including the characteristic sound, perform audio analysis, with the trained analyzer module, on the signal to determine whether the signal includes the characteristic sound, and transmit a notification in response to a result of the audio analysis indicating absence of the characteristic sound in the signal.

[0013] Those skilled in the art will recognize still other aspects of the present application upon reading and understanding the attached description.

[0014] BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The invention may take form in various components and arrangements of components, and in various steps and arrangements of steps. The drawings are only for the purpose of illustrating the embodiments and are not to be construed as limiting the invention.

[0016] FIG. 1 diagrammatically illustrates an example system including an apparatus with one or more systems, one or more sensors, a controller, and memory with instructions and data, in accordance with one or more embodiments herein.

[0017] FIG. 2 diagrammatically illustrates an example system of the apparatus including one or more sub-systems, in accordance with one or more embodiments herein.

[0018] FIG. 3 diagrammatically illustrates an example sub-system system of the system of the apparatus including one or more components, in accordance with one or more embodiments herein.

[0019] FIG. 4 diagrammatically illustrates an example configuration of the components and the sensors, in accordance with one or more embodiments herein.

[0020] FIG. 5 diagrammatically illustrates another example configuration of the components and the sensors, in accordance with one or more embodiments herein.

[0021] FIG. 6 diagrammatically illustrates yet another example configuration of the components and the sensors, in accordance with one or more embodiments herein.

[0022] FIG. 7 diagrammatically illustrates still another example configuration of the components and the sensors, in accordance with one or more embodiments herein.

[0023] FIG. 8 diagrammatically illustrates a variation in which one or more sensors are inside a system of the apparatus, in accordance with one or more embodiments herein. FIG. 9 diagrammatically illustrates a variation in which one or more sensors are inside a sub-system of the system of the apparatus, in accordance with one or more embodiments herein.

[0024] FIG. 10 diagrammatically illustrates a variation in which one or more sensors are on the outside of the apparatus, in accordance with one or more embodiments herein.

[0025] FIG. 11 diagrammatically illustrates a variation in which one or more sensors are located remote from the apparatus, in accordance with one or more embodiments herein.

[0026] FIG. 12 illustrates an example analyzer module, in accordance with an embodiment(s) herein.

[0027] FIG. 13 illustrates a variation of the analyzer module, in accordance with an embodiment(s) herein.

[0028] FIG. 14 illustrates another variation of the analyzer module, in accordance with an embodiment s) herein.

[0029] FIG. 15 illustrates an example method including training the analyzer module using supervised training in-factory and on site, in accordance with an embodiment s) herein.

[0030] FIG. 16 illustrates an example method including training the analyzer module using supervised training on site, in accordance with an embodiment s) herein.

[0031] FIG. 17 illustrates an example method employing the analyzer module trained with the supervised training, in accordance with an embodiment s) herein.

[0032] FIG. 18 illustrates an example method including training the analyzer module using unsupervised training in-factory and on site, in accordance with an embodiment s) herein.

[0033] FIG. 19 illustrates an example method including training the analyzer module using unsupervised training on site, in accordance with an embodiment s) herein.

[0034] FIG. 20 illustrates an example method employing the analyzer module trained with the unsupervised training, in accordance with an embodiment s) herein.

[0035] DESCRIPTION OF EMBODIMENTS

[0036] The following describes a sound and / or vibration predictive maintenance approach(s). This approach employs sound and / or vibration sensors to sense sound and / or vibration in connection with components that produce characteristic sound and / or vibration during normal, healthy operation and other sound and / or vibration otherwise. In one instance, artificial intelligence (Al) and / or models designed and trained for sound and / or vibrational analysis are employed to process signals from the sensors and predict abnormal behavior. The predictive maintenance can run under human control and / or autonomously without human interaction, producing notification when relevant, and, optionally, reports on component s) and / or apparatus health, including, in some instances, sound / vibration landscape information.

[0037] FIG. 1 diagrammatically illustrates an example system 100. The system 100 includes an apparatus 102. Examples of suitable apparatuses include apparatuses with one or more elements (e.g., electrical and / or mechanical) that produce characteristic sounds and / or vibrations during normal, healthy operation, and other sounds and / or vibrations outside of normal, healthy operation. Examples of such apparatuses include medical imaging scanners such as Magnetic Resonance (MR), Computer Tomography (CT), Positron Emission Tomography (PET), Single Photon Emission Computed Tomography (SPECT), X-ray, Ultrasound (US), etc., personal care / hygienic devices such as electric toothbrushes, shavers, etc., and / or other apparatuses that include such elements.

[0038] The apparatus 102 includes a system 104i, ..., 104i, ..., and a system 104N (where N is an integer greater to or equal to one, and 1 < I < N), collectively referred to herein as systems or at least one system 104. One or more of the systems 104 includes a set of subsystems, and one or more of the sub-systems includes a set of components. For example, FIG. 2 diagrammatically illustrates an embodiment in which at least the system 104i includes a subsystem 106i, . . ., a sub-system 106i, . . . and 106K (where K is an integer greater to or equal to one, and 1 < I < K), collectively referred to herein as sub-systems or at least one sub-system 106, and FIG. 3 diagrammatically illustrates an embodiment in which at least the sub-system 106i includes a component 108i, . . ., a component 108i, . . . and a component 108M (where M is an integer greater to or equal to one, and 1 < I < M), collectively referred to herein as components or at least one component 108.

[0039] By way of non-limiting example, an MR scanner apparatus generally includes the following systems: a gantry (housing the acquisition portion), an operator console (e.g., with a processor, memory, application software, etc.), and a reconstructor (e.g., with a graphics and / or other processor). The gantry system may include the following sub-systems: a main magnet, a gradient coil, a gradient amplifier, an RF amplifier, RF electronics, a data acquisition system, a cooling system, etc. A sub-system such as the cooling system includes cold head that produces characteristic sound (e.g., “chirping”) during normal, healthy operation as the components therein expand and contract as helium gas enters and is compressed. Examples of other components in apparatuses in general that may produce characteristic sound and / or vibration include control electronics, a motor, a gear, a belt, a drive shaft, a pump, a compressor, an MRI bore component with characteristic vibrations in response to Lorentz forces, a moving part inside of a rotating CT gantry, etc. The apparatus 102 further includes a sensor 110i, . . ., a sensor 110i, . . . and a sensor 110L (where L is an integer greater to or equal to one, and 1 < I < L), collectively referred to herein as sensors or at least one sensor 110. The sensors 110 include acoustic sensors such as one or more sound sensors and / or one or more vibration sensors. In general, a sound sensor measures sound pressure / electromagnetic radiation in the audible range of the electromagnetic spectrum - 20 Hertz (Hz) to 20 kHz. A suitable sound sensor can be dedicated to sensing sound of one or more components or also further configured for another purpose, e.g., communication between a patient in an examination room and a technologist at the console. A non-limiting example of a sound sensor is a microphone. Vibration sensors include displacement sensors, velocity sensors, and / or acceleration (Piezoelectric, micro-electromechanical systems (MEMS), etc.) sensors. In general, these sensors can measure vibration with a frequency from a few Hz to a few thousand Hz.

[0040] For clarity and brevity, the following describes the sensors 110 in connection with sensing sound. However, one of ordinary skill in the art would readily understand differences in placement, etc. For example, whereas a sensor configured to sense sound may be placed such that it is not in physical contact with a component, a corresponding vibration sensor may have to be in physical contact with the component, sub-system, system, and / or apparatus to sense vibration of the component.

[0041] FIGS. 4-7 diagrammatically illustrates non -limiting configurations of the sensors 110 in connection with the components 108. In FIG. 4, the sensor 110i is located within proximity to sense a sound wave 402 produced by the component 108i and is configured to sense the sound wave 402. Such configuration may include low, band and / or high pass filtering as needed to sense a particular sound, signal conditioning, signal processing, etc. In FIG. 5, more than one of the sensors 110 (i.e., the sensor 110i, ... the sensor 110i) are configured to sense the same sound wave 402 produced by the component 108i. Such redundancy may be employed to confirm a detected component failure, identify a sensor that may not be functioning properly, provide backup in case a sensor fails, etc.

[0042] The embodiment disclosed in FIG. 6 is substantially similar to the embodiment disclosed in FIG. 5 except that the sensor 110i is configured to sense a sound wave 602 produced by the component 1081 and the sensor 110i is configured to sense a different sound wave 604 produced by the component 108i. In FIG. 7, the sensor 110i is within proximity to sense a sound wave 702 produced by the component 108i, . . ., and a sound wave 704 produced by a component 108i. In general, the configuration of one or more of the components 108 and one or more of the sensors 110 can include one or more of one-to-one, one-to-many, many-to-many, and many -to- one, where one or more of the components 108 can be configured to sense one or more sound waves from each one or more of the sensors 110. Another embodiment includes a combination of FIGS. 4-7.

[0043] FIGS. 8 and 9 diagrammatically illustrates variations of the placement of the sensors 110, including one or more of the sensors 110 inside of one or more of the systems 104, and, additionally, or alternatively, one or more of the sensors 110 inside of one or more of the sub-systems 106. Initially referring to FIG. 8, the system 104i includes one or more of the subsystems (i.e., the sub-system 106i, ..., the sub-system 106i, ... the sub-system 106K) and one or more of the sensors 110 (i.e., the sensor 110i, . . ., the sensor 110i, . . . and the sensor 110L). Next, in FIG. 9, the sub-system 106i includes one or more of the components 108 (i.e., the component 108i, . . ., the component 108i, . . . and the component 108M) and one or more of the sensors 110 (i.e., the sensor 110i, . . ., the sensor 110i, . . . and the sensor 110L). Another embodiment includes a combination of FIGS. 1, 8 and / or 9.

[0044] In FIG 1, the sensors 110 are shown inside of the apparatus 102. In another instance, at least one of the sensors 110 can be disposed otherwise, with the at least one of the sensors 110 located within a proximity where it can sense a sound and / or vibration of a component it is intended to monitor. For example, FIG. 10 diagrammatically illustrates an embodiment in which at least the sensor 110i is disposed outside of the apparatus 102, e.g., on an outside surface of the apparatus 102. In another example, FIG. 11 diagrammatically illustrates an embodiment in which at least the sensor 110i is located remote from the apparatus 102 and supported by a sensor support 1102 (a portable stand, a wall bracket, etc.). Another example combines FIGS. 1, 8 and 9.

[0045] Returning to FIG. 1, the apparatus 102 further includes a computer readable storage medium (“memory”) 112, which includes non-transitory medium (e.g., a storage cell, device, etc.) and excludes transitory medium (i.e., signals, carrier waves, and the like). The computer readable storage medium 112 is encoded with computer instructions (“instructions”) 114 and configured to store data 116. The apparatus 102 further includes a controller 118 with a processor such as a micro-processing unit (MPU), a central processing unit (CPU), a graphics processing unit (GPU), etc. The controller 118 is configured to execute the computer instructions 114 and read / write the data 116.

[0046] The computer instructions 114 include instructions for processing signals from the sensors 110 to determine a health state of the apparatus 102, e.g., by determining a health of one or more of at least one of the components 108, at least one of the systems 104, and / or at least one of the sub-systems 106. As described in greater detail below, in one instance the computer instructions 114 include artificial intelligence (Al) based instructions trained to analyze sensed information (i.e., sound and / or vibration) by one or more of the sensors 110 during operation and determine whether the apparatus 102 is functioning in a normal, healthy state or malfunctioning. The Al is trained at least to learn the sounds and / or vibrations of the apparatus 102 (i.e., of one or more particular components) during regular operation, and / or distinguish sound and / or vibration from persistent and / or transient background noise, including human voice.

[0047] In one instance, where it is determined that a component is not functioning in a normal, healthy state, the controller 118 invokes transmission of a notification indicating a component may not be functioning properly / is likely malfunctioning. The notification can also include a time stamp of when the component was determined to be malfunctioning and / or other information such as information about the signal and / or the signal itself, an expected signal, a difference therebetween, and / or decision criteria used to evaluate the difference, etc. In another instance, the notification may include an identification of the component, etc. Other information includes an identification of a user(s) of the apparatus 102, the type of examination, parameters used for the examination, etc.

[0048] In one instance, the controller 118 estimates a location of the component in the apparatus based on the sensor(s) sensing the signal and a spatial mapping of the sensors 110 relative to the components 108 in the data 116. Additionally, or alternatively, the controller 118 identifies an issue with the component based on the signal and a mapping between sounds and issues. In one instance, the controller 118 additionally retrieves a description of the issue and / or a possible solution to the issue based on the signal and descriptions and / or solutions to issues. In these instances, the data 116 may include multiple different sound snippets for healthy operation of a same component, but based on different uses of the apparatus 102. The estimated location, the identified issue, and / or the retrieved description and / or the possible solution can be provided with the notification.

[0049] A notification can be transmitted to a system and / or sub-subsystem of the apparatus 102, and / or to a computing device remote from the apparatus 102 such as a workstation, a cloud service, a smartphone, etc. where the notification can be displayed in human readable form and / or stored. A user with suitable authorization (e.g., personnel of the manufacturer, etc.) to read the notification can read the notification, e.g., based on their level of authorization, which can be from full access, to varying degrees of partial access, to no access. The notification, including any information therein and / or therewith, can be archived and / or otherwise stored in a database, log files, in the data 116, etc. FIG. 12 diagrammatically illustrates a non -limiting example of the instructions 114. In this example, the instructions 114 include instructions for an analyzer module 1202. The analyzer module 1202 receives, as input, one or more signals from the sensors 110. The analyzer module 1202 is configured to process the one or more signals to predict a health state of the apparatus 102, the systems 104, the sub-systems 106, and / or the components 108. In one instance, the analyzer module 1202 is configured to process the one or more signals using artificial intelligence (Al) to detect outliers / changes in the sound landscape and log information about the timing of events, the audio changes, and / or an estimate of where the changes originated, e.g., based on the positions of the sensors 110. The Al can be trained via unsupervised and / or supervised training.

[0050] Examples of suitable Al include, but are not limited to, one or more of: a Neural Network (NN) such as a Recurrent Neural Network (RNN) (e.g., Long Short Term Memory (LSTM), etc.), a Variational Autoencoder (VAE), a transformer (e.g., encoders / decoders, foundation models, large language models (LLMs),etc.) with an attention mechanism (e.g., selfattention, multi-head attention, etc.), a sliding window and other temporal window analysis approach, and / or other approach(s). In general, with a traditional NN the inputs and the outputs are independent of each other. With a RNN, an output from a previous step is fed back as an input to a current step in one or more hidden layers, and a hidden state stores the previous input to the network. The LSTM reads and writes information important in predicting the output, with less focus on the information that is not important in predicting the output.

[0051] In general, a transformer is a deep learning model that uses attention to process data. With one transformer, the input is tokenized into individual tokens, which are encoded via an embedding layer. A positional encoding vector is added to each embedding, and the embeddings then go through a multi-head self-attention layer of an encoder. An add and normalize step performs a layer normalization and adds the original embeddings via a skip connection. The added and normalized embeddings are processed by a multilayer perceptron consisting of multiple fully connected layers with a nonlinear activation function in between. Another add and normalize step performs a layer normalization and adds the processed embeddings via a skip connection. The output embeddings are then fed to a decoder, which has a processing architecture that is similar to the encoder, except it includes a masked multi-head self-attention layer and processes the output of the encoder. In this configuration, the encoder extracts relevant information from the input and the decoder makes a prediction based thereon.

[0052] With one multimodal transformer model, e.g., an X second sound snippet can be sensed and stored. The transformer can take a first sub-set (i.e., all or less than X seconds) of the snippet and predict how a second snippet, a third snippet, . . . will look N seconds later. This involves embedding the sound snippets with an appropriately trained embedding model to tokenize the sound input. The predicted snippet can be compared with a sensed snippet to determine any deviation of the prediction from the sense snippet. The deviation may be readily noticeable because the transformer has metrics of the hidden states so any sound snippet would be transformed to a latent space vector where encoded vectors are compared with a scalar product to determine similarity. The deviation can be logged, output as a graph, etc. A user and / or the analyzer module 1202 can determine whether the deviation is within an acceptable tolerance or outside of the tolerance and indicative of background noise. A user and / or the analyzer module 1202 can then determine whether the apparatus 102 is operating in a normal, healthy state or possibly malfunctioning, based on criteria such as a predetermined threshold.

[0053] In one instance, if the scalar product is too small (e.g., below a learned threshold, considering natural, “healthy” variance), the analyzer module 1202 would subtract the predicted feature vector from the feature vector of the measurement, and the result will contain a latent space representation of the unexpected deviation. In one instance, the decoder decodes the difference feature vector and analyzes the result. In another instance, a trained classifier is directly used on the difference feature vector.

[0054] For either approach, the analyzer module 1202 can be trained using supervised learning with examples of “acceptable” noise, which is expected to show some consistent characteristics in the high-dimensional feature space. In general, a component failure and / or unacceptable functioning would produce difference vectors in feature space that are distinct from noise. If the scalar product threshold triggers the difference vector analysis, an outlier detection can be used to discriminate problematic operation signals from the “acceptable noise” distribution, which was learned during classifier training.

[0055] With another transformer model, an autoregressive, generative model is employed to learn sound continuations from a given operation sound (e.g., once an apparatus has started running, predict how the sound should continue). Transformers with attention mechanisms are utilized on the frequency spectra (or some transformation thereof), self-attention is utilized for sound prediction, and cross attention is utilized to determine similarity. During operation, similarity between the predicted and the recorded sounds are monitored. If deviations failing a predetermined threshold arise, the deviations can be reported, e.g., via a notification and / or a report.

[0056] Training can be achieved by recording healthy apparatus operations and training the autoregressive model therewith. Unhealthy apparatus operation sounds can also be included with a corresponding label, including different failure and deterioration modes, if available. To adapt a generative sound model from factory training to the in-house sound landscape of a customer, the apparatus can be calibrated at the customer site to adapt the generative model to the “dialect” of the customer’s apparatus, e.g., via fine-tuning, i.e., adjusting a subset of transformer model parameters with an appended training to the pre-trained model from factory. During the healthy state, calibration scans can be run to adjust the parameters of the underlying neural networks to produce the recorded sounds, through back-propagation on the pre-configured network from factory training.

[0057] In general, a foundation model and a LLM are forms of generative Al that are trained (e.g., self-supervised learning, semi-supervised learning, etc.) on large amounts of data and are adaptable to a wide range of tasks such as predicting, classifying, etc. In general, a sliding window analysis allows for the construction of models that can automatically learn and improve over time. They can be used to develop models that are able to automatically learn and improve from data, without the need for manual intervention. One algorithm works by first dividing the data into a series of smaller windows, and then training a model on each window. The model is then used to predict the output for the next window, and the process is repeated. As the model is trained on more data, it becomes more accurate at predicting the output for future windows.

[0058] In another instance, the analyzer module 1202 is configured to process the signal(s) based on Mel Spectrograms (e.g., visualization of sound on the Mel scale, which is a logarithmic transformation of a signal’s frequency) and Mel Frequency Cepstral Coefficients (MFCCs). In one instance, this is achieved by converting from frequencies in Hertz to the Mel scale, taking the logarithm of Mel representation of audio, taking logarithmic magnitude and using Discrete Cosine Transformation (DCT), aggregating a spectrum over Mel frequencies as opposed to time (MFCCs), and applying an image analysis neural network framework (e.g., visual attention, and convolutional neural networks) to extract latent features of the analyzed sound and perform classification. In general, any approach that analyzes the spectrograms are contemplated herein, e.g., where sound is transformed to frequency components to determine which frequency components are present in one sample that are absent from another sample.

[0059] In one instance, the FCC includes a two dimensional (2-D) matrix where time over short-time frequency components are represented in short time frames of that time. The 2- D matrix can be linearized into a vector, which is a representation of the corresponding sound snippet. The vector can be fed through a neural network or a transformer with attention mechanisms where certain different frequency components in different parts of the spectrum and / or in different times would be re-occurring and disappearing over time. In one instance, an X-axis of the 2-D matrix represents a timeframe and a Y axis of the 2-D matrix represents frequencies within a short time snippet. With a retention mechanism, the neural network or the transformer connects to different parts of the 2-D matrix to produce new snippets, new spectrograms, to continue the sound.

[0060] In one instance, the sensors 110 are installed with the apparatus 102 in-factory. In this instance, the analyzer module 1202 can be trained at the factory and then installed at a customer site. The analyzer module 1202 can then be further trained at the customer site to update the analyzer module 1202 for the customer site. Alternatively, the analyzer module 1202 is trained at the factory, installed at a customer site, and then further trained at the customer site to update the analyzer module 1202 for the customer site. In another instance, the sensors 110 are installed at the customer site and the analyzer module 1202 is trained at a customer site. It is to be appreciated that an installed sensor can later be re-located, e.g., after training where the analyzer module 1202 is re-trained based on the new location.

[0061] In the illustrated embodiment, the instructions 114 and hence the analyzer module 1202 is part of the apparatus 102. In another instance, the analyzer module 1202 is located remote from the apparatus 102, e.g., a “cloud” based service, part of another computing system (e.g., local and / or remote), etc. The analyzer module 1202 can operate with or without human surveillance, and produce notifications and / or reports for personnel (e.g., a user, staff, technical support, and / or vendor) when a relevant event has been detected. In one instance, the analyzer module 1202 is configured to detect abnormal functioning and produce notifications. In another instance, the analyzer module 1202 is further configured to describe the problem in a notification.

[0062] FIG. 13 diagrammatically illustrates a variation of the instructions 114 discussed in connection with FIG. 12. In this variation, the instructions 114 further include a report generating module 1302. The report generating module 1302 is configured to generate a report with one or more of the notifications, the estimated location of the component, the identified issue of the component, the retrieved description of and / or the solution to the issue, and / or other information. The report generating module 1302 can be part of a system and / or sub-system of the apparatus 102, a report generating system external to the apparatus 102, etc.

[0063] In one instance, to determine a description of the failure, relevant features are extracted from the signal, e.g., via MFCCs spectrum and / or a one or multidimensional timeseries analysis. The analyzer module 1202 passes the extracted features as an embedding to a LLM trained on a vast base of technical / service information to map the input to a particular malfunction. A description of the malfunction is obtained from the mapping. In this instance, the report includes the particular malfunction along with the description thereof.

[0064] In another instance, the report can be generated as described in US 63 / 604,918, filed on 1 December 2023, and entitled “maintenance service system for medical imaging systems,” which is incorporated herein by reference in its entirety. As described in 63 / 604,918, the maintenance service system generates a maintenance report with information about required service actions that are tailored to a specific user and / or user group, such as an end user, e.g., by tailoring a report based on a profile of a specific individual and / or experience for self-help in medical equipment fixing / repair.

[0065] In one instance, for describing a problem and generating a report, the analyzer module 1202 is trained to learn healthy and faulty apparatuses for different faults and with labels for problems. In this manner, for example, the analyzer module 1202 learns a particular sound and / or vibration corresponds to a loose screw, another sound and / or vibration corresponds to a bearing break, etc. With this approach, the sound spectrum is processed with a classifier that classifies the sound spectrum as a particular problem. In another instance, the analyzer module 1202 may provide a list of possible problems, each with probability, likelihood, confidence internal, etc. The data 116 and / or other storage accessible to the analyzer module 1202 can store descriptions for each of the problems, which are used for describing a problem.

[0066] Additionally, or alternatively, the data 116 and / or other storage accessible to the analyzer module 1202 can store information that maps sounds to normal and / or abnormal functioning. For normal sounds, the information can be further mapped to different available uses of the apparatus 102. For example, the power requirements may be different for two different uses, and the sound and / or vibration of a gradient amplifier may be different depending on the power such that in one use the sound and / or vibration from the gradient amplifier is as expected and represents normal, healthy functioning, where that same sound and / or vibration from the same gradient amplifier in a different use either is not expected or indicates an unhealthy apparatus. By way of another example, with an MR scanner, the sound and / or vibration will be different between an inversion recovery sequence and diffusion imaging sequence.

[0067] FIG. 14 diagrammatically illustrates another variation of the instructions 114. In this variation, the instructions 114 further include a de-identification module 1402. In one instance, the de-identification module 1402 is configured to evaluate the signal and determine whether the signal includes sound other than the desired sound of the systems 104, sub-systems 106 and / or components 108, e.g., human voice, machinery other than the apparatus 102, etc. In response to determining there is such sound in the signal, the de-identification module 1402 removes the sound, e.g., via filtering and / or otherwise. In one instance, a model is trained by overlaying different vocal recordings onto “clean” recorded machine sounds, e.g., acquired as discussed herein. The combined track is the input to the model and the separate, original tracks are the output. The splitter model can be prepended with a classifier that recognizes whether human speech is present in the recording. If human speech is determined to be present, the splitter can be triggered, otherwise the recording is just processed as described herein. Known or other techniques can be employed. For example, known Al models such as those used to split a music track into an instrumental track and a vocal track using specifically trained neural network based models can be employed to spit human sounds from recorded machine sounds. Another variation includes a combination of the variations of FIGS. 13, 14, and / or other variations.

[0068] Additionally, or alternatively, known and / or other noise cancellation and / or suppression approaches can be employed. In one instance, the sensors 110 are employed to sense sounds while the apparatus 102 is “off’ and the analyzer module 1202 analyzes the sounds to determine a background sound landscape. Different background sound landscapes can be determined for different time periods throughout a day to capture time dependent background noise such as a nearby entity with machines that produce sound sensed by the sensor 110 where the entity operates the machines only one shift each day. In one instance, background sound can be analyzed before, as part of and / or after a calibration of the analyzer module 1202 to determine if there is any background sound not in the background sound landscape that may impact the calibration. This can be performed in-factory and / or at a customer site, and can be performed more than once, e.g., one or more times in the future to update the background sound landscape, if needed. Additionally, or alternatively, reference sensors placed farther away from the machine to monitor can be used to determine background sounds through comparison of the signals from the main monitoring sensors with the reference sensor signals. Background sounds are expected to be found in the signals from both sensor types, but should be more prominent in the reference sensor signal.

[0069] As briefly discussed above, the analyzer module 1202 may include Al trained under unsupervised and / or supervised training. The following describes such training and use of the trained Al for predictive maintenance.

[0070] FIG. 15 discloses a computer-implemented method for training the analyzer module 1202 using unsupervised learning in-factory and on site. It is to be appreciated that the ordering of the acts of the method is not limiting. As such, other orderings are contemplated herein. In addition, one or more acts may be omitted, and / or one or more additional acts may be included.

[0071] At 1502, the apparatus 102 is assembled in-factory, including installation of the sensors 110. At 1504, the apparatus 102 is confirmed to be a healthy functioning apparatus. At 1506, the apparatus 102 is operated in a defined use of operation. At 1508, the sensors 110 sense characteristic sounds of the components 108 during the normal, healthy operation. At 1510, the analyzer module 1202 is trained to predict a temporal continuation of a sound snippet representing healthy machine function, as recorded by the sensors 110. At 1512, the apparatus 102 is installed at a customer site and confirmed to be a healthy functioning apparatus.

[0072] At 1514, the apparatus 102 is operated in one or more different operational modes. At 1516, the sensors 110 sense characteristic sounds of the components 108 during normal, healthy operation. At 1518, the analyzer module 1202 is updated, calibrated, fine-tuned, etc. based on the sounds sensed at the customer site. In one instance, this includes adjusting the model to the local sound landscape, echo characteristics, background sounds, etc. This can be achieved with predefined calibration runs, iterating through different operational modes of the machine, and / or during regular usage where the apparatus is functioning correctly.

[0073] FIG. 16 discloses a computer-implemented method for training the analyzer module 1202 using unsupervised learning on site. It is to be appreciated that the ordering of the acts of the method is not limiting. As such, other orderings are contemplated herein. In addition, one or more acts may be omitted, and / or one or more additional acts may be included.

[0074] At 1602 the apparatus 102 is installed at a customer site. At 1604, the apparatus 102 is confirmed to be a healthy functioning apparatus. At 1606, the apparatus 102 is operated in a defined use of operation. At 1608, the sensors 110 sense characteristic sounds of the components 108 during the normal, healthy operation. At 1610, the analyzer module 1202 is trained to predict a temporal continuation of a sound snippet representing healthy machine function, as recorded by the sensors 110.

[0075] FIG. 17 discloses a computer-implemented method employing the analyzer module 1202 trained in connection with the method of FIG. 15, FIG. 16, and / or otherwise. It is to be appreciated that the ordering of the acts of the method is not limiting. As such, other orderings are contemplated herein. In addition, one or more acts may be omitted, and / or one or more additional acts may be included.

[0076] At 1702, the apparatus 102 is set up for operating in a particular mode. At 1704, the analyzer module 1202 predicts a sound snippet for the particular mode. At 1706, the sensors 110 sense sounds of the apparatus 102 during the procedure and generate signals indicative thereof. In one instance, the signals are pre-processed, e.g., by the de-identification module 1402 to remove human voice from the sensed sound, and / or otherwise, e.g., via noise cancelation, suppression, etc. In another instance, at least some of the pre-processing is omitted, e.g., the removal of human voice.

[0077] At 1708, the analyzer module 1202 compares the predicted sound snippet with the newly recorded snippet during continued operation, e.g. with techniques found in the Al (e.g., multi -head attention, scalar products, cosine similarity of latent vectors and projections thereof, etc.), and determines whether a difference between the predicted sound snippet and the newly recorded snippet satisfies a predetermined criteria, e.g., predefined / 1 earned threshold or if latent space components associated with certain failure modes arise.

[0078] At 1710, in response to the difference satisfying the predetermined criteria, the analyzer module 1202 generates a notification, as described herein and / or otherwise. For example, the notification can simply indicate a particular component may be malfunctioning or include additional information. Furthermore, in one instance, a report is also generated, as described herein and / or otherwise, with at least some of the information transmitted in and / or with the notification.

[0079] FIG. 18 discloses a computer-implemented method for training the analyzer module 1202 using supervised learning in-factory and on site. It is to be appreciated that the ordering of the acts of the method is not limiting. As such, other orderings are contemplated herein. In addition, one or more acts may be omitted, and / or one or more additional acts may be included.

[0080] At 1802, the apparatus 102 is assembled in-factory, including installation of the sensors 110. At 1804, the apparatus 102 is confirmed to be a healthy functioning apparatus. At 1806, the analyzer module 1202 is trained by presenting normal operation and different failure modes to the analyzer module 1202. The training data can be acquired in-factory and / or from one or more customer sites. At 1808, the analyzer module 1202 runs predicted and expected snippets (or their latent vector representations) through a neural network classifier, which determines probabilities for the different failure modes in a multi-class classification problem (among the modes included in supervised training).

[0081] This can include projecting out or computing differences between expected and measured sound data in latent space, akin to the algebraic transformations and similarity determinations used in attention mechanisms, with a neural network to link the latent space information with failure / issue states (N classes for N failure modes included in training). Each of the N failure modes is assigned a probability of presence by passing the output of the last neural network layer (with N neurons) through a softmax-function (i.e., a normalized exponential function,) and / or similar function to attain a probability. At 1810, the apparatus 102 is installed at a customer site and steps 1804-1808 are repeated as needed to update the analyzer module 1202 for the customer site.

[0082] FIG. 19 discloses a computer-implemented method for training the analyzer module 1202 using supervised learning on site. It is to be appreciated that the ordering of the acts of the method is not limiting. As such, other orderings are contemplated herein. In addition, one or more acts may be omitted, and / or one or more additional acts may be included.

[0083] At 1902, the apparatus 102 is installed at a customer site. At 1904, the apparatus 102 is confirmed to be a healthy functioning apparatus. At 1906, the analyzer module 1202 is trained by presenting normal operation and different failure modes to the analyzer module 1202. The training data can be acquired in-factory and / or from one or more customer sites. At 1908, the analyzer module 1202 runs predicted and expected snippets (or their latent vector representations) through a neural network classifier, which determines probabilities for the different failure modes in a multi-class classification problem (among the modes included in supervised training). At 1910, the analyzer module 1202 is updated as needed at the customer site.

[0084] FIG. 20 discloses a computer-implemented method employing the analyzer module 1202 trained in connection with the method of FIG. 18, FIG. 19, and / or otherwise. It is to be appreciated that the ordering of the acts of the method is not limiting. As such, other orderings are contemplated herein. In addition, one or more acts may be omitted, and / or one or more additional acts may be included.

[0085] At 2002, the apparatus 102 is set up for operating in a particular mode. At 2004, the sensors 110 sense sounds of the apparatus 102 during the procedure. In one instance, the signals are pre-processed, e.g., by the de-identification module 1402 to remove human voice from the sensed sound, and / or otherwise, e.g., via noise cancelation, suppression, etc. In another instance, at least some of the pre-processing is omitted, e.g., the removal of human voice. At 2006, the analyzer module 1202 predicts a sound snippet for the particular mode.

[0086] At 2008, the analyzer module 1202 compares the predicted sound snippet with the newly recorded snippet during continued operation, e.g. with techniques found in the Al (e.g., multi -head attention, scalar products, cosine similarity of latent vectors and projections thereof, etc.), and determines whether a difference between the predicted sound snippet and the newly recorded snippet satisfies a predetermined criteria, e.g., predefined / 1 earned threshold or if latent space components associated with certain failure modes arise. At 2010, in response to the difference satisfying the predetermined criteria, the analyzer module 1202 generates a notification, as described herein and / or otherwise. For example, the notification can simply indicate a particular component may be malfunctioning or include additional information. Furthermore, in one instance, a report is also generated, as described herein and / or otherwise, with at least some of the information transmitted in and / or with the notification. At 2012, the sound profiles and timings of different failure modes are labeled, and, optionally, the origin location of the failure is given.

[0087] At 2014, an incremental report is compiled by the system (e.g., once a month, etc.), with information about the functioning of the apparatus 102 and events detected in the past. The collected information can be stored, e.g., by a maintaining vendor automatically in a database (e.g., possibly a vector database) and / or otherwise. At 2016, statistics are determined from the information (e.g., in the database) about long-term machine behavior in the field. This may also include determining more subtle precursors of approaching failure modes, or more generally attain a detailed understanding of typical weak spots of the apparatus 102 (e.g., what usually breaks first, resolved and sorted for different operational modes in the database). At 2018, where relevant, update the analyzer module 1202 based on the statistics.

[0088] In one instance, post-market surveillance and predictive maintenance enables automated and scalable gathering of statistics about machine failure modes in the field, filterable by different operational modes, with the availability of prior “historical” data on the machine’s functioning sound. This can boost proactive product improvement and the quality of future developments.

[0089] In another embodiment, an LLM is first trained unsupervised to attain a semantic understanding of different objects, interactions, failure modes, and general physics / technical knowledge. Then, supervised learning is used to fine-tune the model to predict machine failure modes from sounds with an appropriately trained embedding model to transform the sound recordings to the latent vector domain of the LLM, as described herein. In one instance, the general knowledge of the LLM (from unsupervised training) will assist the model to recognize new failure modes, which have not been seen during supervised training.

[0090] The above is implemented by way of computer readable instructions, encoded, or embedded on the computer readable storage medium, which, when executed by a processor, cause the processor to carry out the described acts or functions. Additionally, or alternatively, the above is carried out by a signal, carrier wave or other transitory medium, which is not computer readable storage medium. While the invention has been illustrated and described in detail in the drawings and foregoing description, such illustration and description are to be considered illustrative or exemplary and not restrictive; the invention is not limited to the disclosed embodiments. Other variations to the disclosed embodiments can be understood and effected by those skilled in the art in practicing the claimed invention, from a study of the drawings, the disclosure, and the appended claims.

[0091] The word “comprising” does not exclude other elements or steps, and the indefinite article “a” or “an” does not exclude a plurality. A single processor or other unit may fulfill the functions of several items recited in the claims. The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to an advantage.

[0092] A computer program may be stored / distributed on a suitable medium, such as an optical storage medium or a solid-state medium supplied together with or as part of other hardware, but may also be distributed in other forms, such as via the Internet or other wired or wireless telecommunication systems. Any reference signs in the claims should not be construed as limiting the scope.

Claims

CLAIMSClaim 1. An apparatus (102), comprising: at least one component (108) that produces a characteristic sound during normal, healthy functioning and a different sound outside of the normal, healthy functioning; at least one sensor (110) disposed in spatial proximity to the at least one component to sense the characteristic sound and configured to sense the characteristic sound and generate a signal indicative of the characteristic sound; a memory (112) storing instructions (114) for a trained artificial intelligencebased analyzer module (1202) trained to identify the characteristic sound in the signal using audio analysis; and a controller (118) with a processor configured to execute the instructions to process signals from the at least one sensor (110) to predict a health state of the apparatus based on whether the characteristic sound is identified in the signal, wherein the controller is configured to determine the health state is not normal, healthy functioning and at least invokes transmission of a notification in response to not identifying the characteristic sound in the signal.Claim 2. The apparatus of claim 1, wherein at least one sensor further senses environmental background sound, and the trained artificial intelligence-based analyzer module is further trained to distinguish the characteristic sound from the environmental background sound that is additionally sensed by the at least one sensor.Claim 3. The apparatus of any of claims 1 to 2, wherein the trained artificial intelligencebased analyzer module is further trained to identify different failure modes, and the controller at least invokes transmission of the notification in response to the instructions predicting not normal, healthy functioning based on the different failure modes.Claim 4. The apparatus of any of claims 1 to 3, wherein the memory further includes a deidentification module (1402) configured to remove human voice from sound sensed by the at least one sensor.Claim 5. The apparatus of any of claims 1 to 4, wherein the trained artificial intelligencebased analyzer module is further trained to estimate a spatial location of the at least one component in the apparatus based on a spatial location of the at least one sensor, and thenotification includes the estimated spatial location of the at least one component.Claim 6. The apparatus of claim 5, wherein the trained artificial intelligence-based analyzer module is further trained to predict a type of the at least one component, and notification further includes the predicted type of the at least one component.Claim 7. The apparatus of claim 6, wherein the controller is further configured to invoke generation of a report that includes the predicted type of the at least one component, the spatial location of the at least one component, and a description of the not normal, healthy functioning.Claim 8. The apparatus of any of claims 1 to 7, wherein the artificial intelligence-based analyzer module is trained during assembly in one environment and installed and employed in a different environment.Claim 9. The apparatus of claim 8, wherein the trained artificial intelligence-based analyzer module is updated in the different environment based on the sound landscape of the different environment.Claim 10. The apparatus of any of claims 1 to 7, wherein the trained artificial intelligencebased analyzer module is trained and employed in a same environment.Claim 11. A computer-implemented method, comprising: training an analyzer module to identify a characteristic sound produced by a component, a sub-system, or a system of an apparatus during normal, healthy functioning, where the component, the sub-system, or the system produces a different sound outside of the normal, healthy functioning; sensing, with a sensor, sound produced by the component, the sub-system, or the system during operation of the apparatus; generating, with the sensor, a signal indicative of the sensed sound; performing audio analysis, with the trained analyzer module, on the signal to determine whether the signal includes the characteristic sound; and transmitting a notification in response to a result of the audio analysis not identifying the characteristic sound in the signal.Claim 12. The computer-implemented method of claim 11, further comprising: training the analyzer module to distinguish the characteristic sound from the environmental background sound that is additionally sensed by the at least one sensor.Claim 13. The computer-implemented method of any of claims 11 to 12, further comprising: training the analyzer module with different failure modes to identify the different failure modes.Claim 14. The computer-implemented method of any of claims 11 to 13, further comprising: training the analyzer module during assembly with a first sound landscape of the assembly environment.Claim 15. The computer-implemented method of claim 14, further comprising: updating the training of the analyzer module with a second sound landscape of a use environment, which is different from the assembly environment.Claim 16. A computer readable medium encoded with computer executable instructions, which, when executed by a processor, causes the processor to: train an analyzer module to identify a characteristic sound produced by a component, a sub-system, or a system of an apparatus during normal, healthy functioning, where the component, the sub-system, or the system produces a different sound outside of the normal, healthy functioning; receive a signal from a sensor of an apparatus, wherein the apparatus includes a component, a sub-system, or a system that produces a characteristic sound during operation of the apparatus, and the signal sensor is configured to sense sounds of the component, the subsystem, or the system, including the characteristic sound; perform audio analysis, with the trained analyzer module, on the signal to determine whether the signal includes the characteristic sound; and transmit a notification in response to a result of the audio analysis not identifying the characteristic sound in the signal.Claim 17. The computer readable medium of claim 16, wherein the instructions further cause the processor to: train the analyzer module to distinguish the characteristic sound from theenvironmental background sound that is additionally sensed by the at least one sensor.Claim 18. The computer readable medium of any of claims 16 to 17, wherein the instructions further cause the processor to: train the analyzer module with different failure modes to identify the different failure modes.Claim 19. The computer readable medium of any of claims 16 to 17, wherein the instructions further cause the processor to: train the analyzer module during assembly with a first sound landscape of the assembly environment.Claim 20. The computer readable medium of any of claims 16 to 17, wherein the instructions further cause the processor to: update the training of the analyzer module with a second sound landscape of a use environment, which is different from the assembly environment.

Citation Information

Patent Citations

  • Remote maintenance of medical imaging devices

    US20110121969A1

  • System, method and computer-accessible medium for machine condition monitoring

    US20200233397A1

  • Intelligent Audio Analytic Apparatus (IAAA) and Method for Space System

    US20200409653A1

  • Predicting vehicle health

    US20210335062A1

  • Anomalous sound detection with timbre separation

    US20220254366A1