System and method for enhanced data generation in fault diagnosis
By using text-based audio generation technology and leveraging a large language model and audio manipulation module, fault-related audio data is generated, solving the problem of scarce fault data in mechanical systems and improving the accuracy of fault diagnosis and training efficiency.
Patent Information
- Application Number
- CN202511101625.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-08-07
- Filing Date
- 2025-08-07
- Publication Date
- 2026-02-10
AI Technical Summary
Existing technologies face the problem of scarce fault data samples and severe noise interference in fault diagnosis of mechanical systems under harsh conditions, making it difficult to effectively train machine learning models.
We employ text-based audio generation technology, using data augmentation and data generation methods, and utilize a large language model (LLM) and audio manipulation module to generate fault-related audio data. We combine multimodal information for training to generate more fault samples to supplement the training data.
It effectively solves the problem of scarce fault data, improves the training efficiency and diagnostic accuracy of machine learning models, and enhances the fault detection capability under harsh conditions.
Smart Images

Figure CN121506173A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to artificial intelligence technologies for generating fault diagnosis data. Background Technology
[0002] Various systems are configured to perform tasks using machine learning (ML) or other artificial intelligence (AI) techniques. For example, systems configured to perform image or sound recognition, object detection, and / or other automated tasks can implement AI techniques. As an example, audio or sound detection systems and methods use various detection models trained for fault detection / diagnosis. Summary of the Invention
[0003] A method for generating audio to obtain manipulated audio data, the manipulated audio data including one or more audio features indicating a fault, the method comprising: receiving, at one or more processing devices, a text description of audio associated with device operation; receiving audio data associated with device operation; generating, based on the text description, a descriptive text input of audio features associated with device operation, the descriptive text input including at least one of audio characteristics of a fault associated with device operation, contextual information associated with device operation, and a condition associated with device operation; generating manipulated audio data based on the descriptive text input and the audio data, the manipulated audio data including one or more audio features indicating a fault associated with the descriptive text input; training a machine learning (ML) model using the manipulated audio data to diagnose the fault, the ML model being trained to generate an output indicating a fault based on audio data acquired during device operation; and outputting a trained ML model, the trained ML model being configured to generate the output indicating a fault, based on convergence during training.
[0004] Other embodiments include a non-transitory computer-readable storage medium configured to store instructions, which, when executed by a processor included in a computing device, cause the computing device to perform the steps of any of the methods described above. Further embodiments include systems, controllers, computing devices, etc., configured to perform the steps of any of the methods described above. Further embodiments include machines configured to perform the steps of any of the methods described above.
[0005] Other aspects and advantages of the invention will become apparent from the following detailed description taken in conjunction with the accompanying drawings, which illustrate the principles of the described embodiments by way of example. Attached Figure Description
[0006] Figure 1 A system for training a machine learning model based on the principles of this disclosure is generally illustrated.
[0007] Figure 2 The diagram generally illustrates a computer implementation method for training and implementing a machine learning model based on the principles of this disclosure.
[0008] Figure 3A An audio data tagging system based on the principles of this disclosure is generally illustrated.
[0009] Figure 3B A portion of a data capture system based on the principles of this disclosure is generally illustrated.
[0010] Figure 3C An alternative audio data tagging system based on the principles of this disclosure is generally illustrated.
[0011] Figure 4A An example audio generation system configured to perform audio generation and enhancement according to this disclosure is illustrated.
[0012] Figure 4B The illustration shows the steps of an example method for implementing an audio generation model based on the principles of this disclosure.
[0013] Figure 5 The diagram illustrates the interaction between a computer-controlled machine and a control system according to the principles of this disclosure.
[0014] Figure 6 The diagram illustrates the principles based on this disclosure. Figure 5 A schematic diagram of the control system, which is configured to control a vehicle, which may be a partially autonomous vehicle, a fully autonomous vehicle, a partially autonomous robot, or a fully autonomous robot.
[0015] Figure 7 The diagram shows... Figure 5 A schematic diagram of a control system configured to control manufacturing machines, such as stamping presses, cutting machines, or gun drills, of a manufacturing system (such as a part of a production line).
[0016] Figure 8 The diagram shows... Figure 5 A schematic diagram of a control system configured to control power tools, such as a drill or drive with at least a partially autonomous mode.
[0017] Figure 9 The diagram shows... Figure 5 A schematic diagram of the control system configured to control an automated personal assistant.
[0018] Figure 10 The diagram shows... Figure 5 A schematic diagram of a control system configured as a control and monitoring system, such as a control access system or a monitoring system.
[0019] Figure 11 The diagram shows... Figure 5 A schematic diagram of the control system configured to control an imaging system, such as an MRI device, an X-ray imaging device, or an ultrasound device. Detailed Implementation
[0020] This document describes embodiments of the present disclosure. However, it is to be understood that the disclosed embodiments are merely examples, and other embodiments may take various and alternative forms. The figures are not necessarily drawn to scale; some features may be enlarged or reduced to show details of specific components. Therefore, the specific structural and functional details disclosed herein should not be construed as limiting, but merely as a representative basis for teaching those skilled in the art to employ the embodiments in various ways. As will be understood by those skilled in the art, the various features illustrated and described with reference to any of the figures may be combined with features illustrated in one or more other figures to produce embodiments not explicitly illustrated or described. Combinations of illustrated features provide representative embodiments for typical applications. However, for a particular application or implementation, various combinations and modifications of features consistent with the teachings of this disclosure may be desired.
[0021] Unless the context clearly indicates otherwise, the terms “a,” “one,” and “the” used herein refer to both the singular and plural indicators. By way of example, “processor” programmed to perform various functions refers to one processor programmed to perform each function, or more than one processor programmed together to perform each of the various functions.
[0022] As used herein, “content” can refer to raw content corresponding to input data (e.g., data representing captured sound, images, videos, text, etc.) or synthetic content (e.g., synthetic sound or audio, images, videos, text, etc.). In some examples, “content” can include sound, which can correspond to captured sound, synthetic sound, or a combination thereof. Sound can be represented by sound data. In some contexts herein, the terms “sound” and “sound data” are used interchangeably. Similarly, “sound” and “audio” are used interchangeably. In examples, “sound” and / or “sound data” refer to the raw representation of sound, such as a numerical array representing sound level or volume, frequency, etc., which in some examples can include preprocessed data derived from sound or audio sensors. Conversely, “metadata” or “sound metadata” can refer to background or supplementary details about sound, such as size, format, creation date, geographic location data, etc. In various examples, “sound” and “sound data” may, but not necessarily, further include metadata.
[0023] Various systems are configured to perform tasks using machine learning (ML) or other artificial intelligence (AI) techniques. For example, systems configured to perform image or sound recognition, object detection, and / or other automated tasks can implement AI techniques. As an example, audio or sound detection systems and methods use various detection models trained for fault detection / diagnosis.
[0024] In various mechanical systems, operation under harsh conditions can lead to unexpected failures. Therefore, fault diagnosis of these components is a crucial aspect of system operation and maintenance. For example, vibration signals are commonly used to diagnose early failures in bearings and gears, and some techniques may include spectral analysis and handcrafted fault features coupled with classifiers. However, these methods are limited by noise and the rapidly increasing volume of machine data. This has led to the development of data-driven fault diagnosis, particularly using machine learning (ML) or deep learning (DL) techniques that can automatically extract representations from raw data. Several ML / DL models, including convolutional neural networks (CNNs) and recurrent neural networks (such as Long Short-Term Memory (LSTM) models), have shown promising results in fault diagnosis. However, most existing ML or DL-based methods require sufficient fault data, which is often difficult to obtain in industrial applications where failures are infrequent and short-lived.
[0025] In some examples, sound- or audio-based, or acoustic, sensing techniques can provide cost-effective monitoring and fault diagnosis. Acoustic sensing involves measuring sound waves generated by a system or process and using these measurements to estimate other physical quantities. Audio-based sensing provides information about the sound and vibration characteristics of a system, which can be used to detect faults or anomalies in the system and to improve predictive maintenance models. For example, sensed audio data including one or more faults may differ from “health” data (data excluding any faults, representing “normal” or “healthy” system operation), and different faults may have different audio or acoustic signatures. Acoustic signatures can also provide insights into system behavior, such as changes in operating conditions.
[0026] One challenge associated with audio-based sensing for fault detection is the limited number of fault samples in the audio data (e.g., within a given audio stream). For example, in a given dataset of audio data sensed by the system, there exists a very large amount of “healthy” data and a very small amount of “faulty” data (i.e., sound data or data points indicating a fault). Therefore, a very large amount of data must be collected, processed, and analyzed to identify the very small number of faults.
[0027] Various techniques, such as transfer learning and data generation, can be used to address the problem of limited fault samples. For example, transfer learning involves learning knowledge using an additional complete dataset and applying it to the target dataset, while data generation focuses on using oversampling methods or generative adversarial networks (GANs) to generate synthetic samples. GANs can facilitate learning the distribution characteristics of vibration signals and generating synthetic fault data. However, in industrial applications, the ratio of healthy data to fault data is often very high, making it difficult to effectively train GANs.
[0028] The data (e.g., audio) generation system and method according to this disclosure are configured to implement data (e.g., audio data) generation techniques, including audio manipulation based on descriptive text (or text guidance), which may be referred to as "data augmentation." For example, text-guided audio manipulation involves using text input to control and manipulate audio signals. In this way, audio content can be modified and generated based on text descriptions. Various techniques for text-guided audio manipulation include, but are not limited to, audio style transfer, conditional generative models, and audio effects.
[0029] Audio style transfer techniques involve modifying the characteristics of an audio signal based on a specified style or attribute in the input text. By leveraging deep learning techniques such as neural networks and generative models, audio style transfer can transform the timbre, pitch, or emotional content of an audio signal to match the specifications of the desired textual guidance.
[0030] Conditional generative modeling techniques include structured prediction methods that model the complete distribution of probabilities over the joint configuration of the output. These techniques are used for text-guided audio manipulation to generate audio aligned with a given text cue. Conditional generative models learn the relationship between text input and audio output, allowing the generation of novel audio samples based on specific text cues.
[0031] Audio effects technology uses text prompts to control audio effects and processing parameters. By specifying the desired effect or adjustment, such as reverb, echo, or equalization, in the text, the algorithm can apply appropriate audio processing techniques to modify the input audio accordingly. Audio effects technology uses natural language instructions to achieve interactive and expressive manipulation of audio.
[0032] Data augmentation and data generation may require resampling the distribution of existing data to generate additional data, or manual control to adjust input parameters or features to achieve a desired output. These constraints limit the scope of the generated data and restrict the ability to generate realistic data for unseen distributions or scenarios. The systems and methods of this disclosure are configured to implement data generation techniques that incorporate data augmentation by leveraging a broad array of sources such as online audio, text, and knowledge resources, as well as the interrelationships between these sources. Further data can be generated by manipulating input data that meets specified criteria, by providing textual descriptions of physical devices and contextual elements of those devices, including conditions, environment, behavior, etc.
[0033] These systems and methods leverage advancements in foundational models to capture general nuances, factual knowledge, and contextual understanding, providing a dynamic basis for generating coherent, context-sensitive responses. The techniques described in this paper are adaptable to specific applications such as text generation, language translation, and question answering, and are therefore suitable for a wide range of tasks.
[0034] The audio generation system and method disclosed herein can implement one or more types of models. As an example, the CLAP (Contrastive Language-Audio Pre-trained) model includes a neural network trained on various (audio, text) pairs. Given an audio sample, the model can be instructed to predict the most relevant text fragment or content without direct optimization for the task. For example, the CLAP model can use a shift window (SWIN) transformer to obtain audio features from log-Mel spectrogram input and a robustly optimized BERT (Bidirectional Encoder Representation and Transformer) pre-trained method (RoBERTa) model to obtain text features. Both text and audio features are then projected into a latent space with the same dimension. The dot product between the projected audio and text features is then used as a similarity score.
[0035] As another example, Large Language Models (LLMs) are configured to fully understand and generate human-like text. LLMs are trained on a wide range of text data to master the nuances of language and achieve coherent responses. The benefits of LLMs include natural language understanding, text creation, and task automation (e.g., customer support, translation, research assistance, personalization, innovation facilitation, educational support, enhanced creativity across various domains, etc.).
[0036] Figure 1An example system 100 for training ML or other AI models, such as audio generation (or synthesis) models according to this disclosure, is shown. System 100 can be configured (and / or includes circuitry configured as follows) to implement the systems and methods of this disclosure, which are described in more detail below. System 100 may include an input interface for accessing training data 102 of the audio generation model. For example, as Figure 1 As illustrated, the input interface can be comprised of a data storage interface 104, which can access the training data 102 from the data storage device 106. For example, the data storage interface 104 can be a memory interface or a persistent storage interface (e.g., a hard disk or SSD interface), but it can also be a personal area network (PAN), local area network (LAN), or wide area network (WAN) interface, such as a Bluetooth, Zigbee, or Wi-Fi interface, or an Ethernet or fiber optic interface. The data storage device 106 can be an internal data storage device of the system 100, such as a hard disk drive or SSD, but it can also be an external data storage device, such as a network-accessible data storage device.
[0037] In some embodiments, the data storage device 106 may further include a data representation 108 of an untrained version of the audio generation model, which may be accessed by the system 100 from the data storage device 106. However, it will be appreciated that the training data 102 and the data representation 108 of the untrained audio generation model may also be accessed from different data storage devices, for example, via different subsystems of the data storage interface 104. Each subsystem may be of the type of data storage interface 104 as described above.
[0038] In some embodiments, the data representation 108 of the untrained audio generation model can be generated internally by the system 100 based on the design parameters of the audio generation model, and therefore the data representation 108 may not be explicitly stored on the data storage device 106. The system 100 may also include a processor subsystem 110, which can be configured to provide an iterative function as a substitute for the stacked layers of the audio generation model to be trained during operation of the system 100. Here, the corresponding layers in the substituted stacked layers may have mutually shared weights and may receive the output of the previous layer as input, or, for the first layer of the stacked layers, receive initial activation and a portion of the input of the stacked layers.
[0039] Processor subsystem 110 can also be configured to iteratively train the audio generation model using training data 102. Here, the training iterations performed by processor subsystem 110 may include a forward propagation portion and a backward propagation portion. Processor subsystem 110 can be configured to perform the forward propagation portion by determining an equilibrium point of the iterative function in other operations defining the executable forward propagation portion, and by providing the equilibrium point as a substitute for the output of the stacked layers in the audio generation model, at which the iterative function converges to a fixed point, wherein determining the equilibrium point includes using a numerical root-finding algorithm to find the root solution of the iterative function minus its input. Processor subsystem 110 is configured to train the audio generation model according to the systems and methods of this disclosure, which are described in more detail below.
[0040] System 100 may also include an output interface for outputting a data representation 112 of the trained audio generation model. This data may also be referred to as the trained model data 112. For example, as... Figure 1 As illustrated, the output interface can be comprised of a data storage interface 104, wherein in these embodiments, the interface is an input / output (“IO”) interface through which trained model data 112 can be stored in data storage device 106. For example, the data representation 108 defining the “untrained” audio generation model can be at least partially replaced by the data representation 112 of the trained audio generation model during or after training, because the parameters of the audio generation model (such as the weights, hyperparameters, and other types of parameters) can be adapted to reflect training on the training data 102. This also... Figure 1 The figures are illustrated by reference numerals 108 and 112, which refer to the same data records on data storage device 106. In some embodiments, data representation 112 may be stored separately from data representation 108 that defines the "untrained" audio generation model. In some embodiments, the output interface may be separate from data storage interface 104, but it can generally be of the type described above for data storage interface 104.
[0041] Figure 2An example content generation system 200 is depicted, configured (and / or including circuitry configured as follows) to implement a system for annotating, enhancing, and / or generating data. The content generation system 200 may include at least one computing system 202 configured to implement all or part of the systems and methods of this disclosure as explained in more detail below. The computing system 202 may include at least one processor 204 operatively connected to a memory unit 208. The processor 204 may include one or more integrated circuits implementing the functions of a central processing unit (CPU) 206. The CPU 206 may be a commercially available processing unit implementing an instruction set, such as one of the x86, ARM, Power, or MIPS instruction set families. Various components of system 200 may be implemented using the same or different circuitry.
[0042] During operation, CPU 206 may execute stored program instructions retrieved from memory unit 208. The stored program instructions may include software that controls the operation of CPU 206 to perform the operations described herein. In some embodiments, processor 204 may be a system-on-a-chip (SoC) that integrates the functionality of CPU 206, memory unit 208, network interface, and input / output interface into a single integrated device. Computing system 202 may implement an operating system for managing various aspects of operation.
[0043] Memory cell 208 may include volatile and non-volatile memory for storing instructions and data. Non-volatile memory may include solid-state memory, such as NAND flash memory, magnetic and optical storage media, or any other suitable data storage device that retains data when the computing system 202 is deactivated or loses power. Volatile memory may include static and dynamic random access memory (RAM) for storing program instructions and data. For example, memory cell 208 may store one or more machine learning models (e.g., in...). Figure 2 The data are represented as machine learning model 210 or algorithm, training dataset 212 of machine learning model 210, original source dataset 216, etc.
[0044] The computing system 202 may include a network interface device 222 configured to provide communication with external systems and devices. For example, the network interface device 222 may include wired and / or wireless Ethernet interfaces as defined by the Institute of Electrical and Electronics Engineers (IEEE) 802.11 family of standards. The network interface device 222 may include a cellular communication interface for communicating with cellular networks (e.g., 3G, 4G, 5G). The network interface device 222 may also be configured to provide a communication interface to an external network 224 or the cloud.
[0045] External network 224 may be referred to as the World Wide Web or the Internet. External network 224 can establish standard communication protocols between computing devices. External network 224 can allow easy exchange of information and data between computing devices and the network. One or more servers 230 can communicate with external network 224.
[0046] The computing system 202 may include an input / output (I / O) interface 220, which may be configured to provide digital and / or analog inputs and outputs. The I / O interface 220 may include an additional serial interface (e.g., a Universal Serial Bus (USB) interface) for communicating with external devices.
[0047] The computing system 202 may include a human-machine interface (HMI) device 218, which may include any device that enables the system 200 to receive control input. Examples of input devices may include HMI inputs such as a keyboard, mouse, touchscreen, voice input device, and other similar devices. The computing system 202 may include a display device 232. The computing system 202 may include hardware and software for outputting graphical and textual information to the display device 232. The display device 232 may include an electronic display screen, projector, printer, or other suitable device for displaying information to a user or operator. The computing system 202 may also be configured to allow interaction with remote HMIs and remote display devices via a network interface device 222.
[0048] System 200 can be implemented using one or more computing systems. While the example depicts a single computing system 202 implementing all the described features, the various features and functions can be decoupled and implemented by multiple computing units that communicate with each other. The specific system architecture chosen may depend on various factors.
[0049] System 200 may implement a machine learning model 210 for analyzing raw source dataset 216. For example, CPU 206 and / or other circuitry may implement machine learning model 210. Raw source dataset 216 may include raw or unprocessed sensor data that may represent the input dataset for the machine learning system. Raw source dataset 216 may include audio, images, video, video clips, audio, text-based information, and raw or partially processed sensor data (e.g., radar charts of objects). In some embodiments, machine learning model 210 may be a deep learning or neural network algorithm designed to perform a predetermined function. For example, a neural network algorithm may be configured to identify events or objects based on audio data.
[0050] Computer system 202 may store training dataset 212 for machine learning model 210. Training dataset 212 may represent a collection of previously constructed data used to train machine learning model 210. Machine learning model 210 may use training dataset 212 to learn various conditions and other factors (e.g., weighting factors) associated with the ML algorithm. Training dataset 212 may include a collection of source data that has the corresponding results or outcomes that machine learning model 210 attempts to replicate through the learning process.
[0051] Machine learning model 210 can operate in learning mode using training dataset 212 as input. Machine learning model 210 can be executed through multiple iterations using data from training dataset 212. With each iteration, machine learning model 210 can update its internal weighting factors based on the achieved results. For example, machine learning model 210 can compare its output (e.g., generated content) with those included in training dataset 212. Since training dataset 212 includes expected results, machine learning model 210 can determine when performance is acceptable. Once machine learning model 210 achieves a predetermined performance level (e.g., 100% consistency with results associated with training dataset 212), machine learning model 210 can be executed using data not in training dataset 212. The trained machine learning model 210 can be applied to new datasets to generate content. Machine learning model 210 may include an audio generation model trained according to the systems and methods of this disclosure.
[0052] Machine learning model 210 can be configured to identify specific features in raw source data 216. Raw source data 216 may include multiple instances or input datasets (e.g., audio data, audio streams, images, video streams or segments including audio data, etc.) for its desired output. By way of example only, machine learning model 210 can be configured to identify objects, features, or events in audio segments based on audio data. In some examples, machine learning model 210 can be configured to annotate the identified objects, features, or events. Machine learning model 210 can be configured to perform audio generation according to the principles of this disclosure. Machine learning model 210 can be programmed to process raw source data 216 to identify the presence of specific features. Machine learning model 210 can be configured to identify features in raw source data 216 as predetermined features. Raw source data 216 can be derived from various sources. For example, raw source data 216 can be actual input data collected by a machine learning system. Raw source data 216 can be machine-generated for testing the system. By way of example, raw source data 216 may include raw audio data, audio data from a microphone, etc.
[0053] In the example, machine learning model 210 can process raw source data 216 and output audio data including one or more indications of the identified features or events. Machine learning model 210 can generate confidence levels or factors for each generated output. For example, a confidence value exceeding a predetermined high confidence threshold can indicate that machine learning model 210 is confident that the identified event (or feature) corresponds to a specific event. A confidence value below a low confidence threshold can indicate that machine learning model 210 has some uncertainty about the existence of a specific feature.
[0054] like Figure 3A and Figure 3B As generally illustrated herein, example system 300 may include an image (e.g., image and / or video) capture device 302, an audio capture array 304, and a computing system 202. This system may receive video stream data associated with a data capture environment from the image capture device 302. System 202 may be configured to perform video object detection to identify one or more objects in the corresponding image of the video stream data. System 202 may receive audio stream data corresponding to at least a portion of the video stream data from the audio capture array 304. The audio capture array 304 may include one or more microphones 306 or other suitable audio capture devices. The systems and methods described herein may be configured to label at least some objects in the video stream data and / or audio stream data using output from at least a first machine learning model (e.g., machine learning model 210 or other suitable machine learning models configured to provide outputs including predictions of one or more object or event detections).
[0055] System 202 can calculate (e.g., using at least one probability-based function or other suitable technique or function) an offset value for at least a portion of audio stream data corresponding to at least one tagged object in the video stream data, based on at least one data capture characteristic. System 202 can use at least one offset value to synchronize at least a portion of the video stream data with a portion of the audio stream data corresponding to at least one tagged object in the video stream data. The at least one data capture characteristic may include one or more characteristics of at least one image capture device, one or more characteristics of at least one audio capture array, one or more characteristics corresponding to the position of at least one image capture device relative to at least one audio capture array, one or more characteristics corresponding to the movement of objects in the video stream data, one or more other suitable data capture characteristics, or a combination of the above.
[0056] System 202 can label at least a portion of audio stream data corresponding to at least one labeled object of the video stream data using one or more labels and at least one offset value of labeled objects of the video stream data. Each corresponding label may include an event type, an event start indicator, and an event end indicator. System 202 can generate training data using at least some of the labeled portions of the audio stream data. System 202 can use the training data to train a second machine learning model. System 202 can use the second machine learning model to detect one or more sounds associated with the audio data provided as input to the second machine learning model. The second machine learning model may include any suitable machine learning model and can be configured to perform any suitable function, such as those described herein with reference to Figure 4- Figure 11 The functions described.
[0057] In some embodiments, such as Figure 3C As generally illustrated, computing system 202 can be configured to tag audio data based on sensor data received from one or more sensors, such as those described herein or any other suitable sensors or combinations thereof. System 202 can receive audio stream data associated with a data capture environment from audio capture array 354 or any suitable audio capture device (such as one or more of microphones 306 or other suitable audio capture devices). It should be understood that audio capture array 354 may include features similar to those of audio capture array 304 and may include any suitable number of audio capture devices. System 202 can receive sensor data associated with a data capture environment from at least one sensor (e.g., such as sensor 352) asynchronously relative to audio capture array 354. Sensor 354 may include at least one of the following: induction coil, radar sensor, LiDAR sensor, sonar sensor, image capture device, any other suitable sensor, or a combination thereof. Audio capture array 354 may be located remotely from sensor 354, located close to sensor 354, or located in any suitable relationship with sensor 354.
[0058] System 202 can use the output from at least a first machine learning model (such as machine learning model 210 or other suitable machine learning model) to identify at least some events in sensor data. Machine learning model 210 can be configured to provide an output including one or more event detection predictions based on sensor data. System 202 can synchronize at least a portion of sensor data associated with a portion of audio stream data, that portion of audio stream data corresponding to at least one event of sensor data. System 202 can use one or more labels extracted for corresponding events of sensor data values to label at least a portion of audio stream data corresponding to at least one event of sensor data. Each corresponding label may include an event type, an event start indicator, and an event end indicator. System 202 can use at least some of the labeled portions of audio stream data to generate training data. System 202 can use the training data to train a second machine learning model. System 202 can use the second machine learning model to detect one or more sounds associated with audio data provided as input to the second machine learning model. The second machine learning model may include any suitable machine learning model and can be configured to perform any suitable function, such as those described herein with respect to Figures 4 to 5. Figure 11 The functions described.
[0059] The audio generation systems and methods disclosed herein (e.g., any one of systems 100, 200, etc.) are configured to train an audio generation model (e.g., model 210) to modify and generate audio content based on text descriptions (e.g., perform data or audio enhancement), as described in more detail below.
[0060] In various use cases and applications, collecting data under normal operating conditions (i.e., health status) is straightforward. However, acquiring fault data for training machine learning models can be challenging or prohibitively costly. In some instances, collecting such data may be impractical before deploying physical equipment or systems and experiencing actual failures. Many applications still struggle with insufficient data for effective ML model training. To address the lack of audio signals representing fault conditions (or any desired conditions), the techniques disclosed herein introduce data generation techniques for audio signals. These techniques involve manipulating source / reference audio signals using inputs from multiple modalities and conditions. These inputs can encompass textual descriptions, sample-style audio, or various conditions / modalities, thus providing a general solution.
[0061] The systems and methods disclosed herein include a description manager configured to receive input queries or desired conditions of a physical device / machine / system to generate descriptions or information that are more relevant and necessary for an audio manipulator. The audio manipulator (e.g., an audio manipulator or manipulation module, circuitry, etc.) creates audio content by manipulating input reference audio based on the guidance and information provided by the description manager.
[0062] In examples where expert knowledge is unavailable, a description manager can leverage an LLM (Limited Learning Model) as a shared knowledge base to aggregate comprehensive insights into the device and its expected behavior. By providing the LLM with precise instructions or hints about the device, the context, and specific characteristics or behaviors of interest, the LLM can then generate textual descriptions or related information. For example, an LLM can handle queries such as, “What type of sound does a wind turbine produce when there is a fracture fault in its bearings and gears?” This process helps to effectively bridge knowledge gaps.
[0063] Based on the information obtained about the equipment, the instructions provided to the LLM can be enriched by incorporating a wide range of contextual details. These details can cover factors such as the materials the equipment is made of and its operating settings (e.g., within a vehicle or factory). For example, instructions might state: “Locate the fuel pump in the front trunk of the vehicle to ensure that the generated audiometer reflects the surrounding ambient sound.” This approach ensures that the LLM generates output that is more realistically aligned with the specified conditions.
[0064] In examples where expert knowledge or application knowledge bases are accessible, sought-after information about the application can be directly retrieved through queries. The LLM can then be used to restate or restructure this knowledge to align with the requirements of the audio manipulation module. For instance, by analyzing parameters such as operation duration, current system temperature, and other monitored metrics, users can anticipate potential deviations from normal system behavior.
[0065] In some implementations, the description manager can leverage the capabilities of LLM to transform existing shared knowledge into more comprehensive insights that define the system's key physical behaviors and properties. This involves converting information into a format suitable for input to the audio manipulation module, which is achieved by creating prompts. An example instruction might be, "Arrange subsequent details into the subsequent structural outline…".
[0066] In some implementations, the audio manipulation module includes: a reference audio encoder configured to extract audio embeddings; a text encoder configured to extract text embeddings; a style encoder configured to extract style or pattern embeddings; and an ML / DL model (e.g., a diffusion model) trained to manipulate audio by modulating all embeddings shared within the latent space.
[0067] In some variations, the audio manipulation module may also include: an image encoder configured to extract an image embedding; and a signal “X” encoder (e.g., a haptic encoder) configured to extract an X embedding.
[0068] As used in this paper, embedding refers to the numerical representation of an object in a continuous vector space, configured to capture semantic relationships between entities. In other words, embedding maps items from a high-dimensional discrete space (such as vocabulary words) to a low-dimensional continuous space. Here, items with similar meanings are more closely grouped together based on their semantic similarity. Embeddings are typically pre-trained on extensive datasets using techniques such as Word2Vec, GloVe, or FastText, which analyze co-occurrence patterns of words in text, or even across modal extensions such as CLAP (for text and audio) or CLIP (contrastive language-image pre-training; for text and images).
[0069] In some implementations, various base models can be used to derive relevant embeddings for different modalities. For example, models such as wav2vec or the Hierarchical Token Semantic Audio Transformer (HT-SAT) can be used for audio data, while RoBERTa or T5 models can be used for text, and Visual Geometry Groups (VGG) or ResNet models can be used for image data. Some models can be configured to achieve joint processing across multiple modalities, such as CLAP or CLIP.
[0070] In some examples, the reference or original (i.e., unmanipulated / enhanced) audio may refer to recordings taken during typical or optimal operating conditions of the physical equipment or machine, such as the sound of a running fuel pump. Textual descriptions outlining key characteristics, such as “significant high pitch with intermittent crackling noise” or “sound captured when the machine is operating on a street with traffic noise,” can be included / provided with the original audio. Furthermore, styled audio clips can be used, including samples encapsulating sound patterns associated with the machine, such as screams, brittleness, grating, etc. To enhance contextual understanding, images can be introduced as a complementary component, providing relevant insights into the surrounding environment and context. For example, an image depicting the placement of a fuel pump in the front trunk located on a street can add valuable context.
[0071] In some implementations, when sample audio of the surrounding environment or anticipated noise is available, a supplementary or duplicated encoder (such as a “style” encoder) can be used. This supplementary encoder is introduced to extract specific information related to environmental factors or noise characteristics, thereby enhancing the generation process of the target audio.
[0072] In some implementations, when additional modalities (such as tactile or surface sensor data) are available, the corresponding base model can be used to extract the signal embedding. As an example, some base models can be adapted to different modalities, as demonstrated by utilizing a ResNet18 model tailored for processing tactile data.
[0073] In some examples, the audio manipulation module is configured to receive or acquire the spectrum of audio (such as a spectrogram) as input and subsequently produce a manipulated spectrogram as output. In these examples, a vocoder is used to reconstruct the audio waveform from the spectrogram.
[0074] In some examples, diffusion models can be used to provide speech generation, catering to both waveform and mel spectrogram formats. A diffusion model can include two distinct processes: a forward process, which facilitates the transformation of the data distribution to a standard Gaussian distribution by implementing a predefined noise schedule; and a backward process, which is responsible for gradually generating data samples based on the noise, carefully guided by the inference noise schedule. This approach ensures the production of data samples aligned with the desired output.
[0075] In some implementations, all embeddings coexist within a shared space, containing valuable cross-modal information. During the generation phase, audio synthesis is initiated by subjecting the input audio representation to a denoising process, which serves as an initial step in the reverse process. This denoising process is conditioned on other cross-modal representations and progressively refines the generation of the audio output.
[0076] In some generative implementations, including weights within each condition can significantly enhance control over the generated data. This integration of weights not only amplifies the commands to the generated outcomes but also enhances the diversity within the generated dataset.
[0077] In some examples, data generated by manipulating reference audio to simulate different fault conditions can be effectively used to augment training datasets for subsequent (e.g., downstream) machine learning or deep learning models. These downstream models can include tasks such as distinguishing between healthy and faulty data, predicting anomalous behavior, or converting audio into another modality (such as torque data).
[0078] The techniques disclosed herein can be extended beyond generating only fault data to include the generation of any type of negative data. Furthermore, these techniques avoid the need for domain-specific feature extraction or preprocessing, allowing implementation across a wide range of audio applications. In some embodiments, the generated data can be used to train experts or provide insights into previously unseen audio signals.
[0079] As described in more detail below, the audio generation system and method of this disclosure implement one or more of the techniques described above for data generation and fault diagnosis. For example, enhanced audio signals (e.g., synthetic audio data) are generated by manipulating input or reference audio (“x” signal). This manipulation is conditioned on textual descriptions, styles, and / or any conditions that can be extracted from the “x” signal or modality. These techniques are used to create synthetic audio data for training machine learning-based models. Synthetic audio data is generated by applying different data augmentation methods to pre-existing physical data (e.g., the original, input, or reference audio). By using an expanded dataset that includes synthetic data, the accuracy and resilience of predictive maintenance systems can be improved, thereby enhancing the effectiveness and reliability of classifying health and fault states, predicting faults, etc.
[0080] The example audio generation system includes a description manager and an audio manipulation module. The description manager is configured to generate a text description of the audio that outlines the characteristics or properties of the audio to be manipulated. To achieve this, the description manager uses an LLM (Local Level Management), which is configured to generate descriptive text based on (i) instructions derived from the physical device or system generating the audio signal and (ii) the conditions under which the audio is to be manipulated. Furthermore, the description manager can use the LLM to organize or restate information derived from experts or knowledge bases into a desired structure.
[0081] The audio manipulation module includes various encoders, including audio, text, style, and other encoders associated with desired conditions. The audio manipulation module implements a trained model (such as a diffusion model) that operates simultaneously conditioned on text embeddings, style embeddings, and other data embeddings within a continuous latent space. Given a pre-trained model, this process can be reversed to facilitate the generation of manipulated audio.
[0082] Figure 4A The figure illustrates an example audio generation system 400 according to the present disclosure, which is configured to perform audio generation and enhancement. For example, one or more computing devices, processors, or processing devices are configured to execute instructions to implement the functionality of the audio generation system 400, such as one or more processors of the systems described herein (e.g., 100, 200, etc.).
[0083] The audio generation system includes a description manager 402 and an audio manipulation module 404. As described above, the description manager 402 is configured to generate a text description of the audio, which summarizes the characteristics or properties of the audio being manipulated. For example, the LLM 406 is configured to generate the text description based on instructions or prompts derived from a physical device, machine, system, etc., and the LLM 406 generates sound and / or audio signals. In some examples, the LLM 406 may also receive one or more inputs indicating the conditions under which the audio is to be manipulated to further determine the text description generated by the LLM 406. These conditions may include, but are not limited to, information such as the constituent materials of the device, operating settings, location, or environment.
[0084] In some examples, LLM 406 can be configured to organize or restate information from experts or knowledge base 408 into a desired structure. Knowledge base 408 can be a repository of information about specific devices, components, faults, etc., and can include textual descriptions of the behavior of devices and components (“healthy” sounds, fault sounds, sound descriptions of specific faults, etc.). For example, queries / prompts about devices, situations, and specific characteristics or behaviors of interest are entered into knowledge base 408, and can then be provided to LLM 406 to generate textual descriptions or information to be used by description manager 402. For example, LLM 406 can generate textual descriptions in response to queries such as “What type of sound does a wind turbine produce when there is a fracture fault in its bearings and gears?”
[0085] Description manager 402 receives text descriptions from LLM 406 to transform existing shared knowledge into a more comprehensive insight defining the device's key physical behaviors and attributes. For example, description manager 402 is configured to convert text descriptions received from LLM 406 into a format suitable for input into audio manipulation module 404, referred to herein as "descriptive text input." For example, descriptive text input may include one or more categories or types of information, such as characteristics, context, conditions, and styles. For instance, characteristics may indicate one or more characteristics of the audio signal to be generated / manipulated (e.g., "with broken high-pitched whimpering noise"). Conversely, conditions and context may indicate the device, location, environment, and other conditions emitting the sound (e.g., "with engine sounds of road traffic"). Styles may indicate sound characteristics such as timbre, pitch, emotional content, etc.
[0086] Audio manipulation module 404 is configured to generate audio data (e.g., manipulated audio) based on original audio received from descriptive text input from description manager 402, and in some examples, based on one or more other inputs, such as style audio (e.g., user-provided audio / sound samples or examples) and other conditions or modalities (e.g., images). For example, audio manipulation module 404 includes or implements various encoders, such as audio encoders, text encoders, style encoders, etc. Audio manipulation module 404 is configured to implement trained models (such as diffusion models) that operate conditioned on text embeddings, style embeddings, and other data embeddings within a continuous latent space. Audio manipulation module 404 may also include or implement one or more base models 410, such as wav2vec or hierarchical token semantic audio transformer (HT-SAT) models, RoBERTa or T5 models, visual geometry groups (VGG) or ResNet models, etc.
[0087] In this way, the audio manipulation module 404 is configured to generate / synthesize manipulated audio (e.g., a manipulated version of the original audio) using descriptive text input. In other words, the original "healthy" audio is manipulated to include audio features indicating various faults associated with the corresponding device (e.g., as described by the descriptive text input). For example, based on the descriptive text input, the audio manipulation module 404 adds audio features or spectrograms indicating various faults (e.g., grinding sounds, squeaks, whoos, knocking sounds, clicks, higher frequencies, lower frequencies, etc.). In some examples, an audio synthesis device such as a vocoder can be used to generate the manipulated audio.
[0088] Observed audio (e.g., actual audio acquired during device operation, such as raw audio, which may contain both healthy audio and audio including various faults) and manipulated audio (e.g., manipulated audio including fault data indicating one or more faults) can be provided as input to train one or more ML / DL models 412. In this way, ML / DL models 412 can be trained to detect, identify, and diagnose faults in a device or system based on audio data acquired during device operation.
[0089] Figure 4B The figure illustrates steps of an example method 440 for implementing an audio generation model (e.g., training and subsequently using it to perform audio generation) according to the principles of this disclosure. For example, one or more processors or processing devices are configured to execute instructions to implement method 440, such as one or more processors in the system described herein.
[0090] At 442, method 440 includes using LLM to generate a text description of audio / sound associated with the operation of a device, machine, system, etc., which includes a description of a fault associated with the device, the sound caused by the fault, the cause of the fault (e.g., the component that causes the specific fault sound), etc.
[0091] At 444, method 440 includes generating descriptive text input based on text descriptions generated by the LLM. For example, the descriptive text input includes various categories associated with the operation of a particular device and the corresponding audio produced by the operation of that device, such as characteristics (e.g., sound characteristics), conditions, context, style, etc., as described herein.
[0092] At 446, method 440 includes generating manipulated audio based at least on descriptive audio and original audio (e.g., audio signals or audio data acquired during operation of the device or similar device). For example, based on descriptive text input, audio features or spectrograms indicating various faults are added to the original audio to obtain manipulated (synthesized) audio or audio data.
[0093] At 448, method 440 includes training one or more ML or DL models using manipulated audio and the original (e.g., observed) audio.
[0094] At 450, method 440 includes using a trained ML or DL model to detect, identify, and / or diagnose faulty equipment in real time, based on audio data acquired during equipment operation (i.e., observed audio data) and / or using previously recorded audio data.
[0095] At 452, method 440 includes controlling one or more functions of a device, system, machine, etc., based on a detected or diagnosed fault. For example, information about the fault can be used for various downstream tasks, such as controlling or adjusting the operating parameters or functions of the device, stopping the operation of the device, generating alarms to the device operator, storing and / or transmitting data indicating that one or more faults have been diagnosed, etc. In some examples, method 440 includes controlling the following... Figures 5 to 11 The functionality of any system described in the system.
[0096] Figures 5 to 11 Example systems and devices that can implement the audio generation model according to this disclosure are described. Figure 5A schematic diagram depicting the interaction between a computer-controlled machine 500 and a control system 502 is provided. In this example, the control system 502 is configured to control the computer-controlled machine 500 by executing an audio generation model according to the principles of this disclosure. The computer-controlled machine 500 includes actuators 504 and sensors 506. Actuators 504 may include one or more actuators, and sensors 506 may include one or more sensors. Sensors 506 are configured to sense the condition of the computer-controlled machine 500. Sensors 506 may be configured to encode the sensed condition into a sensor signal 508 and transmit the sensor signal 508 to the control system 502. Non-limiting examples of sensors 506 include video, radar, LiDAR, ultrasonic, and motion sensors. In some embodiments, sensor 506 is an audio sensor configured to sense sound (audio or sound data) in the environment approaching the computer-controlled machine 500. The audio generation model according to this disclosure can be used to perform audio generation using audio data as described herein.
[0097] The control system 502 is configured to receive sensor signals 508 from the computer-controlled machine 500. As described below, the control system 502 can also be configured to calculate actuator control commands 510 based on the sensor signals and transmit the actuator control commands 510 to the actuators 504 of the computer-controlled machine 500.
[0098] like Figure 5 As shown, the control system 502 includes a receiving unit 512. The receiving unit 512 can be configured to receive sensor signals 508 from sensor 506 and transform the sensor signals 508 into input signals x. In an alternative embodiment, the sensor signals 508 are received directly as input signals x, without the need for a receiving unit 512. Each input signal x may be a portion of each sensor signal 508. The receiving unit 512 can be configured to process each sensor signal 508 to generate each input signal x. The input signals x may include data corresponding to the image recorded by sensor 506.
[0099] The control system 502 includes a classifier 514. The classifier 514 can be configured to classify an input signal x into one or more labels using a machine learning (ML) algorithm (such as a neural network). For example, classifier 514 corresponds to classifier 408 described above. Classifier 514 is configured to be parameterized by parameters (such as those described above, e.g., parameter θ). Parameter θ can be stored in and provided by non-volatile storage device 516. Classifier 514 is configured to determine an output signal y based on the input signal x. Each output signal y includes information assigning one or more labels to each input signal x. Classifier 514 can transmit the output signal y to conversion unit 518. Conversion unit 518 is configured to convert the output signal y into an actuator control command 510. Control system 502 is configured to transmit actuator control command 510 to actuator 504, actuator 504 being configured to actuate computer-controlled machine 500 in response to actuator control command 510. In some embodiments, actuator 504 is configured to actuate computer-controlled machine 500 directly based on output signal y.
[0100] When actuator 504 receives actuator control command 510, actuator 504 is configured to perform an action corresponding to the associated actuator control command 510. Actuator 504 may include control logic configured to transform actuator control command 510 into a second actuator control command used to control actuator 504. In one or more embodiments, actuator control command 510 may be used to control a display instead of an actuator or other actuator-related device.
[0101] In some embodiments, instead of a computer-controlled machine 500, a sensor 506 may be included, or in addition, a control system 502 may include a sensor 506. Instead of a computer-controlled machine 500, an actuator 504 may be included, or in addition, a control system 502 may also include an actuator 504.
[0102] like Figure 5 As shown, the control system 502 also includes a processor 520 and a memory 522. The processor 520 may include one or more processors. The memory 522 may include one or more memory devices. A classifier 514 (e.g., an ML algorithm) of one or more embodiments may be implemented by the control system 502, which includes a non-volatile storage device 516, a processor 520, and a memory 522.
[0103] Non-volatile storage device 516 may include one or more persistent data storage devices, such as hard disk drives, optical disk drives, magnetic tape drives, non-volatile solid-state devices, cloud storage devices, or any other device capable of persistently storing information. Processor 520 may include one or more devices selected from a high-performance computing (HPC) system, including high-performance cores, microprocessors, microcontrollers, digital signal processors, microcomputers, central processing units, field-programmable gate arrays, programmable logic devices, state machines, logic circuits, analog circuits, digital circuits, or any other device that manipulates (analog or digital) signals based on computer-executable instructions residing in memory 522. Memory 522 may include a single memory device or multiple memory devices, including but not limited to random access memory (RAM), volatile memory, non-volatile memory, static random access memory (SRAM), dynamic random access memory (DRAM), flash memory, cache memory, or any other device capable of storing information.
[0104] Processor 520 may be configured to read memory 522 and execute computer-executable instructions residing in non-volatile storage device 516 and embodying one or more anomaly detection methodologies of one or more embodiments. Non-volatile storage device 516 may include one or more operating systems and applications. Non-volatile storage device 516 may store compilations and / or interpretations of computer programs created using various programming languages and / or techniques, including but not limited to Java, C, C++, C#, Objective C, Fortran, Pascal, JavaScript, Python, Perl, and PL / SQL, whether used individually or in combination.
[0105] When executed by processor 520, computer-executable instructions of non-volatile storage device 516 can cause control system 502 to implement one or more anomaly detection methodologies disclosed herein. Non-volatile storage device 516 may also include data supporting the functionality, features, and processes of one or more embodiments described herein.
[0106] Program code embodying the algorithms and / or methodologies described herein can be distributed individually or collectively as a program product in a wide variety of different forms. The program code can be distributed using a computer-readable storage medium having computer-readable program instructions thereon for causing a processor to implement aspects of one or more embodiments. Computer-readable storage media are inherently non-transitory and can include tangible media, implemented in any way or by any technique, that are volatile and non-volatile, and that are removable and non-removable, for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer-readable storage media can also include RAM, ROM, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other solid-state memory technologies, portable optical disc read-only memory (CD-ROM) or other optical storage devices, magnetic tape cassettes, magnetic tape, disk storage devices or other magnetic storage devices, or any other medium that can be used to store desired information and is readable by a computer. Computer-readable program instructions can be downloaded from the computer-readable storage medium to a computer, another type of programmable data processing device or other device, or via a network to an external computer or external storage device.
[0107] Computer-readable program instructions stored in a computer-readable medium can be used to direct a computer, other type of programmable data processing apparatus, or other device to operate in a particular manner, causing the instructions stored in the computer-readable medium (including instructions implementing the functions, actions, and / or operations specified in the flowchart or figure) to produce an article of writing. In some alternative embodiments, consistent with one or more embodiments, the functions, actions, and / or operations specified in the flowchart and figure can be reordered, processed sequentially, and / or processed simultaneously. Furthermore, any flowchart and / or figure may include more or fewer nodes or blocks than those illustrated consistent with one or more embodiments.
[0108] Processes, methods, or algorithms may be embodied, in whole or in part, using suitable hardware components such as application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), state machines, controllers, or other hardware components or devices, or a combination of hardware, software, and firmware components.
[0109] Figure 6A schematic diagram of a control system 502 is depicted, configured to control a vehicle 600, which may be at least partially autonomous or at least partially autonomous robot. In the example, the control system 502 is configured to control the vehicle 600 and / or perform various diagnostic techniques by executing an audio generation model according to the principles of this disclosure. The vehicle 600 includes actuators 504 and sensors 506. Sensors 506 may include one or more video sensors, cameras, radar sensors, ultrasonic sensors, LiDAR sensors, and / or position sensors (e.g., GPS). One or more of the specific sensors may be integrated into the vehicle 600. As an alternative to or supplement to the one or more specific sensors identified above, sensor 506 may include a software module configured to determine the state of actuator 504 upon execution. A non-limiting example of the software module includes a weather information software module configured to determine the current or future weather conditions near the vehicle 600 or other locations.
[0110] The classifier 514 of the control system 502 of vehicle 600 can be configured to detect objects near vehicle 600 based on the input signal x. In such an embodiment, the output signal y may include information characterizing the proximity of the object to vehicle 600. An actuator control command 510 can be determined based on this information. The actuator control command 510 can be used to avoid collisions with the detected objects.
[0111] In some embodiments, vehicle 600 is at least partially autonomous, and actuator 504 may be embodied in the vehicle 600's brakes, propulsion system, engine, transmission system, or steering mechanism. Actuator control command 510 can be determined to control actuator 504 so that vehicle 600 avoids collisions with detected objects. Classification can also be performed based on what the classifier 514 deems most likely to be (such as a pedestrian or a tree). Actuator control command 510 may be determined based on the classification. In the event of a potential adversarial attack, the system described above can also be trained to better detect objects or changes in lighting conditions or the angles of sensors or cameras on vehicle 600.
[0112] In some embodiments where the vehicle 600 is at least partially autonomous, the vehicle 600 may be a mobile robot configured to perform one or more functions, such as flying, swimming, diving, and stepping. The mobile robot may be at least partially autonomous lawnmower or at least partially autonomous cleaning robot. In such embodiments, actuator control commands 510 may be determined such that the propulsion unit, steering unit, and / or braking unit of the mobile robot can be controlled to prevent the mobile robot from colliding with identified objects.
[0113] In some embodiments, the vehicle 600 is at least a partially autonomous robot in the form of a gardening robot. In this embodiment, the vehicle 600 may use an optical sensor as sensor 506 to determine the state of plants in the environment near the vehicle 600. The actuator 504 may be a nozzle configured to spray chemicals. Depending on the identified plant species and / or the identified plant state, an actuator control command 510 may be determined to cause the actuator 504 to spray an appropriate amount of appropriate chemical onto the plant.
[0114] The vehicle 600 may be at least partially autonomous robots in the form of household appliances. Non-limiting examples of household appliances include washing machines, stoves, ovens, microwave ovens, or dishwashers. In such a vehicle 600, sensor 506 may be an optical sensor or an audio sensor configured to detect the state of the object the household appliance is about to handle. For example, in the case of a washing machine, sensor 506 may detect the state of the clothes inside the washing machine. Actuator control commands 510 may be determined based on the detected state of the clothes.
[0115] Figure 7 A schematic diagram of a control system 502 is depicted, which is configured to control a control system 700 (e.g., a manufacturing machine), such as a punching machine, cutting machine, or gun drill of a manufacturing system 702 (e.g., part of a production line). The control system 502 may be configured to control an actuator 504, which is also configured to control the control system 700 (e.g., the manufacturing machine). In the example, the control system 502 is configured to control the control system 700 and / or implement various diagnostic techniques by executing an audio generation model according to the principles of this disclosure.
[0116] Sensor 506 of system 700 (e.g., manufacturing machine) may be an audio sensor configured to capture one or more properties of the manufactured product 704. Classifier 514 may be configured to determine the state of the manufactured product 704 based on the captured one or more properties. Actuator 504 may be configured to control system 700 (e.g., manufacturing machine) based on the determined state of the manufactured product 704 in order to perform subsequent manufacturing steps of the manufactured product 704. Actuator 504 may be configured to control system 700 (e.g., manufacturing machine) to function on the subsequently manufactured product 706 of system 700 (e.g., manufacturing machine) based on the determined state of the manufactured product 704.
[0117] Figure 8A schematic diagram of a control system 502 is depicted, configured to control a power tool 800 (such as a drill or drive), having at least a partially autonomous mode. The control system 502 may be configured to control an actuator 504, which is also configured to control the power tool 800. In the example, the control system 502 is configured to control the power tool 800 and / or perform various diagnostic techniques by executing an audio generation model according to the principles of this disclosure.
[0118] The sensor 506 of the power tool 800 may be an audio sensor configured to capture one or more properties of the working surface 802 and / or the fastener 804 driven into the working surface 802. The classifier 514 may be configured to determine the state of the working surface 802 and / or the fastener 804 relative to the working surface 802 based on the captured one or more properties. This state may be that the fastener 804 is flush with the working surface 802. Alternatively, this state may be the hardness of the working surface 802. The actuator 504 may be configured to control the power tool 800 such that the drive function of the power tool 800 is adjusted depending on the determined state of the fastener 804 relative to the working surface 802 or the captured one or more properties of the working surface 802. For example, if the state of the fastener 804 is flush with the working surface 802, the actuator 504 may stop the drive function. As another non-limiting example, the actuator 504 may apply additional or less torque depending on the hardness of the working surface 802.
[0119] Figure 9 A schematic diagram of a control system 502 configured to control an automated personal assistant 900 is depicted. The control system 502 may be configured to control an actuator 504, which in turn is configured to control the automated personal assistant 900. The automated personal assistant 900 may be configured to control household appliances such as a washing machine, stove, oven, microwave oven, or dishwasher. In this example, the control system 502 is configured to control the automated personal assistant 900 and / or perform various diagnostic techniques by executing an audio generation model according to the principles of this disclosure.
[0120] Sensor 506 may be an optical sensor and / or an audio sensor. The optical sensor may be configured to receive video images of the gesture 904 of user 902. The audio sensor may be configured to receive voice commands from user 902.
[0121] The control system 502 of the automated personal assistant 900 can be configured to determine an actuator control command 510, which is configured to be received by the control system 502. The control system 502 can be configured to determine the actuator control command 510 based on a sensor signal 508 from a sensor 506. The automated personal assistant 900 is configured to transmit the sensor signal 508 to the control system 502. A classifier 514 of the control system 502 can be configured to execute a gesture recognition algorithm to identify a gesture 904 made by the user 902 in order to determine the actuator control command 510 and transmit the actuator control command 510 to the actuator 504. The classifier 514 can be configured to retrieve information from non-volatile storage in response to the gesture 904 and output the retrieved information in a form suitable for reception by the user 902.
[0122] Figure 10 A schematic diagram of a control system 502 is depicted, which is configured to control a monitoring system 1000. The monitoring system 1000 can be configured to physically control access through a door 1002. A sensor 506 can be configured to detect scenes related to determining whether to grant access. The sensor 506 can be an optical sensor configured to generate and transmit image and / or video data. The control system 502 can use such data to detect a person's face. In this example, the control system 502 is configured to control the monitoring system 1000 and / or implement various diagnostic techniques by executing an audio generation model according to the principles of this disclosure.
[0123] The classifier 514 of the control system 502 of the monitoring system 1000 can be configured to determine the identity of a person by interpreting image and / or video data by matching it with the identities of known persons stored in non-volatile storage device 516. The classifier 514 can be configured to generate actuator control commands 510 in response to the interpretation of the image and / or video data. The control system 502 is configured to transmit the actuator control commands 510 to actuator 504. In this embodiment, actuator 504 can be configured to lock or unlock door 1002 in response to actuator control commands 510. In some embodiments, non-physical logical access control is also possible.
[0124] The monitoring system 1000 can also be a surveillance system. In such an embodiment, sensor 506 may be an optical sensor configured to detect the monitored scene, and control system 502 is configured to control display 1004. Classifier 514 is configured to determine the classification of the scene, for example, whether the scene detected by sensor 506 is suspicious. Control system 502 is configured to transmit actuator control command 510 to display 1004 in response to classification. Display 1004 may be configured to adjust the displayed content in response to actuator control command 510. For example, display 1004 may highlight objects deemed suspicious by classifier 514. Using embodiments of the disclosed system, the surveillance system can predict the appearance of objects at some point in the future.
[0125] Figure 11 A schematic diagram of a control system 502 is depicted, configured to control an imaging system 1100, such as an MRI apparatus, an X-ray imaging apparatus, or an ultrasound apparatus. In this example, the control system 502 is configured to control the imaging system 1100 and / or perform various diagnostic techniques by executing an audio generation model according to the principles of this disclosure. The sensor 506 may be, for example, an imaging sensor and / or an audio sensor. A classifier 514 may be configured to determine a classification of all or part of the sensed image. The classifier 514 may be configured to determine or select an actuator control command 510 in response to a classification obtained by a trained neural network. For example, the classifier 514 may interpret a region of the sensed image as a potential abnormality. In this case, the actuator control command 510 may be determined or selected to cause the display 1102 to display the image and highlight the region of the potential abnormality.
[0126] While exemplary embodiments have been described above, these embodiments are not intended to describe all possible forms covered by the claims. The terms used in this specification are descriptive rather than restrictive, and it is to be understood that various changes may be made without departing from the spirit and scope of this disclosure. As described above, features of various embodiments may be combined to form further embodiments of the invention that may not be explicitly described or illustrated. While various embodiments may have been described as providing an advantage or superiority over other embodiments or prior art implementations with respect to one or more desired features, those skilled in the art will recognize that one or more features or characteristics may be waived to achieve desired overall system properties, depending on the particular application and implementation. These properties may include, but are not limited to, cost, strength, durability, lifecycle cost, merchantability, appearance, packaging, size, suitability, weight, manufacturability, ease of assembly, etc. Thus, any embodiment may be described to some extent as less desirable than other embodiments or prior art implementations with respect to one or more features, but such embodiments do not exceed the scope of this disclosure and may be desirable for a particular application.
Claims
1. A method for generating audio to obtain manipulated audio data, said manipulated audio data including one or more audio features indicating a fault, said method comprising at one or more processing devices: Receive a text description of the audio associated with the operation of the device; Receive audio data associated with the operation of the device; Based on the text description, a descriptive text input is generated that relates to the audio features of the operation of the device, wherein the descriptive text input includes at least one of the following: audio characteristics of a fault associated with the operation of the device, contextual information associated with the operation of the device, and conditions associated with the operation of the device. Based on the descriptive text input and the audio data, the manipulated audio data is generated, wherein the manipulated audio data includes one or more audio features indicating a fault associated with the descriptive text input; A machine learning (ML) model is trained to diagnose the fault using the manipulated audio data, wherein the ML model is trained to generate an output indicating the fault based on audio data obtained during the operation of the device. as well as Based on convergence during training, a trained ML model is output, which is configured to generate the output indicating the fault.
2. The method of claim 1, further comprising controlling one or more functions of the device based on the output.
3. The method of claim 2, wherein controlling the one or more functions includes at least one of the following: controlling or adjusting the operating parameters of the device, stopping the operation of the device, and generating an alarm.
4. The method of claim 1 further includes using a Large Language Model (LLM) to generate the text description.
5. The method of claim 4, further comprising receiving at the LLM a prompt from at least one of the following: (i) a knowledge base and (ii) one or more users.
6. The method of claim 5, wherein the prompt includes a description of audio characteristics associated with a malfunction in the operation of the device.
7. The method of claim 1, wherein training the ML model includes providing the ML model with observed audio data, the observed audio data including (i) healthy audio data, the healthy audio data excluding audio features indicating the fault, and (ii) fault audio data, the fault audio data including audio features indicating the fault.
8. A computing device configured to generate audio to obtain manipulated audio data, the manipulated audio data including one or more audio features indicating a fault, the computing device including a processing device configured to execute instructions stored in a memory to: Receive a text description of the audio associated with the operation of the device; Receive audio data associated with the operation of the device; Based on the text description, a descriptive text input is generated that relates to the audio features of the operation of the device, wherein the descriptive text input includes at least one of the following: audio characteristics of a fault associated with the operation of the device, contextual information associated with the operation of the device, and conditions associated with the operation of the device. Based on the descriptive text input and the audio data, the manipulated audio data is generated, wherein the manipulated audio data includes one or more audio features indicating a fault associated with the descriptive text input; A machine learning (ML) model is trained to diagnose the fault using the manipulated audio data, wherein the ML model is trained to generate an output indicating the fault based on audio data obtained during the operation of the device. as well as Based on convergence during training, a trained ML model is output, which is configured to generate the output indicating the fault.
9. The computing device of claim 8, wherein the processing device is further configured to execute the instructions to control one or more functions of the device based on the output.
10. The computing device of claim 9, wherein controlling the one or more functions includes at least one of the following: controlling or adjusting operating parameters of the device, stopping operation of the device, and generating an alarm.
11. The computing device of claim 8, wherein the processing device is further configured to execute the instructions to generate the text description using a Large Language Model (LLM).
12. The computing device of claim 11, wherein the processing device is further configured to execute the instructions to receive at the LLM a prompt from at least one of: (i) a knowledge base and (ii) one or more users.
13. The computing device of claim 12, wherein the prompt includes a description of audio characteristics associated with a malfunction in the operation of the device.
14. The computing device of claim 8, wherein training the ML model includes providing the ML model with observed audio data, the observed audio data including (i) healthy audio data, the healthy audio data excluding audio features indicating the fault, and (ii) fault audio data, the fault audio data including audio features indicating the fault.
15. A system configured to generate audio to obtain manipulated audio data, the manipulated audio data including one or more features indicating a fault corresponding to the operation of a computer-controlled machine, the system comprising: The control system is configured as follows Receive a text description of the audio associated with the operation of the computer-controlled machine. Receive audio data associated with the operation of the computer-controlled machine. Based on the text description, a descriptive text input is generated that associates audio features with the operation of the computer-controlled machine, wherein the descriptive text input includes at least one of the following: audio characteristics of malfunctions associated with the operation of the device, contextual information associated with the operation of the computer-controlled machine, and conditions associated with the operation of the computer-controlled machine. The manipulated audio data is generated based on the descriptive text input and observed audio data corresponding to the operation of the computer-controlled machine, wherein the manipulated audio data includes one or more audio features indicating a fault associated with the descriptive text input, and A machine learning (ML) model is trained to diagnose the fault using the manipulated audio data, wherein the ML model is trained to generate output indicating the fault based on audio data acquired during the operation of the computer-controlled machine. Based on the output, output control signals; as well as An actuator configured to control the operation of the computer-controlled machine based on the control signal.
16. The system of claim 15, wherein the operation of controlling the computer-controlled machine includes at least one of the following: (i) controlling or adjusting the operating parameters of the computer-controlled machine and (ii) stopping the operation of the computer-controlled machine.
17. The system of claim 15, wherein the control system is further configured to generate an alarm based on the output.
18. The system of claim 15 further includes a large language model (LLM) configured to generate the text description.
19. The system of claim 18, wherein the LLM is configured to generate the text description in response to a prompt received from at least one of: (i) a knowledge base and (ii) one or more users.
20. The system of claim 19, wherein the prompt includes a description of audio characteristics associated with a malfunction in the operation of the computer-controlled machine.