LLM-supported speech synthesis framework
The controllable speech synthesis framework addresses the limitations of AFMs by generating high-quality, spatially controlled speech-text pairs, enhancing AFM performance in complex audio tasks.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- ROBERT BOSCH GMBH
- Filing Date
- 2025-11-19
- Publication Date
- 2026-06-01
AI Technical Summary
Existing audio foundation models (AFMs) struggle with complex real-world audio environments due to the lack of large, high-quality datasets containing detailed, domain-specific speech information, and they often face challenges in modeling spatial information, localization, signal-to-noise ratio, and handling advanced inference tasks such as temporal causality and complex speech sequences.
A controllable speech synthesis framework is developed, using a large language model (LLM) to generate speech-text pairs with consistent audio quality and spatial attributes, enabling training of AFMs on curated datasets with logical sound combinations and descriptive captions, incorporating impulse response parameters to control spatial characteristics.
The framework enhances AFM capabilities for advanced inference tasks like speech caption generation, question answering, temporal reasoning, and acoustic scene prediction, improving the model's reliability and accuracy in understanding complex audio environments.
Smart Images

Figure 2026089691000001_ABST
Abstract
Description
Technical Field
[0001] Aspects of the present disclosure generally relate to large language model (LLM)-assisted speech synthesis frameworks.
Background Art
[0002] Background A speech generation model may be capable of generating, editing, or even transforming corresponding sound effects for a given language command. Furthermore, it has been shown that an audio foundation model (AFM) can convert audio content into natural language descriptions. Therefore, the AFM can be used for multimodal interactions in advanced audio applications.
Summary of the Invention
Means for Solving the Problems
[0003] Summary In one or more exemplary embodiments, a method for training an audio foundation model (AFM) to interpret digital audio signals is provided. By using a large language model (LLM) as a planner agent, a plurality of digital audio compositions are generated. The planner agent is prompted to generate a composition plan that defines a logical combination of digital foreground sound and digital background sound, event occurrences within the digital audio composition, and digital sound characteristics. The digital foreground sound and digital background sound have a consistent audio quality. An audio composition tool generates a plurality of digital audio compositions according to the composition plan. A summarizer agent is used to generate a description text for each of the plurality of digital audio compositions. The summarizer agent is implemented as an LLM that is prompted to describe the digital audio composition. A combination of the digital audio composition and the corresponding description text forms an audio-text pair. The AFM is trained to interpret digital audio signals using the audio-text pair.
[0004] In one or more exemplary embodiments, the method further includes preparing a set of audio sources by collecting audio clips using one or more speech generation models and / or source datasets, verifying the audio quality of the audio clips using an audio quality checker to ensure consistency based on objective metrics, and storing the verified audio clips in a foreground sound bank and a background sound bank.
[0005] In one or more exemplary embodiments, the method further includes introducing speech spatial characteristics to foreground and background sounds using impulse response (IR) parameters that define attributes including room dimensions, location of sound source, and distance of microphone; convolving the foreground and background sounds using the IR parameters to generate IR-tuned foreground and IR-tuned background sounds; and controlling the spatial characteristics of the foreground and background sounds and background sounds by storing the IR-tuned foreground sounds in a foreground sound bank and the IR-tuned background sounds in a background sound bank for use when generating multiple digital speech compositions.
[0006] In one or more exemplary embodiments, the IR parameters describe characteristics related to one or more of the following: acoustic reflection, energy absorption, and microphone array arrangement.
[0007] In one or more exemplary embodiments, the method further includes using a checker agent to verify that the descriptive text generated by the summarizer agent is consistent with the corresponding digital speech composition, and any mismatched speech-text pairs are flagged for review and / or regeneration, with the checker agent being implemented as an LLM that receives the descriptive text, a composition plan, and IR parameters as input.
[0008] The descriptive text includes question-answer pairs based on captions generated by a summarizer agent, and the AFM is trained for question-answering reasoning tasks using the audio-text pairs.
[0009] In one or more exemplary embodiments, the descriptive text includes a descriptive caption generated by a summarizer agent for describing a digital audio composition with respect to one or more of the following: an audio event, microphone position, sound propagation, signal characteristics, and background scenes, and the AFM is trained for the inference task of audio caption generation using audio-text pairs.
[0010] In one or more exemplary embodiments, the signal characteristics include one or more of the following: loudness level or signal-to-noise ratio (SNR).
[0011] In one or more exemplary embodiments, the descriptive text includes descriptive captions in chronological order, and the AFM is trained to predict subsequent acoustic scenes based on the current digital audio composition.
[0012] In one or more exemplary embodiments, a system for training a speech-based model (AFM) to interpret digital speech signals includes generating a plurality of digital speech compositions, the generation comprising using a large-scale language model (LLM) as a planner agent that is prompted to generate a composition plan that defines a logical combination of digital foreground and digital background sounds having consistent speech quality, event occurrences within the digital speech compositions, and digital sound characteristics; generating compositions according to the composition plan using a speech composition tool; generating descriptive text for each of the plurality of digital speech compositions using a summarizer agent implemented as an LLM that is prompted to describe the digital speech compositions; combining the digital speech compositions and the corresponding descriptive texts to form speech-text pairs; and training the AFM to interpret digital speech signals using the speech-text pairs, comprising one or more computing devices configured to do these things.
[0013] In one or more exemplary embodiments, one or more computing devices are further configured to prepare a set of audio sources by operations including collecting audio clips using one or more speech generation models and / or source datasets, verifying the audio quality of the audio clips using an audio quality checker to ensure consistency based on objective metrics, and storing the verified audio clips in foreground and background sound banks.
[0014] In one or more exemplary embodiments, one or more computing devices are further configured to control the spatial characteristics of foreground and background sounds and background sounds by operations including introducing speech spatial characteristics to foreground and background sounds using IR parameters that define attributes including room dimensions, location of sound sources and distance of microphones; convolving the foreground and background sounds using the IR parameters to generate IR-tuned foreground and IR-tuned background sounds; and storing the IR-tuned foreground sounds in a foreground sound bank and the IR-tuned background sounds in a background sound bank for use when generating multiple digital speech compositions.
[0015] In one or more exemplary embodiments, the IR parameters describe characteristics related to one or more of the following: acoustic reflection, energy absorption, and microphone array arrangement.
[0016] In one or more exemplary embodiments, one or more computing devices are further configured to use a checker agent to verify that the descriptive text generated by the summarizer agent matches the corresponding digital speech composition, and any mismatched speech-text pairs are flagged for review and / or regeneration, and the checker agent is implemented as an LLM that receives the descriptive text, the composition plan, and the IR parameters as input.
[0017] In one or more exemplary embodiments, the descriptive text includes question-answer pairs based on captions generated by a summarizer agent, and the AFM is trained for a question-answering reasoning task using the voice-text pairs.
[0018] In one or more exemplary embodiments, the descriptive text includes a descriptive caption generated by a summarizer agent for describing a digital audio composition with respect to one or more of the following: an audio event, microphone position, sound propagation, signal characteristics, and background scene, and the AFM is trained for the inference task of audio caption generation using audio-text pairs. In one or more exemplary embodiments, the signal characteristics include one or more of the following: loudness level or SNR.
[0019] In one or more exemplary embodiments, the descriptive text includes descriptive captions in chronological order, and the AFM is trained to predict subsequent acoustic scenes based on the current digital audio composition.
[0020] In one or more exemplary embodiments, the system further includes one or more voice sensors configured to capture digital voice from a manufacturing system, and one or more computing devices configured to provide the captured digital voice to the AFM in order to perform an inference task on the captured digital voice.
[0021] In one or more exemplary embodiments, a non-temporary computer-readable medium, when executed by one or more computing devices, generates a plurality of speech compositions, the generation of which includes using a Large Language Model (LLM) as a planner agent that is prompted to generate a composition plan defining a logical combination of digital foreground and digital background sounds having consistent speech quality, event occurrences within the digital speech compositions, and digital sound characteristics; using a speech composition tool to generate digital speech compositions according to the composition plan; using a summarizer agent implemented as an LLM that is prompted to describe the digital speech compositions to generate descriptive text for each of the plurality of digital speech compositions; combining the digital speech compositions and the corresponding descriptive texts to form speech-text pairs; and using the speech-text pairs to train the AFM to interpret the digital speech signals, the instructions cause one or more computing devices to perform such operations.
[0022] In one or more exemplary embodiments, a non-transient computer-readable medium further includes instructions that cause one or more computing devices to prepare a set of audio sources using operations that include, when executed by one or more computing devices, collecting audio clips using one or more speech generation models and / or source datasets; verifying the audio quality of the audio clips using an audio quality checker to ensure consistency based on objective metrics; and storing the verified audio clips in foreground and background sound banks.
[0023] In one or more exemplary embodiments, when executed by one or more computing devices, the non-transitory computer-readable medium uses IR parameters that define attributes including the dimensions of a room, the location of a sound source, and the distance to a microphone to introduce acoustic spatial characteristics into foreground sound and background sound, convolve the foreground sound and background sound using the IR parameters to generate IR-adjusted foreground sound and IR-adjusted background sound, and store the IR-adjusted foreground sound in a foreground sound bank and the IR-adjusted background sound in a background sound bank for use in generating a plurality of digital audio compositions, and further includes instructions to cause one or more computing devices to control the spatial characteristics of the foreground sound and background sound and the background sound using operations including these.
[0024] In one or more exemplary embodiments, the IR parameters describe characteristics related to one or more of acoustic reflection, energy absorption, and the arrangement of a microphone array.
[0025] In one or more exemplary embodiments, when executed by one or more computing devices, the non-transitory computer-readable medium further includes instructions to cause one or more computing devices to use a checker agent to verify that the descriptive text generated by a summarizer agent matches the corresponding digital audio composition, and non-matching audio-text pairs are flagged for review and / or regeneration, and the checker agent is implemented as an LLM that receives the descriptive text, a composition plan, and the IR parameters as inputs.
[0026] In one or more exemplary embodiments, the descriptive text includes question-answer pairs based on captions generated by a summarizer agent, and the AFM is trained for a question-and-answer inference task using the audio-text pairs.
[0027] The descriptive text includes descriptive captions generated by a summarizer agent for describing a digital audio composition with respect to one or more of audio events, microphone positions, acoustic propagation, signal characteristics, and background scenes. The AFM is trained for the inference task of audio caption generation using audio-text pairs. In one or more exemplary embodiments, the signal characteristics include one or more of loudness level or SNR.
[0028] In one or more exemplary embodiments, the descriptive text includes descriptive captions along the passage of time, and the AFM is trained to predict subsequent acoustic scenes based on the current digital audio composition.
Brief Description of the Drawings
[0029] [Figure 1] FIG. shows an exemplary process for utilizing an LLM-assisted speech synthesis framework for training and using AFM. [Figure 2] FIG. shows an exemplary portion of an LLM-assisted speech synthesis framework for preparing a sound source. [Figure 3] FIG. shows an exemplary portion of an LLM-assisted speech synthesis framework for controlling spatial characteristics. [Figure 4] FIG. shows an exemplary portion of an LLM-assisted speech synthesis framework for constructing a high-level audio composition. [Figure 5] FIG. shows an exemplary portion of an LLM-assisted speech synthesis framework for determining controllable language descriptors. [Figure 6] FIG. shows an exemplary portion of an LLM-assisted speech synthesis framework for performing model training to train AFM using audio-text pairs. [Figure 7] FIG. is a schematic diagram of the interaction between a computer-controlled machine and a control system. [Figure 8]This figure shows an exemplary manufacturing system that implements AFM for use in anomaly detection. [Modes for carrying out the invention]
[0030] Detailed explanation Detailed embodiments of the invention are disclosed herein as necessary, but it should be understood that the disclosed embodiments are merely illustrative examples of the invention, which may be realized in various alternative forms. The drawings are not necessarily to scale, and some features may be exaggerated or minimized to illustrate the details of certain components. Accordingly, certain structural and functional details disclosed herein should not be construed as limiting, but rather as merely representative grounds to teach those skilled in the art how to employ the invention in various ways.
[0031] AFM may be useful for a variety of inference tasks. For example, AFM could perform audio question-answering (AQA) by interpreting audio signals (e.g., digital audio signals) based on user queries (e.g., Question: "What sound events are present in the audio clip?", Answer: "A car passes by, people are talking"). However, AFM is typically only reliable for inference or basic semantic comprehension tasks such as generating a single sound event. Many AFMs struggle with the complexity of real-world audio environments.
[0032] For example, speech has unique physical properties such as spatial information (e.g., sound moving from left to right), localization (identifying the direction of the sound source), distance (distinguishing between foreground and background sounds), and signal-to-noise ratio (SNR). However, modeling these properties using existing AFMs can be challenging.
[0033] Beyond these basic speech-specific characteristics, many AFMs are insufficient to handle more advanced inference tasks. These inference tasks include understanding temporal causality (e.g., one event triggering another), counting events, and understanding complex structures in speech sequences.
[0034] A challenge in performing higher-level inference tasks using AFM is the lack of large, high-quality datasets containing detailed, domain-specific speech information. Most of the speech data currently available originates from public platforms that inherently lack control and knowledge regarding critical speech characteristics such as recording equipment and configuration specifications (distance between microphone and sound source). Furthermore, these natural environment recordings often come with inconsistent descriptions, tags, or captions that can be subjective. As a result, the uncontrollability and variability in both speech and language sources hinder progress in developing next-generation AFMs.
[0035] Aspects of this disclosure generally relate to a controllable speech synthesis framework capable of simulating real-world data for use in training an AFM. This simulated data may include speech-text pairs of compositions, each accompanied by descriptive text describing the composition. Compositions may be generated by foreground and background sounds, event occurrences, and logical combinations of sound attributes. Foreground and background sounds may be curated to have consistent quality and spatial attributes, thereby ensuring that the AFM is trained on relevant features of the data. A logical combination and descriptive text can be generated by leveraging a speech-language memory (LLM). Thus, the speech-language data synthesis framework combines low-level speech characteristic control with high-level compositional planning.
[0036] AFM can be trained for a variety of inference tasks using the generated datasets. These inference tasks may include speech caption generation and question answering, temporal inference and acoustic counting, and / or simulation of long-context scenarios and prediction of causal relationships. Further aspects of this disclosure are described in detail herein.
[0037] Figure 1 shows an exemplary process 100 for utilizing an LLM-assisted speech synthesis framework for training and using AFM112. As shown in the figure, process 100 includes preparing a speech source 102, controlling spatial characteristics 104, constructing a high-level speech composition 106, determining controllable language descriptors 108, training AFM112 110, and utilizing AFM112 for an inference task 114.
[0038] Figure 2 shows an exemplary portion 200 of an LLM-assisted speech synthesis framework for preparing a speech source 102. Preparing a speech source 102 may involve operations for collecting simple, short, and high-quality speech clips 202. As shown, portion 200 includes one or more speech generation models 204 and / or one or more source datasets 206. The speech generation models 204 and / or source datasets 206 are used as sources for the speech clips 202. A speech quality checker 208 verifies the aspects of the speech clips 202. The verified speech clips 202 are stored in the foreground sound bank 210 and the background sound bank 212.
[0039] Audio clip 202 may include a digital audio signal in the form of a computer-readable audio file, representing the clean audio of a single clear sound event. As described herein, a digital audio signal refers to a digital representation of a sound wave, created by sampling and quantizing an analog audio signal. The digital audio signal may be recorded using various sampling rates, quantization, optional compression, and encoding. Audio clip 202 may be diversified to function as basic sound elements for creating more complex compositions. These basic sound elements may be referred to as foreground sounds and can be stored in a foreground sound bank 210.
[0040] One source of the audio clip 202 may be a speech generation model 204. The speech generation model 204 may include various pre-trained models such as AudioBox, which are capable of generating outputs such as simple sound effects. These speech generation models 204 can be used to generate a variety of clean audio sources using direct prompt commands.
[0041] Other sources for audio clip 202 may be an existing clean audio source dataset 206. An example source dataset 206 is ESC50. Foreground sound bank 210 can be further expanded by incorporating these audio clips 202 from source dataset 206.
[0042] Preparing the audio sources 102 may also include saving a background sound bank 212. Background sound clips 202 that may require long-term continuous audio where high quality is not a major concern may originate from sounds of distant environments (e.g., urban parks, streets) or acoustic scenes (e.g., a home kitchen). These recordings may be obtained from a variety of sources, such as security cameras (private sources) or public sources such as video websites (e.g., YouTube).
[0043] The speech quality checker 208 may be run on speech clips 202 collected from the speech generation model 204 and the source dataset 206. The speech quality checker 208 ensures consistent quality in the foreground sound bank 210 and background sound bank 212. To do so, the speech quality checker 208 may be configured to calculate various objective metrics of the speech clips 202, such as SNR, to ensure that the speech clips 202 have matching speech quality. Speech clips 202 that do not meet the objective metrics can be discarded and not used. Thus, the result of preparing the speech sources 102 is a set of foreground and background sound clips with matching quality. This quality matching will be useful for downstream tasks. For example, if there are differences in quality, the AFM 112 trained on the speech clips 202 may use quality as a feature instead of learning based on the content of the speech clips 202.
[0044] In one example, the audio quality checker 208 measures the signal-to-noise ratio (SNR), e.g., minimum and / or maximum values, as the ratio of a desired signal to background noise. A higher SNR indicates clearer audio without excessive noise. For example, if a minimum SNR threshold is set, audio clips below this threshold can be flagged for exclusion or modification. Alternatively, if a maximum SNR threshold is set (for use in scenarios such as a street with background noise), audio clips with an SNR above this threshold can be flagged for exclusion or modification. In another example, the audio quality checker 208 measures the difference in dynamic range between the minimum and maximum volume portions of an audio clip 202. This can be compared to the minimum and / or maximum dynamic range. A consistent dynamic range ensures that the audio clip 202 is neither overcompressed nor excessively dynamic. In some examples, loudness normalization can be used to maintain a constant dynamic range. In further examples, the audio quality checker 208 can be used to measure harmonic distortion and ensure that the audio clip 202 falls within the minimum and / or maximum total harmonic distortion (THD) range. In yet another example, the audio quality checker 208 can be used to measure frequency components and ensure that the audio clip 202 falls within the minimum and / or maximum spectral range, or within the spectral balance range.
[0045] Figure 3 shows an exemplary portion 300 of an LLM-assisted speech synthesis framework for controlling spatial characteristics 104. As shown, the audio clips 202 from the foreground sound bank 210 and background sound bank 212 can be processed by the spatial characteristics application unit 302 to control the introduction of speech spatial characteristics into the audio clips 202 by using impulse response (IR) parameters 304. The result of the spatial characteristics application unit 302 may be an IR-adjusted foreground sound 306 and / or an IR-adjusted background sound 308.
[0046] The spatial characteristics assignment unit 302 can be used to introduce audio spatial characteristics. The introduction of audio spatial characteristics can be performed to simulate a desired IR for the audio clip 202. Exemplary IRs may include outdoor environments, reverberant rooms with hard surfaces, etc. IRs can be assigned to the audio clip 202 by manipulating the common audio spatial attributes of the audio clip 202.
[0047] The specific IR to be applied can be defined by the impulse response parameters 304. These impulse response parameters 304 can be provided as inputs to the spatial characteristic application unit 302. The specific attributes specified by the impulse response parameters 304 may include one or more of the following: room dimensions (indoors), reflection / propagation, energy absorption (transmission medium), location and direction of the sound source, microphone distance, and microphone array arrangement.
[0048] In one example, the controlled impulse response parameters 304 may be provided in a spatial configuration file (for example, as a file titled spatial_config.json in JSON (JavaScript Object Notation) format) to inform the spatial characterization unit 302 how to generate the corresponding IRs. The audio clip 202 is then convolved with these IRs to impart spatial sound effects. In one example, the introduction of audio spatial characteristics may be performed using the open-source Pyroomacoustics library. Pyroomacoustics is a package for audio signal processing for indoor applications that generates an artificial indoor impulse response between the sound source and the microphone. For the audio clip 202 from the foreground sound bank 210, the result of introducing audio spatial characteristics is an IR-adjusted foreground sound 306. For the audio clip 202 from the background sound bank 212, the result of introducing audio spatial characteristics is an IR-adjusted background sound 308.
[0049] Figure 4 shows an exemplary portion 400 of an LLM-assisted speech synthesis framework for constructing high-level speech compositions 106. As shown, IR-tuned foreground sounds 306 and IR-tuned background sounds 308 can be applied to a planner agent 404. Using these IR-tuned foreground sounds 306 and IR-tuned background sounds 308 and a composition prompt 402, the planner agent 404 determines a composition plan 406 for assembling speech clips 202 of the IR-tuned foreground sounds 306 and IR-tuned background sounds 308. These composition plans 406 can be applied to a speech composition tool 410 to generate a composition 412 of the IR-tuned foreground sounds 306 and IR-tuned background sounds 308.
[0050] Planner agent 404 may be one of several different LLMs such as GPT-4, Lama, Claude, etc. LLMs may be used for several different high-level reasoning tasks 114 and are considered to be powerful engines of common sense knowledge. Planner agent 404 may receive a composition prompt 402, which may include instructions that instruct planner agent 404 to determine a composition plan 406 for creating a composition 412. The composition plan 406 may define how the IR-tuned foreground sound 306 and the IR-tuned background sound 308 are combined by logical methods and / or using event class labels from an existing corpus.
[0051] The composition parameter 408 can specify other metadata information describing the desired composition 412. Some non-limiting examples of the composition parameter 408 may include the SNR adjustment (e.g., in dB) of the audio clips 202 to be combined, the frequency of events (e.g., the number of occurrences of a particular event, e.g., the number of times a particular audio clip 202 is played), the foreground / background sound combination scheme, and the desired result time span of each sound event and / or the overall composition 412 (e.g., duration, length, rate of the event). These parameters can be extracted and enumerated as a JSON synthesis configuration file (e.g., synth_config.json) to instruct the audio composition tool 410 to generate the final audio composition 412 according to given instructions.
[0052] The planner agent 404 may be used to populate composition parameters 408 so as to compose plausible combinations of speech events for realistic data curation, using common sense knowledge. For example, unlike conventional synthesis frameworks that may uncontrollably and randomly link audio clips 202, the planner agent 404 can prevent unrealistic combinations, such as associating a background urban park scenario with a foreground microwave event.
[0053] Using the audio composition tool 410, a composition 412 can be synthesized from an audio clip 202 and various composition parameters 408. In one example, the SCAPER package can be used as the audio composition tool 410.
[0054] Figure 5 shows an exemplary portion 500 of an LLM-assisted speech synthesis framework for determining a controllable language descriptor 108. With respect to determining a controllable language descriptor 108, various source recipes 502 of information are supplied to a summarizer agent 504 configured to generate a description text 506 that describes a composition 412. A summarizer prompt 508 can be used to instruct the summarizer agent 504 regarding the type of description text 506 to be generated. A description text 506 can be generated for each of the compositions 412, thereby allowing the combinations of the description text 506 and the composition 412 to be organized together as a speech-text pair 510. To minimize the possibility of hallucination in the description text 506, an additional checker agent 512 can be used to cross-verify that the description text 506 is content-consistent with a given source recipe 502.
[0055] The source recipe 502 may contain various pieces of information, such as a composition plan 406 determined by the planner agent 404 (for example, which foreground events in the audio clips 202 of the IR-tuned foreground sound 306 are associated with which background audio clips 202 of the IR-tuned background sound 308), pre-configured controlled impulse response parameters 304, and composition parameters 408.
[0056] The summarizer agent 504, like the planner agent 404, may be one of several different LLMs. In some cases, the summarizer agent 504 may be the same LLM as the planner agent 404, but using a different prompt (e.g., summarizer prompt 508), while in other cases, the summarizer agent 504 may be a different LLM. Using the complete source recipe 502 corresponding to the synthesized speech composition 412, the summarizer agent 504 can generate natural language description text 506 that accurately reflects the acoustic characteristics of each synthesized speech composition 412.
[0057] The summarizer prompt 508 may include instructions to the summarizer agent 504 regarding the type of descriptive text 506 to be generated. In one example, the summarizer agent 508 may instruct the summarizer agent 504 to generate a caption for composition 412. In another example, the summarizer prompt 508 may instruct the summarizer agent 504 to generate a question-and-answer pair about composition 412. A descriptive text 506 can be generated for each of composition 412, so that the combination of the descriptive text 506 and the composition 412 described by the descriptive text 506 is combined as an audio-text pair 510. Thus, the audio-text pairs 510 can serve as a training dataset for training 110 of AFM 112.
[0058] The checker agent 512, like the planner agent 404 and the summarizer agent 504, may be one of several different LLMs. In some cases, the checker agent 512 may be the same LLM as the planner agent 404 or the checker agent 512 (using other different prompts, depending on the case), and in other cases, the checker agent 512 may be a different LLM. In some cases, it is preferable that they be different LLMs so that potential defects in the summarizer agent 504 can be addressed by the checker agent 512.
[0059] If the checker agent 512 determines that the descriptive text 506 does not describe each of the compositions 412, the checker agent 512 may flag the potential audio-text pair 510 for review, instruct it to regenerate the compositions 412 and / or the descriptive text 506, and / or prevent the potential audio-text pair 510 from being included in the audio-text pair 510.
[0060] Figure 6 shows an exemplary portion 600 of an LLM-assisted speech synthesis framework for performing model training 602 to train AFM 112 using speech-text pairs 510. For example, a summarizer agent 504 may describe a speech composition in terms of one or more of the following: sound events, microphone locations, sound propagation, signal characteristics, and background scenes, and AFM 112 may be trained to interpret digital speech signals for a speech caption generation inference task using the speech-text pairs 510. In one or more exemplary embodiments, signal characteristics include one or more of the following: loudness level or signal-to-noise ratio (SNR).
[0061] In one example, model training 602 may include training 110 for speech caption generation and question answering. In such an example, the framework can be used to generate speech-text pairs 510, each containing a set of captions and question-answers for each of the compositions 412. These informational elements would be useful for facilitating the training 110 of advanced AFMs 112, such as contrastive language-audio pretraining (CLAP) and AQA models.
[0062] The following example caption is an example of descriptive text 506 from summarizer agent 504. This caption shows that planner agent 404 picks up an utterance as a foreground event while in an office background environment with the sounds of multiple dial telephones, which correctly follows the common-sense theory of sound composition 412: [Table 1]
[0063] The detailed speech characteristics controlled by the controlled impulse response parameter 304 and the composition parameter 408 can also be precisely incorporated into the caption description text 506, for example, a far-end microphone suggests the microphone's location, a noisy background indicates that the recording is a simulated low SNR recording, and four times reveals the number of telephone event audio clips 202 in composition 412.
[0064] Furthermore, AQA pairs can be curated by prompting the LLM based on the caption description text 506. For example: [Table 2]
[0065] By using these question / answer descriptive texts 506 and compositions 412 as audio-text pairs 510, model training 602 can be performed to instruct the AFM 112 to perform question-answering based on audio files. Such a model can receive audio clips 202 and questions about audio clips 202 in inference mode and generate answers to the questions.
[0066] In other embodiments, model training 602 may include training 110 for interpreting digital audio signals for temporal reasoning and acoustic counting. Here again, by using the framework, compositions 412 containing soundscapes can be curated that conform to a desired event sequence, timing, and frequency or number of occurrences, thereby increasing the complexity of the AQA and audio caption generation tasks. For example, the following captions can generate complex questions incorporating the concept of time: [Table 3]
[0067] This caption can be used to curate AQA pairs by prompting the LLM based on the caption description text 506. For example: [Table 4]
[0068] Therefore, by using the descriptive text 506 and composition 412 of these questions / answers as audio-text pairs 510 to perform model training 602 of AFM112, AFM112 will be able to tackle more complex tasks such as temporal reasoning or understanding acoustic scenes.
[0069] In further embodiments, model training 602 may include training 110 for interpreting digital audio signals for simulating long-context scenarios and predicting causal relationships. In such applications of the framework, the inference capability of AFM112 is to predict future causal scenarios based on an understanding of the current acoustic environment. Similarly, the synthetic framework can be used to curate long-context acoustic scenarios that conform to desired causal relationships. Based on the above example, the causal relationships in the following scene may be as follows: [Table 5]
[0070] Therefore, synthesis based on causal instructions can be incorporated into the subsequent acoustic scene prediction task, enabling the AFM112 to demonstrate its capabilities.
[0071] Figure 7 shows a schematic diagram of the interaction between the computer-controlled machine 702 and the control system 712. The computer-controlled machine 702 can implement embodiments of the trained AFM 112 as described herein, using a framework for interpreting digital voice signals to perform various inference tasks 114.
[0072] Referring to Figures 7 and 1 to 6, the approach described herein can be implemented in connection with such a computer-controlled machine 702 and control system 712. The computer-controlled machine 702 includes an actuator 714 and a sensor 716. The actuator 714 may include one or more actuators, and the sensor 716 may include one or more sensors. The sensor 716 is configured to detect the state of the computer-controlled machine 702. The sensor 716 may be configured to encode the detected state into a sensor signal 718 and transmit the sensor signal 718 to the control system 712. Non-limiting examples of the sensor 716 include a microphone, an accelerometer, and the like. In one embodiment, the sensor 716 is a voice sensor configured to detect audio data of the environment near the computer-controlled machine 702.
[0073] The control system 712 is configured to receive a sensor signal 718 from the computer-controlled machine 702. As described below, the control system 712 may be further configured to calculate an actuator control command 720 depending on the sensor signal 718 and to transmit the actuator control command 720 to the actuator 714 of the computer-controlled machine 702.
[0074] As shown in Figure 7, the control system 712 includes a receiving unit 722. The receiving unit 722 may be configured to receive a sensor signal 718 from the sensor 716 and convert the sensor signal 718 into an input signal X. In an alternative embodiment, the sensor signal 718 is received directly as an input signal X without using the receiving unit 722. Each input signal x may be a part of each sensor signal 718. The receiving unit 722 may be configured to process each sensor signal 718 to generate each input signal X. The input signal X may include data corresponding to sound recorded by the sensor 716.
[0075] The control system 712 includes a machine learning (ML) processing unit 724. The ML processing unit 724 may be configured to perform learning, classification, inference, generation, etc., using one or more models as detailed above. In one example, the ML processing unit 724 is configured to determine an output signal Y from an input signal X. Each output signal Y contains information to assign one or more labels to each input signal X. The ML processing unit 724 can transmit the output signal Y to a conversion unit 728. The conversion unit 728 is configured to convert the output signal Y into an actuator control command 720. The control system 712 is configured to transmit the actuator control command 720 to an actuator 714, which is configured to operate a computer-controlled machine 702 in response to the actuator control command 720. In another embodiment, the actuator 714 is configured to operate a computer-controlled machine 702 directly based on the output signal Y.
[0076] When actuator 714 receives an actuator control command 720, actuator 714 is configured to perform an action corresponding to the associated actuator control command 720. Actuator 714 may include control logic configured to translate actuator control command 720 into a second actuator control command 720 used to control actuator 714. In one or more embodiments, the actuator control command 720 can be used to control a display instead of or in addition to actuator 714.
[0077] In other embodiments, the control system 712 includes a sensor 716 in place of or in addition to a computer-controlled machine 702 which includes a sensor 716. The control system 712 may also include an actuator 714 in place of or in addition to a computer-controlled machine 702 which includes an actuator 714.
[0078] As shown in Figure 7, the control system 712 also includes a processor 730 and a memory 732. The processor 730 may include one or more processors. The memory 732 may include one or more memory devices.
[0079] Non-volatile storage 726 may include one or more persistent data storage devices, such as hard drives, optical drives, tape drives, non-volatile solid-state devices, cloud storage, or any other devices capable of permanently storing information. Processor 730 may include one or more devices selected from the following, which process (analog or digital) signals based on computer executable instructions residing in memory 732: namely, high-performance computing (HPC) systems including high-performance cores, microprocessors, microcontrollers, digital signal processors, microcomputers, central processing units (CPUs), field-programmable gate arrays (FPGAs), programmable logic devices, state machines, logic circuits, analog circuits, digital circuits, or any other devices. Memory 732 may include one or more memory devices, but are not limited to, random-access memory (RAM), volatile memory, non-volatile memory, static random-access memory (SRAM), dynamic random-access memory (DRAM), flash memory, cache memory, or any other devices capable of storing information.
[0080] The processor 730 may reside in the non-volatile storage 726 and be configured to load and execute computer executable instructions that embody one or more ML algorithms and / or methodologies of one or more embodiments into memory 732. The non-volatile storage 726 may include one or more operating systems and applications. The non-volatile storage 726 can store compiled and / or interpreted computer programs written using a variety of programming languages and / or technologies, including, but not limited to, Java, C, C++, C#, Objective C, Fortran, Pascal, JavaScript, Python, Perl, and PL / SQL, either alone or in combination.
[0081] When executed by the processor 730, the computer-executable instructions of the non-volatile storage 726 can cause the control system 712 to implement one or more of the ML algorithms and / or methodologies disclosed herein. The non-volatile storage 726 may also contain ML data (including data parameters) that support the functions, features and processes of one or more embodiments described herein.
[0082] Program code embodying the algorithms and / or methodologies described herein may be distributed individually or collectively as various different forms of program products. The program code may be distributed using computer-readable storage media containing computer-readable program instructions for causing a processor to execute one or more embodiments of the program code. Computer-readable storage media that are essentially non-transient may include volatile and non-volatile removable and non-removable tangible media implemented in any way or technique for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer-readable storage media may further include RAM, ROM, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other solid-state memory technologies, portable compact disk read-only memory (CD-ROM), or other optical storage, magnetic cassettes, magnetic tapes, magnetic disk storage, or other magnetic storage devices, or any other media that can be used to store and read information of a desired information. Computer-readable program instructions may be downloaded from a computer-readable storage medium to a computer, another type of programmable data processing device, or another device, or they may be downloaded via a network to an external computer or external storage device.
[0083] Computer-readable program instructions stored on a computer-readable medium may be used to cause a computer, other types of programmable data processing devices, or other devices to function in a particular manner so that the instructions stored on the computer-readable medium produce a product containing instructions that perform functions, behaviors, and / or actions specified in a flowchart or diagram. In certain alternative embodiments, multiple functions, behaviors, and / or actions specified in a flowchart and diagram can be rearranged, processed serially, and / or processed simultaneously according to one or more embodiments. Furthermore, both flowcharts and / or diagrams may contain more or fewer nodes or blocks than those illustrated according to one or more embodiments.
[0084] The process, method, or algorithm may be implemented, in whole or in part, using appropriate hardware components such as application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), state machines, controllers, or other hardware components or hardware devices, or combinations of hardware components, software components, and firmware components.
[0085] Figure 8 shows an exemplary manufacturing system 800 that implements the AFM112 for use in anomaly detection. The system 800 may be configured to control manufacturing machinery 802, such as a punch cutter, cutter, or gun drill, which may be part of a production line.
[0086] System 800 may be configured to control an actuator 714 configured to control a manufacturing machine 802. Sensor 716 of System 800 may be configured to capture one or more characteristics of a manufactured product 804. ML processing unit 724 may be configured to determine the state of the manufactured product 804 from one or more captured characteristics. Actuator 714 may be configured to control System 800 (e.g., manufacturing machine 802) depending on the determined state of the manufactured product 804 for subsequent manufacturing steps of the manufactured product 804. In particular, actuator 714 may be configured to control the function of System 800 (e.g., manufacturing machine 802) with respect to subsequent manufactured products 806 depending on the determined state of the manufactured product 804.
[0087] For example, system 800 can use the AFM 112, trained using the framework as described herein, to explain the reasons for potential problems in the manufacturing system 800. This may be done based on abnormal sounds collected from sensor 716. In another example, system 800 can use the AFM 112 to predict the next expected outcome to be addressed, based on sounds collected from sensor 716, particularly when the next action may be related to a manufacturing problem. In yet another example, the AFM 112 can be used to respond to user questions regarding sounds captured by sensor 716.
[0088] The processes, methods, or algorithms disclosed herein may be deliverable to or implemented by an processing unit, controller, or computer, and the processing unit, controller, or computer may include any existing programmable electronic control unit or a dedicated electronic control unit. Similarly, the processes, methods, or algorithms may be stored as data and instructions executable by a controller or computer in various forms, including, but not limited to, information permanently stored in a non-writable storage medium such as a read-only memory (ROM) device and information variably stored in a writable storage medium such as a floppy disk, magnetic tape, compact disk (CD), RAM device, and other magnetic and optical media. The processes, methods, or algorithms may be implemented in a software-executable object. Alternatively, the processes, methods, or algorithms may be embodied in whole or in part using appropriate hardware components, such as an ASIC, FPGA, state machine, controller, or other hardware component or hardware device, or a combination of hardware components, software components, and firmware components.
[0089] While exemplary embodiments have been described above, these embodiments are not intended to describe all possible forms that are covered by the claims. The terms used herein are descriptive, not limiting, and it should be understood that various modifications are possible without departing from the spirit and scope of this disclosure. As stated above, features of various embodiments can be combined to form further embodiments of the invention not expressly described or illustrated. While various embodiments may be described as offering advantages over or being more preferable to other embodiments or implementations of the prior art with respect to one or more desired characteristics, those skilled in the art will recognize that one or more features or characteristics may be compromised to achieve desired overall system attributes that depend on the particular application and implementation. These attributes may include, but are not limited to, strength, durability, lifespan, marketability, appearance, packaging, dimensions, maintainability, weight, manufacturability, ease of assembly, etc. Therefore, to the extent that any embodiment is described as being less desirable than other embodiments or implementations of the prior art with respect to one or more characteristics, these embodiments are not outside the scope of this disclosure and may even be desirable for a particular application.
Claims
1. A method for training an audio-based model (AFM) to interpret digital audio signals, The aforementioned method, The process of generating multiple digital audio compositions, the process of generating them Using a summarizer agent implemented as an LLM that is prompted to describe the digital speech composition, a descriptive text is generated for each of the plurality of digital speech compositions. The digital audio composition and the corresponding descriptive text are combined to form an audio-text pair, The AFM is trained to interpret digital audio signals using the aforementioned voice-text pairs, A method that includes this.
2. The aforementioned method, Collecting audio clips using one or more speech generation models and / or source datasets, To ensure consistency based on objective metrics, the audio quality of the audio clips is verified using an audio quality checker, The verified audio clips are stored in the foreground sound bank and the background sound bank. This involves preparing a set of audio sources. The method according to claim 1, further comprising:
3. The aforementioned method, The method involves introducing spatial speech characteristics to the foreground and background sounds using impulse response (IR) parameters that define attributes including room dimensions, sound source location, and microphone distance. The foreground sound and background sound are convolved using the aforementioned IR parameters to generate IR-adjusted foreground sound and IR-adjusted background sound. In order to use when generating the aforementioned multiple digital audio compositions, the IR-adjusted foreground sound is stored in a foreground sound bank, and the IR-adjusted background sound is stored in a background sound bank. This allows for the control of the spatial characteristics of the foreground sound and the background sound, as well as the background sound. The method according to claim 1, further comprising:
4. The aforementioned IR parameters describe characteristics related to one or more of the following: acoustic reflection, energy absorption, and microphone array arrangement. The method according to claim 3.
5. The aforementioned method, The checker agent is used to verify that the descriptive text generated by the summarizer agent is consistent with the corresponding digital audio composition. It further includes, Incompatible audio-text pairs are flagged for review and / or regeneration. The checker agent is implemented as an LLM that receives the descriptive text, the composition plan, and the IR parameters as input. The method according to claim 3.
6. The descriptive text includes question-answer pairs based on captions generated by the summarizer agent, The AFM is trained for question-answering reasoning tasks using the voice-text pair. The method according to claim 1.
7. The descriptive text includes a descriptive caption generated by the summarizer agent for describing the digital audio composition with respect to one or more of the following: sound events, microphone positions, sound propagation, signal characteristics, and background scenes. The AFM is trained for the inference task of generating speech captions using the speech-text pair. The method according to claim 1.
8. The aforementioned signal characteristics include one or more of the following: loudness level or signal-to-noise ratio (SNR). The method according to claim 7.
9. The aforementioned descriptive text includes descriptive captions that chronologically describe the passage of time. The AFM is trained to predict subsequent acoustic scenes based on the current digital audio composition. The method according to claim 1.
10. A system for training an audio-based model (AFM) to interpret digital audio signals, The aforementioned system, The process of generating multiple digital audio compositions, the process of generating them Using a summarizer agent implemented as an LLM that is prompted to describe the digital speech composition, a descriptive text is generated for each of the plurality of digital speech compositions. The digital audio composition and the corresponding descriptive text are combined to form an audio-text pair, To train the AFM to interpret digital audio signals using the aforementioned voice-text pairs and One or more computing devices configured to perform A system equipped with these features.
11. The one or more computing devices mentioned above are: Collecting audio clips using one or more speech generation models and / or source datasets, To ensure consistency based on objective metrics, the audio quality of the audio clips is verified using an audio quality checker, The verified audio clips are stored in the foreground sound bank and the background sound bank. The operation, which includes this, is further configured to prepare a set of audio sources. The system according to claim 10.
12. The one or more computing devices mentioned above are: Introducing spatial audio characteristics to the foreground and background sounds using IR parameters that define attributes including room dimensions, sound source location, and microphone distance, The foreground sound and background sound are convolved using the aforementioned IR parameters to generate IR-adjusted foreground sound and IR-adjusted background sound. In order to use when generating the aforementioned multiple digital audio compositions, the IR-adjusted foreground sound is stored in a foreground sound bank, and the IR-adjusted background sound is stored in a background sound bank. The operation, which includes, is further configured to control the spatial characteristics of the foreground sound and the background sound and the background sound, The system according to claim 10.
13. The aforementioned IR parameters describe characteristics related to one or more of the following: acoustic reflection, energy absorption, and microphone array arrangement. The system according to claim 12.
14. The one or more computing devices mentioned above are: The checker agent is further configured to verify that the descriptive text generated by the summarizer agent is consistent with the corresponding digital speech composition, Incompatible audio-text pairs are flagged for review and / or regeneration. The checker agent is implemented as an LLM that receives the descriptive text, the composition plan, and the IR parameters as input. The system according to claim 12.
15. The descriptive text includes question-answer pairs based on captions generated by the summarizer agent, The AFM is trained for question-answering reasoning tasks using the voice-text pair. The system according to claim 10.
16. The descriptive text includes a descriptive caption generated by the summarizer agent for describing the digital audio composition with respect to one or more of the following: sound events, microphone positions, sound propagation, signal characteristics, and background scenes. The AFM is trained for the inference task of generating speech captions using the speech-text pair. The system according to claim 10.
17. The aforementioned signal characteristics include one or more of the loudness level or SNR. The system according to claim 16.
18. The aforementioned descriptive text includes descriptive captions that chronologically describe the passage of time. The AFM is trained to predict subsequent acoustic scenes based on the current digital audio composition. The system according to claim 10.
19. The aforementioned system, One or more audio sensors configured to capture digital audio from a manufacturing system. Furthermore, The one or more computing devices mentioned above are: The system is configured to provide the captured digital audio to the AFM in order to perform an inference task on the captured digital audio. The system according to claim 10.
20. A non-temporary computer-readable medium, The non-temporary computer-readable medium, when executed by one or more computing devices to train a speech-based model (AFM) to interpret digital speech signals, The process of generating multiple digital audio compositions, the process of generating them Using a summarizer agent implemented as an LLM that is prompted to describe the digital speech composition, a descriptive text is generated for each of the plurality of digital speech compositions. The composition and the corresponding descriptive text are combined to form an audio-text pair, To train the AFM to interpret digital audio signals using the aforementioned voice-text pairs and A non-temporary computer-readable medium comprising instructions for causing one or more computing devices to perform an operation that includes the above.
21. The non-temporary computer-readable medium, when executed by one or more computing devices, Collecting audio clips using one or more speech generation models and / or source datasets, To ensure consistency based on objective metrics, the audio quality of the audio clips is verified using an audio quality checker, The verified audio clips are stored in the foreground sound bank and the background sound bank. The instructions further include causing one or more computing devices to perform an operation that includes preparing a set of audio sources. The non-temporary computer-readable medium according to claim 20.
22. The non-temporary computer-readable medium, when executed by one or more computing devices, Introducing spatial audio characteristics to the foreground and background sounds using IR parameters that define attributes including room dimensions, sound source location, and microphone distance, The foreground sound and background sound are convolved using the aforementioned IR parameters to generate IR-adjusted foreground sound and IR-adjusted background sound. In order to use when generating the aforementioned multiple digital audio compositions, the IR-adjusted foreground sound is stored in a foreground sound bank, and the IR-adjusted background sound is stored in a background sound bank. The system further comprises instructions for causing one or more computing devices to perform actions that include controlling the spatial characteristics of the foreground sound and the background sound and the background sound, using actions that include the following: The non-temporary computer-readable medium according to claim 20.
23. The aforementioned IR parameters describe characteristics related to one or more of the following: acoustic reflection, energy absorption, and microphone array arrangement. The non-temporary computer-readable medium according to claim 22.
24. The non-temporary computer-readable medium further comprises instructions to cause one or more computing devices to use a checker agent to verify that the descriptive text generated by the summarizer agent is consistent with the corresponding digital speech composition, when executed by one or more computing devices. Incompatible audio-text pairs are flagged for review and / or regeneration. The checker agent is implemented as an LLM that receives the descriptive text, the composition plan, and the IR parameters as input. The non-temporary computer-readable medium according to claim 22.
25. The descriptive text includes question-answer pairs based on captions generated by the summarizer agent, The AFM is trained for question-answering reasoning tasks using the voice-text pair. The non-temporary computer-readable medium according to claim 20.
26. The descriptive text includes a descriptive caption generated by the summarizer agent for describing the digital audio composition with respect to one or more of the following: sound events, microphone positions, sound propagation, signal characteristics, and background scenes. The AFM is trained for the inference task of generating speech captions using the speech-text pair. The non-temporary computer-readable medium according to claim 20.
27. The aforementioned signal characteristics include one or more of the loudness level or SNR. The non-temporary computer-readable medium according to claim 26.
28. The aforementioned descriptive text includes descriptive captions that chronologically describe the passage of time. The AFM is trained to predict subsequent acoustic scenes based on the current digital audio composition. The non-temporary computer-readable medium according to claim 20.