System and method for generating metadata for an audio signal

By jointly training a neural network and combining a transformer model and a CTC model, an encoder-decoder architecture with shared parameters processes audio signals, solving the problem of unifying ASR, AED, and AT tasks in audio processing and achieving high efficiency and accuracy in multi-event recognition and labeling.

CN116324984BActive Publication Date: 2025-12-09MITSUBISHI ELECTRIC CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202180067206.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-10-07
Filing Date
2021-04-27
Publication Date
2025-12-09
Estimated Expiration
2041-04-27

AI Technical Summary

Technical Problem

Existing audio processing technologies cannot effectively unify the execution of automatic speech recognition (ASR), acoustic event detection (AED), and audio tagging (AT) tasks, and they also suffer from the problem of sparse training data.

Method used

By employing a jointly trained neural network, combining a transformer model and a connection time classification (CTC) model, and sharing some parameters, the collaborative execution of ASR, AED, and AT is achieved. Audio signals are processed through an encoder-decoder architecture to generate metadata.

Benefits of technology

It achieves accurate identification and labeling of multiple audio events in different audio scenarios, reduces the training data requirements, and improves the accuracy and efficiency of transcription tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116324984B_ABST
    Figure CN116324984B_ABST
Patent Text Reader

Abstract

An audio processing system is provided. The audio processing system includes an input interface configured to accept an audio signal. Further, the audio processing system includes a memory configured to store a neural network trained to determine attributes of different types for a plurality of concurrent audio events of different origins, where the types of attributes include temporal-dependent attributes and temporal-agnostic attributes for speech audio events and non-speech audio events. Further, the audio processing system includes a processor configured to process the audio signal with the neural network to generate metadata for the audio signal, the metadata including one or more attributes for one or more audio events in the audio signal.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates generally to audio processing, and more particularly, to a system for generating metadata for audio signals using neural networks. BACKGROUND

[0002] Speech recognition systems have progressed to the point where people can rely on speech to interact with computing devices. These systems employ techniques that identify words spoken by a user based on various parameters of received audio input. Speech recognition combined with natural language understanding processing techniques enable voice-based user control of computing devices to perform tasks based on a user's spoken commands. The combination of speech recognition and natural language understanding processing techniques is often referred to as speech processing. Speech processing can also convert a user's speech to text data, which can then be provided to various text-based software applications. The conversion of audio data associated with speech to text representing that speech is referred to as automatic speech recognition (ASR).

[0003] Further, acoustic event detection (AED) techniques can be used to detect certain sound events, such as regular household sounds (door closing, sink running, etc.), speech sounds (but not speech transcription), mechanical sounds, or other sound events, along with corresponding timing information, such that each sound event is associated with a start time and an end time. For example, in an automotive repair shop, AED can be configured to detect the sound of a drill from audio input, along with a corresponding start time and end time of the drill sound. Additionally, audio tagging (AT) techniques can be used to detect the presence of a sound event (e.g., identifying the event as tagged "drill") without timing, such that a start time and end time are not detected. Additionally or alternatively, AT can include audio captioning, in which natural language sentences describing the acoustic scene are generated. For example, in an automotive repair shop, audio captions such as "a person is speaking while operating a drill" can be generated.

[0004] However, audio tagging (AT), acoustic event detection (AED), and automatic speech recognition (ASR) are treated as separate problems. Additionally, task-specific neural network architectures are used to perform each of the ASR, AED, and AT tasks. Some approaches use an attention-based encoder-decoder neural network architecture, where an encoder extracts acoustic cues, an attention mechanism acts as a relay, and a decoder performs perception, detection, and recognition of audio events. However, for event classification, the use of an encoder-decoder neural network architecture is limited to a non-attention-based recurrent neural network (RNN) solution, where an encoder compresses an acoustic signal into a single embedding vector, and a decoder detects audio events encoded in such a vector representation.

[0005] Accordingly, there is a need for a system and method for unifying ASR, AED, and AT. SUMMARY

[0006] It is an object of some embodiments to enable synergy in performing different transcription tasks on audio signals of an audio scene by jointly training a neural network for different transcription tasks. Alternatively, it is an object of some embodiments to provide a system configured to use a neural network to perform different transcription tasks such as automatic speech recognition (ASR), acoustic event detection (AED), and audio tagging (AT) to generate metadata of an audio signal. The metadata includes attributes of different types of multiple concurrent audio events in the audio signal. According to some embodiments, the neural network includes a transformer model and a connectionist temporal classification (CTC) based model, and can be trained to perform ASR, AED, and AT transcription tasks on the audio signal. Additionally, it is an object of some embodiments to train the transformer model jointly with the CTC based model for ASR and AED tasks. Additionally or alternatively, it is an object of some embodiments to use an attention based transformer model for AT tasks.

[0007] Some embodiments aim to analyze an audio scene to identify (e.g., detect and classify) audio events that form the audio scene. The detection and classification of audio events includes determining attributes of different types of audio events that an audio signal of the audio scene carries. The audio signal can carry multiple audio events. Examples of audio events include: speech events, including words spoken by a user; non-speech events, including various exclamations and non-human sound, such as regular household sounds (door closing, sink running, etc.), industrial processing sounds, or other sounds. Furthermore, the audio scene can include different types of audio events that occur simultaneously (i.e., overlap in time) or sequentially (i.e., do not overlap in time).

[0008] The attributes of different types of audio events define metadata of the audio events that form the audio scene. In other words, the metadata includes attributes of the audio events in the audio signal. For example, in a car repair shop, the audio scene can include a drill sound, and an attribute of the drill sound can be an identification label “drill”. Additionally or alternatively, the audio scene in the car repair shop can include a voice command by a human voice to enable a diagnostic tool. Thus, the same audio scene can also include a speech event, and a corresponding attribute can be an identification of the speech event as a voice command, a transcription of the voice command, and / or an identification of the speaker. Additionally or alternatively, the audio scene in the car repair shop can include a conversation between a mechanic and a customer, and an attribute of the conversation can be a transcription of the conversation (i.e., non-command voice utterance). Additionally or alternatively, an attribute of an audio event can be a natural language sentence that describes the scene, such as “the mechanic talks to the customer before using the drill”.

[0009] Accordingly, the attributes of audio events can be time-dependent (e.g., automatic speech recognition or acoustic event detection) and / or time-agnostic (e.g., audio labeling of acoustic scenes (“car repair shop”), audio captioning of audio scenes, or other sound events in acoustic scenes (e.g., drill sounds, speakers, or any other speech / non-speech sound)). Thus, an audio signal can carry multiple audio events including speech events and non-speech events. Time-dependent attributes include one or a combination of transcription of speech, translation of speech, and detection of audio events with their temporal locations. Time-agnostic attributes include labels or captions of audio events. Additionally, attributes can have multiple levels of complexity such that audio events can be labeled differently. For example, an engine sound can be labeled roughly as engine or mechanical noise, or in more detail as car engine, bus engine, large engine, small engine, diesel engine, electric engine, accelerating engine, idling engine, knocking engine, etc., where multiple labels / attributes can be active at the same time. Likewise, a speech event can be labeled as speech / no speech, female / male human voice, speaker ID, singing, screaming, shouting, angry, happy, sad, etc. Additionally or alternatively, automatic speech recognition (ASR) transcription can be considered as an attribute of a speech event.

[0010] Some implementations are based on the recognition that the complexity of audio scenes erases the boundaries between different transcription tasks such as automatic speech recognition (ASR), acoustic event detection (AED), and audio labeling (AT). ASR is an artificial intelligence and linguistics field that involves transforming audio data associated with speech into text that represents the speech. AED involves detecting audio events including speech and non-speech sounds, such as regular household sounds, sounds in a car repair shop, or other sounds present in an audio scene. Additionally, AED involves detecting the temporal locations of these audio events. Further, AT provides label labels for audio events, where only the presence of audio events is detected in an audio signal. Some implementations are based on the recognition that the ASR, AED, and AT tasks can be performed using task-specific neural networks, respectively. Some implementations are based on the recognition that these task-specific neural networks can be combined to unify ASR, AED, and AT to enable synergy in performing ASR, AED, and AT.

[0011] However, in this approach, there is a training data sparsity problem in each transcription task because these task-specific neural networks cannot take advantage of the fact that sound events of other tasks can have similar sound event characteristics. Some implementations are based on the recognition that viewing transcription tasks as different types of attribute estimation of audio events allows for the design of a single mechanism that aims to perform transcription tasks on an audio signal of an audio scene regardless of the complexity of the audio scene.

[0012] Some implementations are based on the recognition that a single neural network can be jointly trained to perform one or more transcription tasks. In other words, by jointly training a neural network (NN) for ASR, AED, and AT, synergy can be achieved in performing ASR, AED, and AT on an audio signal. According to implementations, the neural network includes a transformer model and a connectionist temporal classification (CTC) based model, where the CTC based model shares at least some model parameters with the transformer model. Such a neural network can be used to jointly perform ASR, AED, and AT, i.e., the neural network can simultaneously transcribe speech, identify audio events occurring in an audio scene, and generate an audio caption for the audio scene. To this end, the neural network can be utilized to process an audio signal to determine different types of attributes of audio events to generate metadata for the audio signal. In addition, using such a neural network (or achieving synergy) eliminates the training data sparsity problem in individual transcription tasks, in addition to providing more accurate results. Furthermore, achieving synergy allows for generating customized audio outputs, i.e., allows for generating desired acoustic information from an audio signal.

[0013] According to implementations, the models of the neural network share at least some parameters for determining time-dependent and time-agnostic attributes of speech events and non-speech audio events. The models of the neural network include an encoder and a decoder. In implementations, the parameters shared for determining different types of attributes are parameters of the encoder. In alternative implementations, the parameters shared for determining different types of attributes are parameters of the decoder. While the neural network is jointly trained for transcription tasks, some parameters (e.g., weights of the neural network) are reused for performing the transcription tasks. Reusing some parameters of the neural network for such joint training of the neural network requires less training data to train individual transcription tasks, allows for using weakly-labeled training data, and produces accurate results in individual tasks even with a small amount of training data.

[0014] Some implementations are based on the recognition that the neural network can be configured to selectively perform one or more of the ASR, AED, and AT transcription tasks to output desired attributes of audio events. According to implementations, the output of the transformer model depends on an initial state of a decoder of the transformer model. In other words, the initial state of the decoder determines whether the decoder will output according to an ASR, AT, or AED task. To this end, some implementations are based on the recognition that the initial state of the decoder can be varied based on a desired task to be performed to generate desired attributes.

[0015] Some implementations are based on the recognition that neural networks based on transducer models with an encoder-decoder architecture can be used to perform AED and AT tasks. Neural networks based on an encoder-decoder architecture provide a decisive advantage. For example, in AED and AT tasks, the decoder of the encoder-decoder architecture directly outputs the symbols (i.e., labels). Thus, with the encoder-decoder architecture, the cumbersome process of setting detection thresholds for individual classes during inference, which AED and AT systems often use, is eliminated. In addition, neural networks based on an encoder-decoder architecture do not require monotonic ordering of labels, so weakly labeled audio recordings (without annotations with temporal or sequential information) can be easily used to train the neural network. However, AED and ASR tasks require temporal information of the audio signal. Some implementations are based on the recognition that transducer models can be augmented with a connectionist temporal classification (CTC) based model, where some neural network parameters are shared between the two models to leverage the temporal information of the audio signal. Furthermore, neural networks with transducer models and CTC based models enforce monotonic ordering and learn temporal alignment. To this end, neural networks with transducer model outputs and CTC based model outputs can be used to jointly perform ASR, AED, and AT transcription tasks for generating metadata of the audio signal.

[0016] According to an embodiment, the model of the neural network comprises a transducer model and a CTC based model. The transducer model comprises an encoder and a decoder. The encoder is configured to encode the audio signal and provide the encoded audio signal to the decoder. The CTC based model is configured to process the encoded audio signal of the encoder to generate a CTC output. Since ASR and AED tasks are temporal information dependent tasks, the transducer model and the CTC based model are jointly used to perform the ASR and AED tasks. To this end, to jointly perform the ASR and AED tasks, the encoded audio signal is processed with the decoder to perform ASR decoding and AED decoding. Furthermore, the encoded audio signal is processed with the CTC based model to generate the CTC output. The CTC output is combined with the output of the ASR decoding to generate a transcription of the speech event. Similarly, the CTC output is combined with the output of the AED decoding to generate a label of the audio event.

[0017] The neural network is further configured to perform a temporal dependent AT task, i.e., the temporal information of the audio event is not explicitly determined. For the AT task, the decoder is configured to perform AT decoding. Furthermore, the audio signal is labeled based on the AT decoding. To this end, the neural network can perform both temporal dependent (ASR and AED) and temporal independent tasks (AT).

[0018] Some implementations are based on the recognition that, in order to perform time- dependent tasks (ASR and AED), the transducer model is jointly trained with a CTC-based model to exploit the monotonic alignment property of CTC. Thus, the ASR and AED tasks are jointly trained and decoded with the CTC-based model, while the AT task is trained with the transducer model only.

[0019] In some implementations, the CTC output is also used to compute the time position of detected acoustic events for AED decoding (e.g., using CTC-based forced alignment) so that the start time and end time of an audio event is estimated.

[0020] During training of the neural network, a weight factor is used to balance the loss of the transducer model and the loss of the CTC-based model. The weight factor is assigned to the training samples of ASR, AED, and AT, respectively. According to the implementation, the neural network is trained for jointly performing the ASR, AED, and AT tasks using a multi-objective loss function that includes the weight factor. Thus, an embodiment discloses an audio processing system. The audio processing system includes an input interface configured to accept an audio signal. Further, the audio processing system includes a memory configured to store a neural network trained to determine different types of properties of a plurality of concurrent audio events of different origins, wherein the types of properties include time-dependent properties and time-agnostic properties of speech audio events and non-speech audio events, and wherein models of the neural network share at least some parameters for determining both types of properties. Further, the audio processing system includes a processor configured to process the audio signal with the neural network to generate metadata of the audio signal including one or more properties of one or more audio events in the audio signal. Further, the audio processing system includes an output interface configured to output the metadata of the audio signal.

[0021] Thus, another embodiment discloses an audio processing method. The audio processing method includes accepting, via an input interface, an audio signal; determining, via a neural network, different types of properties of a plurality of concurrent audio events of different origins in the audio signal, wherein the different types of properties include time-dependent properties and time-agnostic properties of speech audio events and non-speech audio events, and wherein models of the neural network share at least some parameters for determining both types of properties; processing, via a processor, the audio signal with the neural network to generate metadata of the audio signal including one or more properties of one or more audio events in the audio signal; and outputting, via an output interface, the metadata of the audio signal.

[0022] The presently disclosed embodiments will be further described with reference to the drawings. The drawings shown are not necessarily to scale, emphasis generally being placed upon illustrating the principles of the presently disclosed embodiments. BRIEF DESCRIPTION OF DRAWINGS

[0023] [ Figure 1A ] Figure 1A An audio scene of a car repair shop is shown according to some embodiments.

[0024] [ Figure 1B ] Figure 1B A schematic diagram showing the principle of an audio scene analysis transform used by some embodiments is shown.

[0025] [ Figure 2A ] Figure 2A A schematic diagram showing a combination of different transcription tasks that can be performed by a neural network to generate different types of attributes according to some embodiments is shown.

[0026] [ Figure 2B ] Figure 2B A schematic diagram showing a combination of automatic speech recognition (ASR) and acoustic event detection (AED) transcription tasks that can be performed by a neural network to generate different types of attributes according to some embodiments is shown.

[0027] [ Figure 2C ] Figure 2C A schematic diagram showing a combination of automatic speech recognition (ASR) and audio tagging (AT) transcription tasks that can be performed by a neural network to generate different types of attributes according to some embodiments is shown.

[0028] [ Figure 3 ] Figure 3 A schematic diagram showing a model of a neural network comprising an encoder-decoder architecture according to some embodiments is shown.

[0029] [ Figure 4 ] Figure 4 A schematic diagram showing a model of a neural network based on a transformer model with an encoder-decoder architecture according to some embodiments is shown.

[0030] [ Figure 5 ] Figure 5 A schematic diagram showing a model of a neural network comprising a transformer model and a model based on connectionist temporal classification (CTC) according to some embodiments is shown.

[0031] [ Figure 6 ] Figure 6 A schematic diagram showing a model of a neural network with a state switcher according to some embodiments is shown.

[0032] [ Figure 7 ] Figure 7 A schematic diagram showing training of a neural network for performing ASR, AED or AT on an audio signal according to some embodiments is shown.

[0033] [ Figure 8 ] Figure 8A block diagram illustrating an audio processing system for generating metadata of an audio signal according to some embodiments.

[0034] [ Figure 9 ] Figure 9 A flowchart illustrating an audio processing method for generating metadata of an audio signal according to some embodiments.

[0035] [ Figure 10 ] Figure 10 Illustrating utilization of an audio processing system to analyze a scene according to some embodiments.

[0036] [ Figure 11 ] Figure 11 Illustrating anomaly detection by an audio processing system according to example embodiments.

[0037] [ Figure 12 ] Figure 12 Illustrating a collaborative operating system using an audio processing system according to some embodiments. DETAILED DESCRIPTION

[0038] In the following description, for the purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. It will be apparent, however, to one skilled in the art that the present disclosure can be practiced without some or all of these specific details. In other instances, devices or methods are shown in block diagram form in order to avoid obscuring the present disclosure.

[0039] As used in this specification and claims, the terms “for example,” “e.g.,” and “such as,” and the verbs “comprising,” “having,” “including,” and their other verb forms, when used to connect the

[0040] Computer auditory (CA) or machine listening is a general area of research for algorithms and systems for audio understanding by machines. Since the notion of what it means for a machine to “listen” is quite broad and somewhat vague, computer auditory tries to bring together multiple disciplines that originally should address specific problems or have a concrete application in mind.

[0041] Similar to computer vision, computer auditory aims to analyze an audio scene to identify (e.g., detect and classify) the audio objects that form the audio scene. In the context of machine listening, such detection and classification includes determining the attributes of the audio events carried by the audio signals measured by the sensors that “listen” to the audio scene. These attributes define the metadata of the audio events that form the audio scene.

[0042] Figure 1A An audio scene of an automotive repair shop 100 is shown in accordance with some embodiments. Figure 1B A schematic diagram showing the principle of an audio scene analysis transform used by some embodiments. Figure 1A and Figure 1B are described in conjunction with each other.

[0043] The audio scene of the automotive repair shop 100 can include multiple concurrent audio events of different origins, such as the sound of a drill operated by the mechanic 102, a voice command to a human voice to enable a diagnostic tool by the mechanic 104, the sound of a person 106 running, a conversation between the mechanic 108 and the customer 110, the sound of the engine 112, etc. Additionally or alternatively, the audio scene can include audio events of different types that occur sequentially (i.e., not overlapping in time). Some embodiments aim to analyze the audio scene to identify (e.g., detect and classify) the audio events that form the audio scene. The detection and classification of the audio events includes determining attributes of different types of the audio events.

[0044] For example, the audio scene of the automotive repair shop 100 can include a non-speech event, such as the sound of a drill operation, and an attribute of the drill sound can be an identification label “drill”, “hydraulic drill”, “electric drill”, “vibrating drill”, etc. Additionally or alternatively, the audio scene of the automotive repair shop 100 can include a speech event, such as a voice command to a human voice to enable a diagnostic tool, and a corresponding attribute can be an identification of the speech event as a voice command, a transcription of the voice command, and / or an identification of the speaker. Additionally or alternatively, for the audio event of a conversation between the mechanic 108 and the customer 110, an attribute of the conversation can be a transcription of the conversation, i.e., a non-command voice utterance. Additionally or alternatively, an attribute of an audio event can be a natural language sentence that describes the audio scene. For example, an audio caption “the mechanic talks to the customer before using the drill” that describes the audio scene can be an attribute of the audio event.

[0045] Accordingly, different types of attributes of audio events include time-dependent attributes and time-agnostic attributes of speech audio events and non-speech audio events. Examples of time-dependent attributes include one or a combination of transcription of speech and detection of temporal location of audio events. Examples of time-agnostic attributes include labels of audio events and / or audio captions of audio scenes. Additionally, attributes can have multiple levels of complexity such that audio events can be labeled differently. For example, the sound of the engine 112 can be labeled roughly as engine or mechanical noise, or more detailed as car engine, bus engine, large engine, small engine, diesel engine, electric engine, accelerating engine, idling engine, knocking engine, etc., where multiple labels / attributes can be active at the same time. Likewise, speech events can be labeled as speech / non-speech, female / male human voice, speaker ID, singing, screaming, shouting, etc.

[0046] Some implementations are based on the understanding that a fragmentation analysis 114 of an audio scene can be performed to determine different types of attributes of audio events 122. In the fragmentation analysis 114 of an audio scene, each of the different types of audio events 116 is considered individually. Accordingly, the determination of attributes of individual audio events is treated as a separate problem. For example, determining a transcription of speech of the mechanic 108 is treated as a separate problem, and determining attributes of the sound of the engine 112 is treated as a separate problem. Furthermore, the determination of each of the different types of attributes of audio events is treated as a separate problem. For example, determining a transcription of speech of the mechanic 108 is treated as a separate problem, and determining a label (e.g., speaker ID) of the mechanic 108 is treated as a separate problem. To this end, there are different audio problems 118 in determining different types of attributes of audio events 122.

[0047] Some implementations are based on the recognition that different attribute-specific audio solutions 120 can be utilized for different audio problems 118 in determining different types of attributes of audio events 122. For example, a neural network trained only for automatic speech transcription can be used to determine a transcription of speech of the mechanic 108, and a neural network trained only for audio event detection can be used to detect audio events such as the sound of the engine 112, the sound of the person 106 running, etc. Additionally, audio labeling techniques can be used separately to determine labels of detected audio events. Some implementations are based on the recognition that these task-specific neural networks can be combined to unify transcription tasks (e.g., automatic speech transcription, audio event detection, etc.) for performing different types of transcription tasks (e.g., automatic speech transcription, audio event detection, etc.). However, in this approach, there can be a training data sparsity problem in individual transcription tasks because these task-specific neural networks cannot take advantage of the fact that audio events of other tasks can have similar audio event characteristics.

[0048] Further, some implementations are based on the understanding that a neural network trained for one type of transcription task can also facilitate performance of another transcription task. For example, a neural network trained only for automatic speech transcription can be used for different speech events. For example, a neural network trained only for automatic speech transcription can be used to transcribe human voice commands and conversations between the repair technician 108 and the customer 110. Some implementations are also based on the understanding that different attribute-specific audio resolutions can be applied to an audio event to determine different attributes of the audio event. For example, automatic speech transcription techniques and audio event techniques can be applied to a human voice command that enables a diagnostic tool to determine corresponding attributes, such as identifying the speech event as a human voice command, transcription of the human voice command, and identification of the speaker. To this end, attribute-specific audio resolutions can be used for similar audio events to determine the same type of attributes, and different attribute-specific audio resolutions 120 can be used for an audio event to determine different types of attributes of the audio event. However, different attribute-specific audio resolutions 120 cannot be used for different types of audio events 116 to determine different attributes of the audio events 122.

[0049] Some implementations are based on the recognition that different types of audio events can be treated uniformly (or as similar audio events) under the assumption that metadata of the audio events is determined at different transcription tasks, regardless of the type of the audio events and the type of the attributes of the audio events. The metadata includes attributes that describe the audio events. The attributes can have different types, such as time-dependent attributes and time-agnostic attributes, but regardless of their types, the attributes are merely descriptions that form the metadata of the audio events. In this way, determination of different types of attributes of the audio events can be treated as a single audio problem. Since only a single audio problem is considered, a single resolution can be formulated for determining different types of attributes of the audio events, regardless of the complexity of the audio scene. This recognition allows the fragmented analysis 114 to be transformed into a uniform analysis 124 of the audio scene. According to implementations, the single resolution can correspond to a neural network trained to determine different types of attributes of multiple concurrent and / or sequential audio events of different origins.

[0050] To this end, in the uniform analysis 124 of the audio scene according to some implementations, an audio signal 126 carrying multiple concurrent audio events 126 is input to a neural network 128 trained to determine different types of attributes of the multiple concurrent audio events 126. The neural network 128 outputs different types of attributes, such as time-dependent attributes 130 and time-agnostic attributes 132. One or more attributes of one or more audio events in the audio signal 126 are referred to as metadata of the audio signal 126. Thus, the audio signal 126 can be processed with the neural network 128 to determine different types of attributes of the audio events to generate the metadata of the audio signal 126.

[0051] In contrast to the fragmented analysis of the scene 114 with multiple neural networks trained to perform different transcription tasks, the single neural network 128 trained to perform the uniform analysis 124 has a model that shares at least some parameters for determining different types of attributes 130 and 132 to transcribe different audio events 126. In this way, the training and execution of a single solution (neural network 128) is synergized. In other words, in contrast to using different attribute-specific audio solutions to generate different types of attributes, a single neural network (i.e., neural network 128) can be used to generate different types of attributes and can be trained with less data than would be required to train separate neural networks that perform transcription of the same task. To this end, synergy can be achieved in determining different types of attributes of audio events. In addition, achieving synergy reduces the training data sparsity problem in individual transcription tasks, in addition to providing more accurate results. Furthermore, achieving synergy allows for the generation of customized audio outputs, i.e., allows for the generation of desired acoustic information from the audio signal 126.

[0052] Figure 2A A schematic diagram illustrating a combination of different transcription tasks that can be performed by the neural network 128 to generate different types of attributes 204, according to some embodiments, is shown. According to embodiments, the neural network 128 can perform one or a combination of different transcription tasks such as automatic speech recognition (ASR), acoustic event detection (AED), and / or audio tagging (AT). ASR is a field of artificial intelligence and linguistics that involves transforming audio data associated with speech into text that represents the speech. AED involves detecting audio events with corresponding timing information, e.g., detecting regular household sounds, sounds in an office environment, or other sounds present in an audio scene, and the time start and end locations of each detected audio event. Furthermore, AT provides label tags for individual audio events or sets of audio events of an audio scene without the need to detect explicit start and end locations or time ordering of audio events in an audio signal. Additionally or alternatively, AT can include audio captioning, where natural language sentences are determined that describe an audio scene. Such audio captions can be attributes of audio events. For example, the audio caption “the repairman talks to the customer before using the drill” describing an audio scene can be an attribute of an audio event. Another example of an audio captioning attribute can be “the repairman asks the customer if they want to pay with a credit card or cash while operating the drill,” where identifying the speech content and audio events of an audio scene requires synergy between ASR and AT transcription tasks.

[0053] An audio signal 200 of an audio scene is input to the neural network 128. The audio signal 200 can carry a plurality of audio events including speech events and non-speech events. The neural network 128 can perform a combination 202 of ASR 200a, AED 200b, and AT 200c transcription tasks. According to some implementations, the neural network 128 can jointly perform the ASR 200a, AED 200b, and AT 200c transcription tasks to determine different types of attributes 204 of the plurality of audio events. The different types of attributes 204 include speech attributes of the speech events and non-speech attributes of the non-speech events. The speech attributes are time-dependent attributes and include speech transcription in the speech events, and the non-speech attributes are time-agnostic attributes and include labels of the non-speech events.

[0054] In alternative implementations, the neural network 128 can jointly perform the ASR 200a, AED 200b, and AT 200c transcription tasks to determine different types of attributes 204 including time-dependent attributes and time-agnostic attributes of the plurality of audio events. The time-dependent attributes include one or a combination of transcription of the speech and detection of time locations of the plurality of audio events, and the time-agnostic attributes include labeling of the audio signal with one or more of labels or audio captions describing the audio scene using natural language sentences.

[0055] Figure 2B A schematic diagram illustrating a combination 206 of ASR and AED transcription tasks that can be performed by the neural network 128 to generate different types of attributes 208 according to some implementations is shown. According to some implementations, the neural network 128 can jointly perform the ASR 200a and AED 200b transcription tasks to determine different types of attributes 208 of the plurality of audio events. The different types of attributes 208 include transcription of the speech in the speech events and detection of time locations of the plurality of audio events. In other words, the neural network 128 can simultaneously transcribe the speech and recognize the plurality of audio events with corresponding timing information, i.e., time start and end locations of each of the plurality of audio events.

[0056] Figure 2C A schematic diagram illustrating a combination 210 of ASR and AT transcription tasks that can be performed by the neural network 128 to generate different types of attributes 212 according to some implementations is shown. According to some implementations, the neural network 128 can jointly perform the ASR 200a and AT 200c transcription tasks to determine different types of attributes 212 of the plurality of audio events. The different types of attributes 212 include one or a combination of detection of the speech events, transcription of the speech in the speech events, labels of the audio signal, and audio captions describing the audio scene using natural language sentences.

[0057] According to some implementations, the model of the neural network 128 shares at least some parameters for determining different types of attributes. In other words, the neural network 128 shares at least some parameters for performing different transcription tasks to determine different types of attributes. Some implementations are based on the recognition that a neural network 128 that shares some parameters for performing different transcription tasks to determine different types of attributes can be better aligned with a single human auditory system. Specifically, in the auditory pathway, an audio signal goes through multiple processing stages, whereby early stages extract and analyze different acoustic cues, while final stages in the auditory cortex are responsible for perception. This processing is similar in many ways to an encoder-decoder neural network architecture, where an encoder extracts important acoustic cues for a given transcription task, an attention mechanism acts as a relay, and a decoder performs perception. To this end, the model of the neural network includes an encoder-decoder architecture.

[0058] Figure 3 A schematic diagram illustrating a model of a neural network including an encoder-decoder architecture 302 according to some implementations is shown. The encoder-decoder architecture 302 includes an encoder 304 and a decoder 306. An audio signal 300 carrying multiple audio events of different origins is input to the encoder 304. The neural network including the encoder-decoder architecture 302 shares some parameters for performing different transcription tasks to determine different types of attributes 308. In implementations, the parameters shared for determining different types of attributes are parameters of the encoder 304. In alternative implementations, the parameters shared for determining different types of attributes are parameters of the decoder 306. In some other implementations, the parameters shared for determining different types of attributes are parameters of the encoder 304 and the decoder 306. According to implementations, the shared parameters correspond to weights of the neural network.

[0059] While the neural network is jointly trained for different transcription tasks, some parameters (e.g., weights of the neural network) are reused for performing the transcription tasks. Reusing some parameters of the neural network for this joint training of the neural network requires less training data to train individual transcription tasks, allows the use of weakly-labeled training data, and produces accurate results in individual tasks even with a small amount of training data.

[0060] Some implementations are based on the recognition that a neural network based on a transformer model with an encoder-decoder architecture can be used to perform AED and AT transcription tasks. Figure 4 A schematic diagram illustrating a model of the neural network 128 based on a transformer model 400 with an encoder-decoder architecture according to some implementations is shown. An audio signal 402 is input to a feature extraction 404. The feature extraction 404 is configured to obtain different acoustic features, such as spectral energy, power, pitch, and channel information, from the audio signal 402.

[0061] The transducer model 400 includes an encoder 406 and a decoder 408. The encoder 406 of the transducer model 400 is configured to encode the audio signal 402 and provide the encoded audio signal to the decoder 408. Further, for the AED transcription task, the decoder 408 is configured to process the encoded audio signal to perform AED decoding to detect and recognize a plurality of audio events present in the audio signal 402 without determining temporal information of the detected audio events. Further, for the AT task, the decoder 408 is configured to perform AT decoding. The audio signal 402 is labeled with tags based on the AT decoding. Additionally, the decoder 408 can provide an audio caption to the audio signal 402.

[0062] The neural network 128 based on the encoder-decoder architecture provides decisive advantages. For example, in the AED and AT tasks, the decoder 408 of the encoder-decoder architecture directly outputs the symbols (i.e., the tags). Thus, the utilization of the encoder-decoder architecture eliminates the cumbersome process of setting detection thresholds for individual classes during inference, which would otherwise be used by the AED and AT systems. Additionally, the neural network based on the encoder-decoder architecture does not require monotonic ordering of the tags, and thus can easily utilize weakly-labeled audio recordings (annotated without temporal or sequential information) to train the neural network 128.

[0063] However, the AED and ASR transcription tasks require temporal information of the audio signal. Some implementations are based on the recognition that the transducer model 400 can utilize a connectionist temporal classification (CTC) based model enhancement to utilize the temporal information of the audio signal 300.

[0064] Figure 5 A schematic diagram illustrating the model of the neural network 128 including the transducer model 400 and a CTC based model 504 is shown, in accordance with some implementations. The CTC based model 504 corresponds to one or more additional layers added to the encoder 406 trained with a CTC objective function, in accordance with implementations. Further, the neural network 128 with the transducer model 400 and the CTC based model 504 enforces monotonic ordering and learns temporal alignment. To this end, the neural network 128 with the transducer model 400 and the CTC based model 504 can be used to jointly perform the ASR, AED, and AT tasks. The model of the neural network 128 with the transducer model 400 and the CTC based model 504 can be referred to as an all-in-one (AIO) transducer.

[0065] The audio signals 500 of the audio scene are input to a feature extraction 502. The feature extraction 502 is configured to obtain different acoustic features, e.g. spectral energy, power, pitch and / or channel information, from the audio signals 500. The encoder 406 of the transducer model 400 encodes the audio signals 500 and provides the encoded audio signals to the decoder. The CTC-based model 504 is configured to process the encoded audio signals to generate a CTC output. Since ASR and AED are time information dependent tasks, the transducer model 400 and the CTC-based model 504 are jointly used to perform ASR and AED. To this end, in order to jointly perform the ASR and AED tasks, the encoded audio signals are processed with the decoder 408 to perform ASR decoding and AED decoding. Further, the encoded audio signals are processed with the CTC-based model 504 to generate a CTC output. The CTC output is combined with the output of the ASR decoding to generate a transcription of the speech events. Similarly, the CTC output is combined with the output of the AED decoding to generate a transcription of the sound events, i.e. the labels of the audio events. In some embodiments, the CTC output is also used to compute the temporal positions of the detected audio events for the AED decoding, e.g. using CTC-based forced alignment, such that the start and end times of the audio events are estimated.

[0066] The neural network 128 is further configured to perform a time-independent AT task, i.e. without explicitly determining the temporal information of the audio events. For the AT task, the decoder 408 is configured to perform AT decoding. Further, the audio signals are labeled based on the AT decoding. In particular, acoustic elements (sound events) associated with audio objects are labeled and / or an audio caption describing the audio scene. To this end, the neural network 128 can perform both time-dependent (ASR and AED) and time-independent tasks (AT).

[0067] To this end, some embodiments are based on the recognition that the neural network 128 can be configured to perform an ASR task to generate a transcription of the speech events in the audio signals 500. In addition, some embodiments are based on the recognition that the neural network 128 can be configured to jointly perform ASR and AED to generate labels and transcriptions of the audio events. Depending on the embodiment, the neural network 128 can be configured to jointly perform ASR, AED and AT transcription events to generate metadata of the audio signals 500.

[0068] Some implementations are based on the recognition that the neural network 128 can be configured to selectively perform one or more of the ASR, AED, and AT transcription tasks to output the desired attribute of the audio event. According to the implementations, the output of the transducer model 400 depends on the initial state of the decoder 408 of the transducer model 400. In other words, the initial state of the decoder 408 determines whether the decoder 408 will output according to ASR, AT, or AED. To this end, some implementations are based on the recognition that the initial state of the decoder 408 can vary based on the desired transcription task to be performed to generate the desired attribute. Accordingly, the model of the neural network 128 is provided with a state switcher.

[0069] Figure 6 A schematic diagram of the model of the neural network 128 with the state switcher 600 is shown according to some implementations. A mapping between the initial states and the different transcription tasks is provided. When the user inputs an input symbol indicating the desired transcription task, the state switcher 600 is configured to switch the initial state corresponding to the desired transcription task to perform the desired transcription task. Accordingly, the desired attribute of the audio event can be output by the model of the network 128.

[0070] Furthermore, Figure 6 An example output is shown by switching the initial feed of the task to the decoder 408 is shown in the angle brackets <asr> 、 <aed> 、 <at1> 、… <at7>) in the input sequence. The symbol <s> denotes a stop symbol for decoding, and the label suffixes S, E, and C denote start boundary and end boundary and continuation of a sound event. The ASR and AED are jointly performed with the CTC-based model, while the AT only uses the decoder output.

[0071] Figure 7 A schematic diagram illustrating training of the neural network 128 for performing ASR, AED, or AT on an audio signal according to some embodiments is shown. In block 700, the parameter settings of the AIO transformer are calibrated. For example, the parameter settings of the AIO transformer are d model = 256, d ff = 2048, d h = 4, E = 12, and D = 6. The training is performed using an Adam optimizer with β1= 0.9, β2= 0.98, and ∈ = 10 -9 for 25000 warm-up steps. In addition, the initial learning rate is set to 5.0, and the number of training epochs reaches 80.

[0072] Further, in block 702, a weight factor is assigned to the ASR sample set and the AED sample set, respectively, to balance the loss of the transformer model and the loss of the CTC-based model while training. For example, the CTC / decoder weight factor is set to 0.3 / 0.7 for the ASR sample set, to 0.4 / 0.6 for the AED sample set, and to 0.0 / 1.0 otherwise. The same weight factors are also used for decoding. In addition, in an alternative embodiment, a weight factor is assigned to the AT sample set. The weight factor is used to control the weighting between the transformer objective function and the CTC objective function. In other words, the weight factor is used to balance the transformer objective function and the CTC objective function during training. These samples assigned with the respective weight factors are used to train the neural network 128. For ASR inference, a neural network-based language model (LM) is applied via shallow fusion using a LM weight of 1.0. For the AED task, the time information of the recognized sound event sequence is obtained using CTC-based forced alignment.

[0073] At block 704, the transducer model is jointly trained with the CTC-based model to perform the ASR, AED, and AT transcription tasks. The ASR sample set, the AED sample set, and the AT sample set are used with the transducer model and the CTC-based model to train the neural network to jointly perform the ASR, AED, or AT transcription tasks. Time-independent tasks, such as the AT task, do not require temporal information. Thus, the AT sample set is used only with the transducer model to learn the AT transcription task. Some implementations are based on the recognition that, to perform the time-dependent tasks (ASR and AED), the transducer model is jointly trained with the CTC-based model to leverage the monotonic alignment property of CTC. Thus, the transducer model is jointly trained with the CTC-based model using the ASR sample set, the AED sample set to perform the ASR and AED transcription tasks.

[0074] The AIO transducer utilizes two different types of attention, namely, encoder-decoder attention and self-attention. The encoder-decoder attention uses the decoder state as the query vector to control the attention to the input value sequence and the encoder state sequence. In self-attention (SA), the query, value, and key are computed from the same input sequence, which results in an output sequence of the same length as the input. The two attention types of the AIO transducer are based on the scaled dot-product attention mechanism

[0075]

[0076] where and are the query, key, and value, where d * represents the dimension, n * represents the sequence length, d q = d k , and n k = n v . Instead of using a single attention head, each layer of the AIO transducer model uses multiple attention heads, where

[0077]

[0078]

[0079] where and are the inputs to the multi-head attention (MHA) layer, Head i represents the output of the i-th attention head out of a total of d h heads, and are trainable weight matrices that satisfy d k = d v = d model / d h .

[0080] The encoder of the AIO transformer comprises two layers of CNN modules ENCCNN and E layers of transformer encoder layers with self-attention stacked in ENCSA:

[0081] X0= ENCCNN(X), (4)

[0082] X E = ENCSA(X0), (5)

[0083] where X = (x1,...,x T ) denotes the acoustic input feature sequence, which is an 80-dimensional log-mel-spectral energy (LMSE) plus three additional features for pitch information. The two CNN layers of ENCCNN use a stride of size 2, a kernel size of 3x3, and a ReLU activation function, which reduces the frame rate of the output sequence X0by a factor of 4. The ENCSA module (5) consists of E layers, where for e = 1,..., E, the e-th layer is a multi-head self-attention layer combined with two ReLU separated feed-forward neural networks of inner dimension d ff and outer dimension d model :

[0084] X′ e = X e-1 + MHA e (X e-1 , X e-1 , X e-1 ), (6)

[0085] X e = X′ e + FF e (X′ e ), (7)

[0086]

[0087] where and are trainable weight matrices and bias vectors.

[0088] The transformer objective function is defined as

[0089]

[0090] where the label sequence Y = (y1,...,y L ), the label subsequence y 1:l-1 = (y1,...,y l-1 ), and the encoder output sequence X E . The term p(y l |y 1:l-1 , X E ) denotes the transducer decoder model, which can be written as

[0091] p(y l |y 1:l-1 , X E ) = DEC(X E , y 1:l-1) , (10)

[0092] where

[0093]

[0094]

[0095]

[0096]

[0097] for d = 1,..., D, where D denotes the number of decoder layers. The function EMBED maps the input label sequence (y <s> θ y1,..., y l-1 ) into a sequence of trainable embedding vectors where <s> θ ∈Θ represents the sequence Θ = ( <asr> , <aed> , <at1> ,…… <at7>The task-specific start symbol (or input symbol) for indexing, such as Figure 3 As shown. The function DEC, through... A fully connected neural network 128 is applied to the softmax distribution on the output to finally predict the label y. l The posterior probability. Sine position codes are added to sequences X0 and Z0.

[0098] For ASR and AED tasks, the transformer model is jointly trained with the CTC objective function.

[0099]

[0100] Where B represents a one-to-many mapping that expands the label sequence Y to the set of all possible frame-level label sequences using the CTC transition rule. π represents the frame-level label sequence. Multi-objective loss function.

[0101]

[0102] This is used to train neural network 128, where the hyperparameter γ controls two objective functions p. ctc and p att Weighted average between them.

[0103] For joint decoding, some implementations use a CTC-based model p. ctc (Y|X E ) and attention-based decoder model p att (Y|X E The decoding target is defined by the sequence probability of the label sequence in order to find the most likely label sequence.

[0104]

[0105] Where λ represents the weighting factor balancing the probabilities of CTC and attention-based decoders, and where p ctc (Y|X E The CTC prefix decoding algorithm can be used to calculate this.

[0106] Figure 8 A block diagram of an audio processing system 800 for generating metadata for audio signals, according to some embodiments, is shown. The audio processing system 800 includes an input interface 802. The input interface 802 is configured to receive audio signals. In an alternative embodiment, the input interface 802 is also configured to receive input symbols indicating a desired transcription task.

[0107] The audio processing system 800 can have multiple interfaces that connect the audio processing system 800 with other systems and devices. For example, a network interface controller (NIC) 814 is adapted to connect the audio processing system 800 to a network 816 through the bus 812, which connects the audio processing system 800 with an operatively connected set of sensors. Through the network 816, the audio processing system 800 receives audio signals wirelessly or through a wire.

[0108] The audio processing system 800 includes a processor 804 configured to execute instructions stored in a memory 806, and the memory 806 stores the instructions executable by the processor 804. The processor 804 can be a single core processor, a multi-core processor, a computing cluster, or any number of other configurations. The memory 806 can include random access memory (RAM), read only memory (ROM), flash memory, or any other suitable memory systems. The processor 804 connects to one or more input devices and output devices through the bus 812.

[0109] According to some embodiments, the instructions stored in the memory 806 implement a method for generating metadata about audio signals received via the input interface 802. To this end, the storage 808 can be adapted to store different modules that store executable instructions for the processor 804. The storage 808 can be implemented using a hard disk drive, an optical drive, a thumb drive, a drive array, or any combination thereof.

[0110] The storage 808 is configured to store parameters of a neural network 810 trained to determine different types of attributes of multiple concurrent audio events of different origins. The different types of attributes include time-dependent attributes and time-agnostic attributes of speech events and non-speech audio events. The models of the neural network 810 share at least some parameters for determining attributes of both types. The models of the neural network 810 include a transducer model and a connectionist temporal classification (CTC) based model. The transducer model includes an encoder configured to encode the audio signals and a decoder configured to perform ASR decoding, AED decoding, and AT decoding for the encoded audio signals. The CTC based model is configured to perform ASR decoding and AED decoding for the encoded audio signals to generate CTC outputs. According to embodiments, the transducer model is jointly trained with the CTC model to perform ASR and AED transcription tasks. The storage 808 is further configured to store a state switcher 824 configured to switch an initial state of the decoder according to input symbols to perform a desired transcription task.

[0111] In some implementations, the processor 804 of the audio processing system 800 is configured to process the audio signal with the neural network 810 to generate metadata of the audio signal. The processor 804 is further configured to process the audio signal with an encoder of the neural network 810 to generate an encoding, and process the encoding multiple times with a decoder initialized to different states corresponding to different types of attributes to generate different decodings of attributes of different audio events. Additionally, in alternative implementations, the processor 804 is further configured to switch decoder states according to input symbols to perform a desired transcription task, and generate an output of the desired task using the neural network 810. The generated output is part of the multi-level information.

[0112] Further, the audio processing system 800 includes an output interface 820. In some implementations, the audio processing system 800 is further configured to submit the metadata of the audio signal to a display device 822 via the output interface 820. Examples of the display device 822 include a computer monitor, a camera, a television, a projector, or a mobile device, among others. In implementations, the audio processing system 800 can also be connected to an application interface suitable for connecting the audio processing system 800 to external devices for performing various tasks.

[0113] Figure 9 A flowchart of an audio processing method 900 for generating metadata of an audio signal according to some implementations of the present disclosure is shown. At block 902, the audio processing method 900 includes accepting an audio signal. In implementations, the audio signal is accepted via the input interface 802.

[0114] Further, at block 904, the audio processing method 900 includes processing the audio signal with the neural network 810 to generate metadata of the audio signal. The metadata includes one or more attributes of one or more audio events in the audio signal. The one or more attributes include time-dependent attributes and time-agnostic attributes of speech events and non-speech audio events.

[0115] Further, at block 906, the audio processing method 900 includes outputting the metadata of the audio signal. In implementations, the multi-level information is outputted via the output interface 820.

[0116] Figure 10 An auditory analysis of a scene 1000 using the audio processing system 800 is shown in accordance with some embodiments. The scene 1000 includes one or more audio events. For example, the scene 1000 includes audio events such as speech of a person 1002 moving in a wheelchair 1004, a sound of a cat 1006, an entertainment device 1008 playing music, and footsteps of a second person 1012. An audio signal 1010 of the scene 1000 is captured via one or more microphones (not shown in the figure). The one or more microphones can be placed at one or more suitable places in the scene 1000 such that they capture the audio signal including the audio events present in the scene 1000.

[0117] The audio processing system 800 is configured to accept the audio signal 1010. The audio processing system 800 is further configured to perform ASR, AED, or AT tasks on the audio signal 1010 using the neural network 128 to generate attributes associated with the audio events in the audio signal. For example, the audio processing system 800 can generate a speech transcription of the person 1002. Further, the audio processing system 800 can identify sound events in the scene 1000 such as the moving wheelchair 1004, the speech of the person 1002, the sound of the cat 1006, the music played in the entertainment device 1008, the footsteps of the second person 1012, and the like.

[0118] Additionally, in accordance with embodiments, the audio processing system 800 can provide labels for the speech of the person 1002 such as male / female voice, singing, and speaker ID. These different types of attributes generated by the audio processing system 800 can be referred to as metadata of the audio signal 1010. The metadata of the audio signal 1010 can be further used to analyze the scene 1000. For example, the attributes can be used to determine various activities taking place in the scene 1000. Similarly, the attributes can be used to determine various audio events taking place in the scene 1000.

[0119] Additionally or alternatively, in accordance with embodiments of the present disclosure, the audio processing system 800 can be used in one or more of an in-car infotainment system including a voice search interface and a hands-free phone, a voice interface for an elevator, a service robot, and factory automation.

[0120] Figure 11 Anomaly detection of the audio processing system 800 in accordance with example embodiments is shown. In Figure 11 In particular, a scenario 1100 is shown that includes a manufacturing production line 1102, a training data pool 1104, a machine learning model 1106, and an audio processing system 800. The manufacturing production line 1102 includes a plurality of engines that work together to manufacture a product. In addition, the production line 1102 uses sensors to collect data. The sensors can be digital sensors, analog sensors, and combinations thereof. The collected data is used for two purposes, some of the data is stored in the training data pool 1104 and used as training data to train the machine learning model 1106, and some of the data is used as operational time data for the audio processing system 800 to detect anomalies. Both the machine learning model 1106 and the audio processing system 800 can use the same data.

[0121] To detect anomalies in the manufacturing production line 1102, training data is collected. The machine learning model 1106 uses the training data in the training data pool 1104 to train the neural network 810. The training data pool 1104 can include labeled data or unlabeled data. Labeled data is labeled with a label (e.g., anomaly or normal), and unlabeled data is without a label. Based on the type of training data, the machine learning model 1106 applies different training methods to detect anomalies. For labeled training data, supervised learning is typically used, and for unlabeled training data, unsupervised learning is applied. In this way, different implementations can handle different types of data. In addition, detecting anomalies in the manufacturing production line 1102 includes detecting anomalies in each of the plurality of engines included in the manufacturing production line 1102.

[0122] The machine learning model 1106 learns features and patterns of the training data, which includes normal data patterns and anomaly data patterns associated with audio events. The audio processing system 800 uses the trained neural network 810 and collected operational time data 1108 to perform anomaly detection, where the operational time data 1108 can include a plurality of concurrent audio events associated with the plurality of engines.

[0123] On receiving the operating time data 1108, the audio processing system 800 can use the neural network 810 to determine the metadata of the audio events associated with the respective engines. The metadata of the audio events associated with the engines can include attributes such as accelerating engine, idling engine, knocking engine, banging engine, and the like. These attributes can enable the user to analyze the sound of the respective engine among the plurality of engines, thereby enabling the user to analyze the manufacturing line 1102 at a granular level. Further, the operating time data 1108 can be identified as normal or abnormal. For example, using the normal data patterns 1110 and 1112, the trained neural network 810 can classify the operating time data as normal data 1114 and abnormal data 1116. For example, the operating time data XI 1118 and X2 1120 are classified as normal, and the operating time data X3 1122 is classified as abnormal. Once the abnormality is detected, necessary actions 1124 are taken.

[0124] In particular, the audio processing system 800 uses the neural network 810 to determine at least one attribute of the audio events associated with the audio source (e.g., engine) among the plurality of audio sources. Further, the audio processing system 800 compares the at least one attribute of the audio events associated with the audio source with at least one predetermined attribute of the audio events associated with the audio source. Further, the audio processing system 800 determines the abnormality in the audio source based on the comparison result.

[0125] Figure 12 A collaborative operation system 1200 using the audio processing system 800 is shown, in accordance with some embodiments. The collaborative operation system 1200 can be arranged in a part of a product assembly / manufacturing line. The collaborative operation system 1200 includes the audio processing system 800 having the NIC 814 connected to a display 1202, a camera, a speaker, and an input device (microphone / pointing device) via a network. In this case, the network can be a wired network or a wireless network.

[0126] The NIC 814 of the audio processing system 800 can be configured to communicate with a robot 1206 such as a robotic arm via the network. The robot 1206 can include a robotic arm controller 1208 and a sub-robotic arm 1210 connected to a robotic arm state detector 1212, wherein the sub-robotic arm 1210 is configured to assemble a workpiece 1214 for manufacturing a product part or a complete product. Further, the NIC 814 can be connected to an object detector 1216 via the network. The object detector 1216 can be arranged to detect the state of the workpiece 1214, the sub-robotic arm 1210, and the robotic arm state detector 1212 connected to the robotic arm controller 1208 arranged in the robot 1206. The robotic arm state detector 1212 detects and sends a robotic arm state signal to the robotic arm controller 1208. The robotic arm controller 1208 then provides a process flow or instructions based on the robotic arm state signal.

[0127] Display 1202 can display a process flow or instructions representing the process steps of assembling a product based on a (pre-designed) manufacturing method. The manufacturing method can be received via a network and stored in memory 806 or storage device 808. For example, when operator 1204 inspects the condition of assembled product parts or assembled products (while performing quality control processing according to a format such as a process record format), audio input can be provided via the microphone of the collaborative operating system 1200 to record the quality inspection. The quality inspection can be performed based on the product manufacturing process and product specifications indicated on display 1202. Operator 1204 can also provide instructions to robot 1206 to perform operations on the product assembly line. Audio processing system 800 can perform an ASR transcription task on the audio input to generate a speech transcript of operator 1204. Alternatively, audio processing system 800 can jointly perform ASR and AED transcription tasks to generate a speech transcript of operator 1204 and determine attributes such as speaker ID and gender of operator 1204.

[0128] The collaborative operating system 1200 can use a voice-to-text program to transcribe the results confirmed by the operator 1204 and their corresponding data as text data into the memory 806 or storage device 808. Additionally, the collaborative operating system 1200 can store defined attributes. Furthermore, the results, along with the item numbers assigned to the various assembled parts or assembled products, can be stored with timestamps for product manufacturing records. Moreover, the collaborative operating system 1200 can transmit the records to the manufacturing center computer via a network. Figure 12 (Not shown in the image), which allows the entire process data of the assembly line to be integrated to maintain / record the quality of operators and products.

[0129] The following description provides exemplary embodiments only and is not intended to limit the scope, applicability, or configuration of this disclosure. Rather, the following description of exemplary embodiments will provide those skilled in the art with a feasible description for implementing one or more exemplary embodiments. Various changes to the function and arrangement of the elements will be conceived without departing from the spirit and scope of the subject matter set forth in the appended claims.

[0130] Specific details are set forth in the following description to provide a thorough understanding of the embodiments. However, it will be understood by those skilled in the art that embodiments may be practiced without these specific details. For example, systems, processes, and other elements in the disclosed subject matter may be shown as components in block diagram form to avoid obscuring the embodiments with unnecessary detail. In other instances, well-known processes, structures, and techniques may be shown without unnecessary detail to avoid obscuring the embodiments. Furthermore, similar reference numerals and designations in the various figures indicate similar elements.

[0131] Additionally, various embodiments can be implemented as a process that is depicted as a flowchart, data flow diagram, structure diagram, or block diagram. Although a flowchart can describe operations as a sequential process, many of the operations can be performed in parallel or concurrently. In addition, the order of the operations can be re-arranged. A process is terminated when its operations are completed, but could be terminated without yet completing its operations due to, for example, a system shutdown. Operations described can be stored as computer-executable instructions in a non-transitory computer-readable medium such as a hard disk or a removable memory. Further, any of the described operations can be performed, at least in part, by specific hardware components that contain hardwired logic for conducting the operations.

[0132] Furthermore, embodiments of the subject matter disclosed can be implemented, at least in part, manually or automatically. Manual or automatic implementation can occur contemporaneously with the execution of software, hardware, firmware, middleware, microcode, hardware description languages, or any combination thereof. When implemented in software, firmware, middleware, or microcode, the program code or code segments to perform the necessary tasks can be stored in a machine readable medium. Processors(s) can execute the program code.

[0133] The various methods or processes outlined herein can be coded as software that is executable on one or more processors that employ any one of a variety of operating systems or platforms. Additionally, such software can be written using any of a number of suitable programming languages and / or programming or scripting tools, and also can be compiled as executable machine language code or intermediate code that is executed on a framework or virtual machine. Typically, the functionality of the program modules can be combined or distributed as desired in various embodiments.

[0134] Embodiments of the present disclosure can be embodied as a method, of which an example has been provided. The acts performed as part of the method can be ordered in any suitable way. Accordingly, embodiments can be constructed in which acts are performed in an order different than illustrated, which can include performing some acts simultaneously, even though shown as serial process acts in illustrative embodiments. The present disclosure is not limited to the illustrative embodiments described herein but can be modified within the spirit and scope of this disclosure. Accordingly, the aspects of the appended claims are intended to encompass all such changes and modifications that fall within the true spirit and scope of this disclosure. < / at1> < / aed> < / asr> < / s> < / s> < / at1> < / aed> < / asr>

Claims

1. An audio processing system, the audio processing system comprising: an input interface configured to receive an audio signal; a memory configured to store a neural network trained to determine different types of attributes of a plurality of concurrent audio events of different origins, wherein the attributes define metadata of the plurality of concurrent audio events of different origins, wherein the different types of attributes include time-dependent attributes and time-agnostic attributes of speech audio events and non-speech audio events, wherein the time-dependent attributes include a transcription of speech, the time-agnostic attributes include a label of an audio event and / or an audio caption of an audio scene, wherein a model of the neural network shares at least some parameters for determining both types of the attributes, wherein the model of the neural network includes an encoder and a decoder, and wherein the parameters shared for determining the different types of the attributes include parameters of the encoder; a processor configured to process the audio signal with the neural network to generate metadata of the audio signal, the metadata including one or more attributes of one or more audio events in the audio signal; and an output interface configured to output the metadata of the audio signal; wherein the neural network is jointly trained to perform a plurality of different transcription tasks using the shared parameters to perform individual transcription tasks; wherein the transcription tasks include an automatic speech recognition, ASR, task and an acoustic event detection, AED, task; wherein the model of the neural network includes a transformer model and a connectionist temporal classification, CTC, based model, wherein the transformer model includes an encoder configured to encode the audio signal and a decoder configured to perform ASR decoding, AED decoding, and AT decoding to generate a decoder output for the encoded audio signal, and wherein the CTC based model is configured to perform the ASR decoding and the AED decoding for the encoded audio signal to generate a CTC output for the encoded audio signal, and wherein the decoder output and the CTC output of the ASR decoding and the AED decoding are jointly scored to generate a joint decoding output.

2. The audio processing system of claim 1, wherein, the audio signal carries a plurality of audio events including speech events and non-speech events, and wherein the processor determines speech attributes of the speech events and non-speech attributes of the non-speech events using the neural network to generate the metadata.

3. The audio processing system of claim 1, wherein, the parameters shared for determining the different types of the attributes include parameters of the decoder.

4. The audio processing system of claim 1, wherein, the CTC based model is configured to generate temporal information of one or more of an ASR transcription task or an AED transcription task.

5. The audio processing system of claim 1, wherein, the transformer model is jointly trained with the CTC based model to perform an ASR transcription task and an AED transcription task.

6. The audio processing system of claim 1, wherein, the audio signal includes a plurality of audio events associated with a plurality of audio sources, and wherein the processor is further configured to: determine at least one attribute of at least one audio event of the plurality of audio sources using the neural network; comparing the at least one attribute of the at least one audio event to a predetermined at least one attribute of the at least one audio event; and determining an anomaly in the audio source based on a result of the comparing.

7. An audio processing method, the audio processing method comprising the steps of: accepting, via an input interface, an audio signal; determining, via a neural network, different types of attributes of different causes of multiple concurrent audio events in the audio signal, wherein the attributes define metadata of different causes of multiple concurrent audio events, wherein the different types of attributes comprise time-dependent attributes and time-agnostic attributes of speech audio events and non-speech audio events, wherein the time-dependent attributes comprise transcriptions of speech, the time-agnostic attributes comprise labels of audio events and / or audio captions of audio scenes, and wherein models of the neural network share at least some parameters for determining both types of the attributes, wherein the models of the neural network comprise an encoder and a decoder, and wherein the parameters shared for determining different types of the attributes comprise parameters of the encoder; processing, via a processor, the audio signal with the neural network to generate metadata of the audio signal, the metadata comprising one or more attributes of one or more audio events in the audio signal; and outputting, via an output interface, the metadata of the audio signal; wherein the neural network is jointly trained to perform multiple different transcription tasks using the shared parameters to perform individual transcription tasks; wherein the transcription tasks comprise automatic speech recognition (ASR) tasks and acoustic event detection (AED) tasks; wherein the models of the neural network comprise a transformer model and a connectionist temporal classification (CTC) based model, wherein the transformer model comprises an encoder configured to encode the audio signal and a decoder configured to perform ASR decoding, AED decoding, and AT decoding to generate decoder outputs for the encoded audio signal, and wherein the CTC based model is configured to perform the ASR decoding and the AED decoding for the encoded audio signal to generate CTC outputs for the encoded audio signal, and wherein the decoder outputs and the CTC outputs of the ASR decoding and the AED decoding are jointly scored to generate a joint decoding output.

Citation Information

Patent Citations

  • Method and device for generating interpretation data according to video data and performing data synthesis as well as electronic equipment

    CN107707931A